Just repeating rumors (sorry, should have been clearer): GLM 5.3 is 5.2 with extra post training, their big upcoming model is 5.5. It kind of makes sense too; pretraining on a new architecture/size is expensive and it's natural for these companies to try to wring an extra minor version or two out of each one.
Pretty tired of every article on Chinese AI being framed entirely in terms of geopolitical plots and nothing else. From what I can tell, Chinese companies releasing open weights was hatched in its private sector ecosystem, one major factor being that the leadership at Deepseek consists of actual open source zealots. Deepseek releasing their models under MIT license then forced the hands of other Chinese companies (including Moonshot AI, which initially wanted to be closed). It wasn't some grand plot hatched in Beijing (even if the CCP belatedly gave their consent after the fact).
Gonna go out on a limb here and say that getting into AI was the best thing Zuckerberg ever did for humanity, specifically the decision to release Llama as open weight (albeit obnoxious license) early on. That was the thing that showed the whole world open weights could work. Even if modern open weight models aren't direct descendents of Llama, and even if Meta subsequently pivoted to closed weights, that's still something that shouldn't be forgotten.
The coding skills are pretty meh relative to its size, but apparently its image processing skills are turning out to be top notch.
Interesting... Google's Gemini, which similarly boasts strong multimodal capabilities, is also having problems being competitive on coding. Wonder if it's just a coincidence.
Weird how the article doesn't mention a key factor: the Chinese labs have really low headcount compared to their Western competitors, and organizationally they're much more focused on model building than sidequests. Deepseek had only a couple hundred employees until recently, and accordong to leaks only had one or two people maintaining their consumer facing app.
He took over a failing Dutch tech company and turned it around. Nexperia was on the path to bankruptcy, that's why it was on the market to be sold back in 2017. His company injected capital and made it profitable. Even last year, his parent company even announced a $200M expansion of Nexperia's Hamburg plant, which totally goes against the narrative that they were moving production out of Europe.
The Dutch govt is trying to spin this, but they have like 5 different storylines and none of them make sense.
Bringing tariffs down in the past literally took decades. Politically it's very hard, because the benefits of trade go to everyone and the benefits of tariffs go to special interests. As a rule, special interests tend to win out.
Because it's arbitrary discrimination against their exporters. This is putting aside the forced purchases and investments Trump is demanding of various countries.
All these TACO memes are wearing pretty thin. Trump has instituted a minimum of 10% tariffs on all trading partners, with scant prospect of them ever lowering in the future even after his presidency (once tariffs go up, it's very hard to bring them down because of the special interests that come to depend on them). He has strongarmed the EU, Japan, and other countries into accepting these permanently elevated tariff levels without retaliation. Only 3 countries have shown any sort of backbone against this: China, Canada, and Brazil (maybe India, but it's too soon to say). In all the other cases, Trump ain't the one chickening out, it's the other side that abjectly folded.
(It doesn't matter, by the way, if the Europeans intend to slow walk the investments and weapon purchases they promised to Trump, or whatever. That's copium. The point is that they bent to Trump's will, and once you cave to a bully, he'll be back for more.)
In some dimensions, current day LLMs are already superintelligent. They are extremely good knowledge retrieval engines that can far outperform traditional search engines, once you learn how properly to use them. No, they are not AGIs, because they're not sentient or self-motivated, but I'm not sure those are desirable or useful dimensions of intellect to work towards anyway.
The kneejerk reaction is gonna be "Meta bad", but it's actually a bit more complicated.
Whatever faults Meta has in other areas, it's been mostly a good player in the AI space. They're one of the major reasons we have strong open-weight AI models today. Mistral, another maker of open AI models and Europe's only significant player in AI, has also rejected this code of conduct. By contrast, OpenAI a.k.a. ClosedAI has committed to signing it, probably because they are the incumbents and they think the increased compliance costs will help kill off competitors.
Personally, I think the EU AI regulation efforts are a big missed opportunity. They should have been used to force a greater level of openness and interoperability in the industry. With the current framing, they're likely to end up entrenching big proprietary AI companies like OpenAI, without doing much to make them accountable at all, while also burying upstarts and open source projects under unsustainable compliance requirements.
The EU AI Act is the thing that imposes the big fines, and it's pretty big and complicated, so companies have complained that it's hard to know how to comply. So this voluntary code of conduct was released as a sample procedure for compliance, i.e. "if you do things this way, you (probably) won't get in trouble with regulators".
It's also worth noting that not all the complaints are unreasonable. For example, the code of conduct says that model makers are supposed to take measures to impose restrictions on end-users to prevent copyright infringement, but such usage restrictions are very problematic for open source projects (in some cases, usage restrictions can even disqualify a piece of software as FOSS).
Just repeating rumors (sorry, should have been clearer): GLM 5.3 is 5.2 with extra post training, their big upcoming model is 5.5. It kind of makes sense too; pretraining on a new architecture/size is expensive and it's natural for these companies to try to wring an extra minor version or two out of each one.
My crackpot theory is that a lot of Western tech bros grew up reading Tolkien, and now they all think of themselves as Sauron forging the One Ring
Pretty tired of every article on Chinese AI being framed entirely in terms of geopolitical plots and nothing else. From what I can tell, Chinese companies releasing open weights was hatched in its private sector ecosystem, one major factor being that the leadership at Deepseek consists of actual open source zealots. Deepseek releasing their models under MIT license then forced the hands of other Chinese companies (including Moonshot AI, which initially wanted to be closed). It wasn't some grand plot hatched in Beijing (even if the CCP belatedly gave their consent after the fact).
Gonna go out on a limb here and say that getting into AI was the best thing Zuckerberg ever did for humanity, specifically the decision to release Llama as open weight (albeit obnoxious license) early on. That was the thing that showed the whole world open weights could work. Even if modern open weight models aren't direct descendents of Llama, and even if Meta subsequently pivoted to closed weights, that's still something that shouldn't be forgotten.
Very likely the same architecture, just with more post training.
The coding skills are pretty meh relative to its size, but apparently its image processing skills are turning out to be top notch.
Interesting... Google's Gemini, which similarly boasts strong multimodal capabilities, is also having problems being competitive on coding. Wonder if it's just a coincidence.
Weird how the article doesn't mention a key factor: the Chinese labs have really low headcount compared to their Western competitors, and organizationally they're much more focused on model building than sidequests. Deepseek had only a couple hundred employees until recently, and accordong to leaks only had one or two people maintaining their consumer facing app.
Yeah, they are currently engaged in a major spat with Huawei, which is accusing them of price gouging HBM supplies. Sounds familiar...
CXMT is probably the most profitable tech company in China right now.
I feel like at this point we need shade on demand more.
That's Chinese social media.
He took over a failing Dutch tech company and turned it around. Nexperia was on the path to bankruptcy, that's why it was on the market to be sold back in 2017. His company injected capital and made it profitable. Even last year, his parent company even announced a $200M expansion of Nexperia's Hamburg plant, which totally goes against the narrative that they were moving production out of Europe.
The Dutch govt is trying to spin this, but they have like 5 different storylines and none of them make sense.
Bringing tariffs down in the past literally took decades. Politically it's very hard, because the benefits of trade go to everyone and the benefits of tariffs go to special interests. As a rule, special interests tend to win out.
Because it's arbitrary discrimination against their exporters. This is putting aside the forced purchases and investments Trump is demanding of various countries.
All these TACO memes are wearing pretty thin. Trump has instituted a minimum of 10% tariffs on all trading partners, with scant prospect of them ever lowering in the future even after his presidency (once tariffs go up, it's very hard to bring them down because of the special interests that come to depend on them). He has strongarmed the EU, Japan, and other countries into accepting these permanently elevated tariff levels without retaliation. Only 3 countries have shown any sort of backbone against this: China, Canada, and Brazil (maybe India, but it's too soon to say). In all the other cases, Trump ain't the one chickening out, it's the other side that abjectly folded.
(It doesn't matter, by the way, if the Europeans intend to slow walk the investments and weapon purchases they promised to Trump, or whatever. That's copium. The point is that they bent to Trump's will, and once you cave to a bully, he'll be back for more.)
That was the rice minister (really).
In some dimensions, current day LLMs are already superintelligent. They are extremely good knowledge retrieval engines that can far outperform traditional search engines, once you learn how properly to use them. No, they are not AGIs, because they're not sentient or self-motivated, but I'm not sure those are desirable or useful dimensions of intellect to work towards anyway.
Hydropower is literally good for the climate.
The kneejerk reaction is gonna be "Meta bad", but it's actually a bit more complicated.
Whatever faults Meta has in other areas, it's been mostly a good player in the AI space. They're one of the major reasons we have strong open-weight AI models today. Mistral, another maker of open AI models and Europe's only significant player in AI, has also rejected this code of conduct. By contrast, OpenAI a.k.a. ClosedAI has committed to signing it, probably because they are the incumbents and they think the increased compliance costs will help kill off competitors.
Personally, I think the EU AI regulation efforts are a big missed opportunity. They should have been used to force a greater level of openness and interoperability in the industry. With the current framing, they're likely to end up entrenching big proprietary AI companies like OpenAI, without doing much to make them accountable at all, while also burying upstarts and open source projects under unsustainable compliance requirements.
The EU AI Act is the thing that imposes the big fines, and it's pretty big and complicated, so companies have complained that it's hard to know how to comply. So this voluntary code of conduct was released as a sample procedure for compliance, i.e. "if you do things this way, you (probably) won't get in trouble with regulators".
It's also worth noting that not all the complaints are unreasonable. For example, the code of conduct says that model makers are supposed to take measures to impose restrictions on end-users to prevent copyright infringement, but such usage restrictions are very problematic for open source projects (in some cases, usage restrictions can even disqualify a piece of software as FOSS).