ignacyr

Being Second Best In the AGI Race Just Got Way Cheaper

Kimi K3 release was special.

Last week, on July 16th Kimi K3 was released by the Chinese company Moonshot AI. It took the AI world by a storm, and rightfully so. The model pushed the frontier of open-weight models (the weights have not been published as of July 26th 2026, they are scheduled to be released on July 27th), almost matching other SOTA models, while being available at a significantly lower cost than Fable 5. All of this was achieved among accusations of using a method that gathered a lot of attention recently - distillation.

While not officially disclosed, there are several cues for Kimi K3 being trained on outputs of Claude including K3 identifying itself as Claude 15% of time. People have gone on to compare the similarity of answers to the same prompts across models and have seen similar results. White House OSTP Director accused Moonshot AI of distilling Kimi K3 from Fable 5 , but did not provide any proof to the public. Back in February 2026 Anthropic claimed that Chinese AI companies distilled Claude to train their models. While not confirmed, it is plausible that K3 was trained on Claude's data. Kimi addressed the accusations pointing at the timeline, there were 37 days between the release of Fable 5, and Kimi K3, which the company claims to be too short to train a state of the art model.

This method is called distillation, because data is distilled from models who have learned it from other sources.

The method involves pre-training models on data which was sourced from other models instead of the internet, which would likely yield worse results. In the past, we expected that using synthetic data recursively lead to model collapse, an Ouroboros snake eating its own tail approach doesn't give good results when training language models. But those results focused on the data's recursive nature, and in the distillation as we think of it now we are dealing with curated synthetic data.

But a shift has taken place, and if you consider what is happening in reality it does make sense. If you train your models on data scraped from the internet, a huge part of it is already synthetic, i.e. created by bots. With state of the art language models reaching performance of PhD students or higher across multiple disciplines, it is easy to see why they will be of higher quality than an average post on X or Reddit.

This is unethical! We can't let this happen!

Many have raised concerns about the ethics of this tactic, which violates the company's terms of service. It is nothing short of piggybacking on frontier labs' research to train your models and releasing them for a fraction of the price feels especially unfair. While in an ideal world, this wouldn't happen, in reality this approach doesn't seem any more unethical than scraping the whole internet, without having any permission for it whatsoever. Lawsuits like The New York Times Co. v. Microsoft Corp. and OpenAI are possibly only by huge efforts and use of capital by the biggest companies, independent creators have no way of hoping to have any chances at such a lawsuit, which would protect the tech giants from profiting off their work.

More than anything else, this seems to be an incentives problem. Currently, companies providing best quality outputs for cheap have a great shot at leading the race, and this is made possible by significantly cutting the training costs by stealing other's results. Majority of the AI bubble doesn't care about the ethics and factors other than price and quality of output these days.

People were amazed by Kimi K3.

While claiming that K3 is the best open-weight model would be too optimistic, given the amount of different benchmarks out there, it is not far from the truth, with the performance across multiple metrics, like coding and placing high in the Artificial Analysis' LLM Leaderboard, behind only GPT and Claude models as of July 26th 2026. While this release gathered a lot of attention from the general public, I don't think the distillation point was covered enough. I believe it has a possibility to reshape the race towards AGI tremendously.

A day before K3 was released, Thinking Machines revealed Inkling.

This was celebrated as the best American open-weights model to be released as of recently. It was cheap, and claimed to be a generalist model being able to work with audio. Thinking Machines claims to have trained the model from scratch, and there doesn't seem to be any evidence suggesting that it could have been distilled. The hype around it quickly faded as K3 came into the picture, as its performance just wasn't that special.

State of open-weight models and how the dynamics of frontier labs have to change.

What didn't change is that to be the company with the best models publicly available is very costly. The novel thing is that being the company with the second-best model out there has gotten significantly cheaper1 with distillation, if you manage to train your models fast enough you can be only months to weeks behind the leading AI labs, at a fraction of the cost. This is especially true if the U.S. jurisdiction doesn't concern you too much, but even there some of the biggest tech companies, including Nvidia, Microsoft, IBM and OpenAI, have gathered to prevent banning open-weights models. In other words, the incentives to be the first might diminish, because it is so much cheaper to be trailing. The costs do not sound like a real concern for the biggest companies in the world, but if you consider the pressure they face to produce returns for their capital investors, it is a real issue. Companies like Meta, Oracle, Amazon, Microsoft and Alphabet (owner of Google) have been accused of hiding their debt from the balance sheet to pour more money into investment.

What's coming next?

We would expect to see stronger guardrails embedded into models that protect against distillation. Frontier companies have already put a lot of effort into preventing the usage of their models to develop other models using them. So far, these approaches proved to be a double edged sword, as the models with stronger guardrails were deemed less usable. For example Fable 5 had guardrails that often "routed" its users to a less capable and more protected Opus 4.8 when it believed the users are trying to do something illegal, like perform cyber attacks or develop bioweapons. While the intent was sound, the application was dubbed underdeveloped, as it prevented developers from using the model's capabilities to perform any serious work in cybersecurity and other use cases the model flagged as dangerous. In the days of distilled models which do not exert so much effort in terms of safety, people can simply switch to somewhat less capable but more usable models like Kimi K3. This is once again a problem of incentives forming in a bad way, people are more likely to use the dangerous models because they enable them to do the safe work they are trying to do. This is bad news for AI Safety.

The AGI race might finally slow down.

It might look like companies will take more time to develop appropriate guardrails before releasing their frontier models, or else they will feed their competitor's growth. Like in competitive road cycling, the rider in the front exerts the most effort and is unlikely to finish first, it is more likely for the people trailing just behind him to overtake at a crucial moment and win instead. While judging whether the race actually slows down will take at least a few months, Anthropic releasing Opus 5 just two days ago and claiming it to be less restrictive and more affordable than Fable 5 signals that the labs notice their shortcomings.

Another plausible scenario is that the gap between internal models (with fewer guardrails) and external ones, available to the public (with guardrails on) might grow bigger, and we won't learn about it until a scandals like the Hugging Face OpenAI hacking come around. This is, once again, not a good thing for the overall safety, as external auditing will be severely limited if not impossible for models which are gatekept within the AI labs.

Footnotes: Leave a comment, I read every one!

  1. Inferred based on lower usage costs.