Experts Question Claims That Anthropic’s Fable Alone Powered Moonshot’s Kimi K3
AI researchers dispute claims that Moonshot built Kimi K3 primarily by distilling Anthropic’s Fable, saying advanced training methods likely played a larger role.
Artificial intelligence researchers are pushing back against claims that Moonshot AI built its Kimi K3 large language model primarily by copying Anthropic’s Fable model through distillation, arguing that the model’s rapid development and advanced capabilities are unlikely to have been achieved through that technique alone.
The debate follows allegations from White House science advisor Michael Kratsios, who said Moonshot developed Kimi K3 by copying Anthropic’s Fable while also using advanced Nvidia chips that are not approved for export to China. Kratsios described the alleged activity as large-scale industrial distillation aimed at extracting proprietary U.S. technology but did not provide evidence supporting the claims. Moonshot also did not respond to questions regarding its model training process.
Kratsios’ comments echoed earlier remarks from Treasury Secretary Scott Bessent, who said U.S. officials had found “watermarks” from American large language models in several Chinese AI models. However, officials have not explained what those watermarks are or how they were identified.
Researchers question whether distillation explains Kimi K3
Several AI researchers believe the timeline alone makes it unlikely that Fable was the primary source of Kimi K3’s capabilities. Anthropic publicly released Fable on July 1, leaving only a short period before Kimi K3 appeared, which experts argue is insufficient for conducting large-scale distillation, retraining a model and releasing it publicly.
Braden Hancock, a researcher at the Laude Institute and co-founder of Snorkel AI, said he does not believe a model as capable as Kimi K3 could have been produced through straightforward distillation in such a limited timeframe.
"I don't think you get a model this strong and this quickly on the heels of Fable doing strictly distillation. There's just not even frankly time. Fable's only been publicly available since July 1st. You can't distill that much data, train a model, and release it in two weeks," Hancock said.
Nathan Lambert, an AI researcher at the Allen Institute for AI, also questioned whether supervised fine-tuning through distillation remains as effective as it once was. According to Lambert, Chinese AI developers have moved closer to the industry’s frontier, making reinforcement learning increasingly important compared with traditional supervised fine-tuning.
Advanced AI training requires more than fine-tuning
Distillation generally involves repeatedly querying a larger language model to generate training data for a smaller model. In some cases, researchers ask the model to explain its reasoning process, while in others they use prompt-and-response pairs to perform supervised fine-tuning.
Lambert said supervised fine-tuning mainly influences how a model behaves rather than dramatically expanding its capabilities. To reproduce the performance of advanced frontier models, developers would likely need extensive reinforcement learning, where AI systems repeatedly evaluate and improve responses over millions of training iterations.
Such reinforcement learning workloads also require enormous computing resources. Running millions of AI agents through commercial APIs would be extremely expensive. It could become a significant bottleneck because frontier models operate relatively slowly and may not deliver enough additional performance to justify the cost.
Distillation remains common across the AI industry.
Although researchers question whether distillation alone explains KimiK3’s performance, they acknowledge that the practice is common throughout the AI industry. Earlier this year, Anthropic accused Moonshot, DeepSeek and MiniMax of systematically extracting capabilities from its models, saying it had identified millions of interactions linked to those companies through IP addresses and other metadata. Anthropic alleged the activity reflected deliberate capability extraction rather than ordinary customer usage.
Industry participants have also acknowledged that model distillation extends beyond Chinese developers. Elon Musk testified earlier this year that his company SpaceXAI distilled OpenAI models while developing Grok, describing the practice as common across the sector. Researchers note that the distinction between distillation and generating synthetic training datasets can often be difficult to define.
Hancock argued that U.S. observers should avoid underestimating the technical expertise of Chinese AI teams, noting that Moonshot’s founders include experienced researchers capable of independently advancing frontier AI systems.
"In general, Americans are understating the technical expertise of these Chinese teams. These are legitimate researchers and engineers doing solid work. If American models ground to a halt, I think China's progress would slow, but would still continue. They're not just riding coattails here," Hancock said.
Chip access remains a separate concern.
Researchers also distinguish the debate over model distillation from allegations that Moonshot obtained restricted Nvidia Grace Blackwell GB300 chips or accessed GB300-equipped servers outside China. Advanced AI chips remain subject to U.S. export controls, although experts say black-market channels continue to exist.
Sam Bresnick, a research fellow at GeorgetownUniversity’ss Centre for Security and Emerging Technology, said stronger customer verification rules for data centres could improve oversight of companies conducting large AI training runs on advanced hardware. While the Biden administration proposed federal know-your-customer requirements for data centres in 2024, the proposal has not advanced further, even as exporters remain responsible for ensuring restricted chips are used only for approved purposes.
What's Your Reaction?
Like
0
Dislike
0
Love
0
Funny
0
Angry
0
Sad
0
Wow
0