Four days after Kimi K3 was released, the storm is still gathering force.
Four days ago, in the early hours of July 17, the 2.8 trillion-parameter K3 arrived seemingly out of nowhere, topping the Frontend Code Arena leaderboard with a score of 1,679. It was the first open-source model to overtake a slate of closed-source models and claim the top spot on Frontend Code Arena’s front-end coding ranking. Tesla CEO Elon Musk commented “Impressive” under a related benchmark. It was the second time Musk had publicly recognized the Chinese team, after liking Kimi’s technical report on its underlying architecture in March. When previewing his new model, Grok 4.6, he even used Kimi K3 as a comparison target.
But the praise came with turbulence. Forty-eight hours after launch, Kimi said user requests had far exceeded its estimates and were approaching the capacity limits of its existing cluster. Kimi announced it would pause new consumer subscriptions, prompting some industry observers to compare the episode to a “DeepSeek 2.0 moment.”
At the same time, doubts began to surface. On one side, claims circulated that K3 was a distilled model. On the other, Dean Ball, OpenAI’s head of strategic futures, publicly criticized Kimi K3. He acknowledged that K3’s performance could not be achieved through distillation, but argued that open-weight models weaken the commercial returns of frontier models, calling them a “decelerationist force.”
Questions over technical originality, the open-source versus closed-source path, and the sustainability of premium pricing have all landed at Moonshot AI. Huang Zhenxin, the company’s head of enterprise business, faced the media and gave the company’s first systematic response to outside criticism.
Denying That K3 Is a Distilled Small Model
The Performance Leap Comes From Original Architecture Innovation at the Base Layer
Responding to speculation that “K3 is a distilled small model,” Huang Zhenxin, head of enterprise business at Moonshot AI, directly denied the claim in an interview with reporters on the morning of July 21. Huang said K3’s performance leap comes from original architectural innovation at the base layer, not from distilling or replicating existing models.
When Kimi launched the model, it had also pointed to efficiency gains from architectural changes. According to the company, Kimi K3 is built on Kimi Delta Attention, a hybrid linear attention mechanism, and Attention Residuals, or AttnRes. At the MoE, or mixture-of-experts, layer, the model activates 16 of 896 experts. Those structural improvements lifted Kimi K3’s overall scaling efficiency by about 2.5 times compared with K2.
Huang further disclosed three key technologies underpinning K3. Moon Clip is a new second-order optimizer, first applied by Kimi to large model training. “The world’s available training data has basically bottomed out. This technology can make 20T of training data deliver the training effect of 40T, directly halving training cost and computing power consumption at the same performance level.” Kimi Linear Tension is the company’s self-developed linear attention mechanism, designed to address the performance degradation of traditional linear attention in ultra-long tasks. “There is an AI Moore’s Law in the industry: the length of tasks AI can execute doubles every seven months. This mechanism expands the context window by 10 times, while training cost only increases by the same 10 times.”
Attention Residuals, a technology that drew attention when its technical report was released in March, optimizes how information is transmitted and reused across the model’s multilayer network. It allows the 2.8 trillion-parameter model to complete training stably while improving inference efficiency by 25%.
Since the start of the year, Moonshot AI founder Yang Zhilin has repeatedly discussed in public speeches Kimi’s strategy for building at scale: token efficiency, long context, and Agent clusters, all aimed at maximizing intelligence under limited resources. Judging by Kimi’s current technical iteration, the team remains firmly on its established path, with a focus on productivity scenarios such as coding, finance, law, and scientific research.
Huang also responded to the issue of model hallucinations. He said hallucinations cannot be completely eliminated. “Hallucination is essentially a product of the same source as model creativity,” he said, but K3 has sharply reduced its occurrence rate. Its Agent Swarm architecture, equipped with fact-checker sub-agents, can verify reports and industry data one by one, further lowering the probability of content distortion.
On the debate between open-source and closed-source models, Huang’s view was clear: the two are not opposing forms of competition, but products of differentiated enterprise demand. Companies that need privatized deployment and fine-tuning tend to prefer open source, while those seeking stable hosted services choose closed source. Whether a model is open or closed source has no necessary link to its token capability or output quality.
He also made clear that Kimi has stuck with an open-source path since K2, and that this direction will not change. Its core underlying technologies and papers are public, and full model weights are planned for release before July 27, 2026.
User Requests Far Exceeded Forecasts
Kimi, Under Computing Power Pressure, Has to Make Trade-Offs
K3’s popularity after launch exceeded everyone’s expectations, including Moonshot AI’s. Huang said the team had prepared a traffic plan in advance, but the actual surge in visits was far beyond what it had anticipated.
Forty-eight hours after launch, late on July 19, Kimi issued an emergency notice saying it faced an unexpected computing power challenge. User requests had far exceeded its forecast and were approaching the capacity limit of its existing cluster.
Huang said the company faced a dilemma at the time. Opening registration to all new users would seriously hurt the experience of existing paying users. Prioritizing paying customers required temporarily limiting new registrations. The company chose the latter. “We would rather temporarily give up new revenue than lose the baseline experience of paying customers.”
On July 19, Moonshot AI formally announced that it would pause new consumer subscriptions, put all existing computing power into serving current subscribers, and push ahead at full speed with computing power expansion. As new computing power comes online, it will gradually open more subscription slots until normal subscriptions are fully restored. According to Kimi’s official reply, existing plans have not been removed; current users can use them normally and renew automatically. Existing users can also choose to upgrade their plans.
National Business Daily reporters also sent the company questions on the timeline for computing power expansion and the scale of additional computing power, but had received no response as of publication.
On commercialization, Huang said Kimi does not currently offer privatized deployment services. Its business model is benchmarked against mature leading overseas AI companies, with its core effort focused on improving the model’s foundational capabilities. He also said the company will not stop at the current 2.8 trillion-parameter scale, and that the parameter scale of large models will continue to expand.
Alongside the model’s capabilities, its rising price has also drawn attention. According to estimates from third-party firm Artificial Analysis, Kimi K3 costs about $0.94 per single task, close to GPT-5.6 Sol’s $1.04 and roughly half of Claude Opus 4.8’s $1.80. Its pricing has pushed K3 into the same price band as leading overseas models. “Open-source models and Chinese large models should not be labeled cheap. We have built a SOTA, or state-of-the-art, model, and we can match it with reasonable commercial pricing,” Huang said.
The data also supports that view. On July 18, Moonshot AI president Zhang Yutong posted a chart on social media showing Kimi’s latest enterprise ARR, or annualized revenue. The image showed that after Kimi K3 was released on July 17, enterprise ARR grew severalfold in a short period.
Image source: screenshot from Zhang Yutong’s Xiaohongshu account
National Business Daily reporters learned that Moonshot AI’s ARR had already surpassed $300 million in mid-June.
Around the same time, Huang said in a public speech that Kimi’s overseas paying users had grown 400%, API revenue had grown 400%, its products had entered more than 200 countries and regions, and sectors including internet, finance, manufacturing, education, and healthcare were becoming important sources of enterprise customers.
Discussing the current state of the large model industry, Huang also said the sector does have a development bubble. He cited leading overseas company Anthropic, which has already achieved quarterly profitability, and said Moonshot AI is working to catch up in that direction. “In the end, we still hope to explore the upper limit of intelligence, and hope we can go head-to-head with the three overseas model companies.”
Amid the storm, another piece of news is also gaining traction. Market sources told National Business Daily that Moonshot AI, a leading large model company, has sent investors a listing proposal and could complete a Hong Kong Stock Exchange listing in as soon as six months, with a valuation possibly exceeding $30 billion. The company has not yet responded.
Moonshot AI founder Yang Zhilin. Image source: National Business Daily media asset library, file photo
The chain reaction triggered by K3 is far from over. Chinese large models are becoming a new value benchmark in the global market. On July 18, Musk announced that xAI was training Grok 4.6 with a parameter scale of 2 trillion, saying it “will complete its first training run next week and may surpass Kimi.” Moonshot AI later responded on social media: “Welcome to the ‘2 trillion-plus’ club.” Huang also said this time: “Maybe this performance from K3 gave them a bit of a shock. Now we will wait and see. We hope they come out and go head-to-head with us.”
As an open-source model stands for the first time at the top of the world in coding capability and moves into the same price band as leading overseas closed-source products, the real test is only beginning. After shedding the “cheap” label, Kimi and all Chinese open-source players still need to keep exploring how to balance the openness of open source with reasonable commercial returns, and how to turn SOTA-level technical capability into sustainable pricing power.
Comments
00No comments yet. Be the first to weigh in.