Because demand was so intense, Moonshot put Kimi K3 on pause less than a week after launch.

On the night of July 19, Moonshot said it was suspending new Kimi subscriptions because of tight computing power supply, and would reserve its limited computing power for existing subscribers.

On July 20, Moonshot told Jiemian News that the change only affects new subscriptions. Existing member benefits remain unchanged, old plans have not been canceled, and current subscribers can keep using the service and renew automatically under the original rules.

For users who had previously turned off auto-renewal manually, the company will gradually restore renewal access. Upgrade options for legacy users will also be opened step by step. New plans are still unavailable for purchase because of computing power constraints, and no timeline has been set for when they will return.

From the company’s response, this looks more like a rebalancing of the experience between new users and the existing base under limited computing power, rather than a change to the membership system.

After K3 launched, developer communities in China and abroad quickly filled with hands-on tests and reviews. Many compared it with international flagship models from Claude and GPT, and debates over inference costs, GPU resources and service stability began to replace benchmark scores as the AI crowd’s new focus.

How hot is K3?

After launch, K3 quickly became a focal point in global AI model discussions.

Artificial Analysis, an independent AI model evaluation firm, said in its latest Intelligence Index that Kimi K3 scored 57 overall, putting it third globally, behind Claude Fable 5 at 60 and GPT-5.6 Sol at 59, and ahead of Claude Opus 4.8 at 56.

At the same time, K3’s average cost per task was about $0.94, roughly on par with GPT-5.6 Sol and only about half of Claude Opus 4.8’s.

For the industry, the bigger point is that this is the first open-source model to break into the top three of Artificial Analysis’s overall ranking. For years, the consensus was that open-source models lagged closed models by roughly half a generation. K3 has narrowed that gap further.

Beyond the benchmark results, K3’s parameter scale also drew attention. It is Moonshot’s first MoE model at the trillion-parameter level, with a total parameter count of 2.8 trillion.

A person at a domestic AI chip maker told Jiemian News that since the start of this year, as overseas players such as Anthropic have rolled out larger models, competition among Chinese model developers has shifted back to large parameters and long context windows. Trillion-parameter models and million-token contexts are becoming key battlegrounds for leading models.

In his view, K3 went viral not only because of what it can do, but because it marked the first time a domestic open-source model reached this scale. “For the industry, 2.8T itself is a major signal,” he said.

The buzz quickly spilled into developer communities.

On X, Reddit, OpenRouter, Linux.do and V2EX, many developers shared early impressions almost immediately, with some plugging K3 into AI coding tools such as Cursor, Cline and OpenCode for testing. From code generation to agent tasks and tool calling, K3 became one of the most discussed models in the AI world over the past few days.

The conversation centered on three questions: Can it replace Claude for coding? Has it really reached the top international tier? And is it worth switching to?

Many developers said K3’s biggest strength is not chat, but long-horizon agent tasks and code generation. Some testers said that in complex code edits and tool-use tasks, K3 has already moved into the same competitive group as top international models such as Claude and GPT.

But others said that in real engineering work, the model, while strong, still misses details and needs human correction. It also remains unstable on tasks such as complex statistical reasoning.

Another thread of discussion focused on cost. Some users argued that although K3’s unit price is competitive, real-world usage may not be as cheap as expected because complex tasks consume more tokens, especially in long agent workflows.

K3’s momentum also spread quickly into the overseas AI scene.

Gavin Baker, founder of U.S. investment firm Atreides Management, called K3 an inflection point in AI development. He said high-quality open-source models are speeding up the spread of AI capabilities, and that the impact will not be limited to model vendors but will benefit the entire AI application ecosystem.

Researchers were more cautious.

Ethan Mollick, a Wharton professor at the University of Pennsylvania who has long studied how AI affects work, said on X that he had asked Kimi K3 to perform a complex statistical audit of one of his earlier academic studies, but the model “got many things wrong,” including using the wrong statistical methods.

Mollick said K3 is already very close to leading global models, but its reliability on complex professional tasks still needs more real-world validation and cannot be judged by benchmark scores alone.

That view also reflects a broader attitude in the AI community toward new models: K3 has won strong approval from developers for code generation and agent scenarios, but its stability and reliability on complex professional tasks still need sustained testing in real-world settings, not just top scores on leaderboards.

Computing power Is Still the Bottleneck

It was against that backdrop of fast-growing developer traffic and rising usage that Moonshot ran into another problem: computing power.

Over the past two years, the usual story in the large-model industry was how many GPUs it took to train a model, or how many parameters it reached, not that so many users showed up that subscriptions had to be paused.

In practice, training does not end the need for computing power. That is especially true as AI coding and agent use cases grow. Models now need to call tools continuously, read context and run multiple rounds of reasoning, which keeps GPUs occupied far longer than a standard chat session. The stronger the model, the longer the context, and the more often users call it, the more inference computing power it consumes.

“The biggest problem today is still the lack of computing power, especially high-quality inference computing power,” Chen Yuqin, investment director at Shanghai State-owned Capital Futeng Capital, told Jiemian News earlier. Even top model companies have started ending unlimited access and raising service prices, she said, all of which shows that inference capacity remains tight. “If the infrastructure problem is not solved, many application companies actually do not dare to push products to scale.”

“From a supply-demand perspective, domestic AI computing power has long been in short supply, and as leading models roll out one after another, that pressure will only intensify,” the chip maker source told Jiemian News.

He said companies such as Moonshot, Zhipu and MiniMax mainly obtain computing power in two ways: renting computing services from public cloud providers, or building their own computing clusters. But the first option is constrained by limited resources, while the second requires heavy upfront investment and a long buildout cycle, so model developers are always under some degree of supply pressure.

“The industry broadly believes that to support training and inference for trillion-parameter models, a 10,000-GPU cluster has already become the basic threshold,” he said. From 2024 to now, only a handful of vendors have truly had 10,000-GPU deployment capability. Many large cluster projects that have already been announced often take years to come online, so the supply-demand gap is likely to persist.

One detail during this year’s WAIC stood out to the domestic chip maker.

He told Jiemian News that when local chip vendors showcased supernode products, the two questions they were asked most often were: “Can it support a 10,000-GPU cluster?” and “Can it support training trillion-parameter models?”

Compared with earlier conversations about single-GPU performance and chip specs, the industry is now more focused on ultra-large-scale cluster capabilities and the infrastructure needed to support continuous training and inference for flagship models.

That shift also reflects a change in how large-model competition works.

In the past, model companies competed on parameter scale, training capability and benchmark results. Now that more models are moving into real production environments, inference efficiency, service stability, cost control and infrastructure strength are emerging as new dimensions of competition.

K3’s pause on new subscriptions may have been only a temporary resource adjustment, but it gave outsiders a clear look at how the large-model race has entered a new phase. Model capability determines whether a product can attract users. Inference computing power, infrastructure and ongoing service capacity determine whether it can support real-scale adoption.

Unauthorized reproduction prohibited