“Hangzhou speed” is becoming an increasingly weighty phrase.

A pure inference GPU company, just over a year after being spun out as an independent business, has already completed seven funding rounds totaling 4 billion yuan.

With its latest round disclosed a few days ago, it also secured 1 billion yuan, the largest single financing in the sector this year, making it China’s first pure inference GPU unicorn valued at more than 10 billion yuan.

That company is Xiwang.

Inference is becoming the new battleground in the AI computing power race. At this year’s NVIDIA GTC, NVIDIA directly showcased the LPU through its acquired company Groq. In China, domestic versions of Groq are also moving fast.

At a time when almost every Chinese GPU company is racing to build chips that combine training and inference while competing on peak computing power, why has Xiwang’s all-in bet on inference attracted such strong interest from capital markets?

With that question in mind, QbitAI sat down for an in-depth conversation with Xiwang co-CEO Wang Zhan.

Wang, an industry veteran from Baidu’s founding team who lived through the full arc of China’s internet industry from bubble to boom, not only cut to the core logic behind investors’ enthusiasm, but also laid out a clear blueprint for the AI inference era across four dimensions: industry trends, technical architecture, team organization and future judgment.

The Structure of Computing Power Demand Has Reversed

Wind the clock back a year or two, when the “battle of the hundred models” was in full swing. The market’s main concerns were model parameter counts and the scale of training clusters. But now, in 2026, the wind has completely shifted.

At the start of the interview, Wang Zhan set the tone this way:

Whoever controls the lowest inference cost will be the winner.

The essence of agents is that AI is no longer confined to a question-and-answer chatbot. It is meant to become an intelligent entity capable of autonomous analysis, learning and execution of complex tasks.

The underlying fuel that keeps all of this running is inference computing power, or more plainly, tokens.

That has created a major inflection point for the industry: a structural reversal in computing power demand.

The hottest demand in the market is for inference computing power, and it is growing exponentially. Demand for training computing power remains stable, but in the data we are currently seeing, demand for AI inference compute across 2026 will reach four to five times the demand for training computing power.

This is the first time inference computing power has comprehensively surpassed training computing power, and it has done so at remarkable speed.

Why has this reversal happened? The answer lies in how agents operate.

In the past, human-AI interaction was a single conversation. In the agent era, to complete a task, an agent makes frequent, repeated multi-round calls and engages in looping reasoning.

It is like the overseas user a few days ago who simply said “Hi” to a lobster and burned through $80 worth of tokens.

On this point, Wang Zhan stressed:

This mode pushes total token consumption to dozens or even hundreds of times the level of past human-machine interaction. Against this backdrop, the cost per token becomes highly visible.

In other words, companies used to care whether large models could be used at all. Now their biggest concerns are whether they work well and whether they are affordable to use.

That also explains why, from NVIDIA emphasizing “tokens per watt” at GTC to Chinese cloud providers adjusting computing power prices under cost pressure, cost has become a central force pushing technology forward.

In Wang Zhan’s view, lowering costs is not only a commercial demand, but also a prerequisite for mass adoption:

Only when the cost per token falls sharply can massive agent usage truly be activated. Otherwise, no matter how useful this thing is, if it is extremely expensive to use, people still cannot afford it.

And that is exactly why Xiwang made the decisive choice from the start to go all in on inference: inference is the real industrialization of AI.

One Cent per Million Tokens: How Is It Done?

If going all in on inference is the direction, then truly driving down costs at the technical level is the ultimate test of a team’s engineering capability and supply-chain insight.

When asked by customers who need both training and inference, Xiwang’s position is clear:

General-purpose GPUs are very good for large-cluster training, but in large-scale inference scenarios their price-performance ratio is often insufficient. In addition, as agents become widely adopted, inference computing power must withstand high-frequency calls with extremely low latency, support extreme stability for long-context workloads, and keep reducing the cost per token. Apart from a small number of special scenarios that do not consider commercial returns, inference GPUs have a stronger cost-performance advantage from a normal commercialization perspective.

After market development validated its strategic foresight, Xiwang revealed its core card: a new-generation inference GPU chip, Qiwang S3.

This is not just a performance upgrade. It is a system-level reconstruction of the AI inference cost curve: abandon training capability and natively customize the chip deeply for large-model inference. By trimming modules needed for training states, Xiwang redirects the saved transistor and power budgets into inference, increasing effective computing power efficiency per unit area by more than fivefold. Xiwang has set an ambitious target for S3: bringing the cost of one million tokens down to one cent.

To address pain points in the agent era, including surging KV cache demands, complex control flow and multi-model collaboration, S3 makes sweeping architectural changes.

First is deep customization at the compute layer.

General-purpose GPUs often face the awkward problem of underutilized computing power. S3’s inference-native AI Core architecture pushes utilization of core operators such as GEMM and Flash Attention to about 99% and 98%, respectively. S3 also natively supports full-chain low-precision computing from FP16 to FP4, multiplying throughput while keeping model performance close to lossless.

Second is bold innovation at the system layer, with two China firsts designed specifically for long context and agents:

S3 is China’s first inference GPU to use LPDDR6, while also supporting LPDDR5X. Its video memory capacity can reach nearly 600GB, making it the domestic GPU with the largest memory capacity. It is also the first released Chinese GPU to use PCIe Gen6, doubling system communication bandwidth.

Together, these two technologies address the bottleneck of long-context memory: S3 can store more users’ conversation histories at the same time, handle longer contexts, run faster and sharply reduce costs.

Wang Zhan explained it this way: Our goal is very clear: cut the cost per token by 90% and build inclusive inference computing power.

Of course, getting LPDDR6 and PCIe Gen6, two of the industry’s most advanced technologies, tuned, deployed and running at very high performance is no easy task. It depends heavily on full-stack in-house development and exceptional engineering capability.

Wang Zhan said proudly that Xiwang’s hardware AI Core and full software stack are 100% self-developed.

For a GPU to truly deliver performance, it must be balanced. You cannot be extremely strong in one place while a bottleneck sits in the middle. It is precisely because we have full-stack in-house capability that we can deeply tune and optimize around LPDDR6 and PCIe Gen6, and truly squeeze out their performance.

But while insisting on underlying independence and controllability, Xiwang has not closed itself off. Instead, it has achieved more than 99% compatibility with the CUDA ecosystem.

To outsiders, independence and controllability may seem naturally at odds with CUDA compatibility. In Wang Zhan’s view, this is entirely a matter of technical path.

We chose a general-purpose computing architecture, a GPU, rather than a dedicated architecture, an ASIC. A general-purpose architecture ensures very strong adaptability to different customer needs and different agents. On that basis, we write our own low-level code to support compatibility with the CUDA ecosystem. This gives customers the convenience of zero migration cost while preserving our underlying independence and controllability. The two are not contradictory.

Xiwang has maintained first-pass tape-out success and bring-up for every generation of its chips.

Behind that is an extremely large and low-profile verification team working quietly. According to people familiar with the matter, Xiwang’s team independently developed a full suite of simulation and verification tools. Before the chip is actually sent for tape-out, a massive number of operators have already been run on the simulation platform, giving the team a clear understanding of where bottlenecks are and how to fix them.

Hexagonal Warriors and a Trinity

Behind any breakout financing story, the most important asset is always people.

In the conversation with Wang Zhan, one can strongly sense the adrenaline-fueled excitement he feels when coming to work every day. That excitement comes from being part of an intensely aligned and formidable team.

Xiwang’s top-level structure has been jokingly called a “trinity” within the industry:

Chairman Xu Bing, co-founder of SenseTime: responsible for strategic direction and financing, with strong insight into AI development trends;

Co-CEO Wang Yong, formerly of AMD and a core architect at Kunlunxin: focused on chip R&D, with more than 20 years of hard-core semiconductor experience and widely seen as the company’s technical soul;

Co-CEO Wang Zhan, former senior vice president at Baidu: in charge of commercialization, operations and market strategy, bringing the sharp instincts and product playbook of a major internet company into hard tech.

But building AI infrastructure cannot rely on three people alone. As Wang Zhan put it:

Competition in AI chips is an all-around contest, like the all-around event in gymnastics. You have to be good at rings, parallel bars and everything else. No single person can be strong in every area. We have to rely on strong organizational management to bring outstanding people together and build our network of hexagonal warriors.

Xiwang now has more than 400 employees, with R&D staff accounting for over 80%. Its core technical leaders come from major companies including NVIDIA, AMD, Huawei HiSilicon, Alibaba and SenseTime, and have an average of more than 15 years of industry experience.

To retain these top-tier hexagonal warriors, Xiwang has made a concession rarely seen among Chinese startups. Wang Zhan disclosed a striking detail to QbitAI:

Among all Chinese GPU companies, we have given our team and employees the largest ESOP pool.

When Xu Bing brought me in, he said he wanted to set aside the largest ESOP pool to recruit the best talent. As long as we get this done, the value of the talent will be enormous.

This sharing mechanism, reminiscent of early Huawei and Alibaba, has unleashed strong organizational combat power.

Are Agents a Bubble or an Industrial Revolution?

After securing a valuation above 10 billion yuan and more than 1 billion yuan in financing, Wang Zhan, who personally experienced the bursting of the internet bubble in 2000, appears both clear-eyed and resolute amid the current AI capital boom.

Valuations for hard tech in today’s primary and secondary markets are indeed very optimistic. It is not just chip companies. Look at the valuation-to-revenue ratios of those large-model companies; they are truly exaggerated. When facing a once-in-an-era technological breakthrough, capital is willing to bet and take risks. That is the nature of capital.

But this time, AI is fundamentally different from the internet bubble of that era.

Wang Zhan recalled that when the internet was being hyped loudly in 2000, China had only a few million internet users. Even after ten years of development, the number of PC internet users was only a little over 100 million. Penetration took a long time.

But what about AI? After ChatGPT came out, it quickly became the fastest application in human history to reach 100 million users. And it is not like Zibo barbecue, where people try it once for novelty and leave. Over the past few years, user numbers have risen rapidly, and the more people use it, the more they depend on it.

Wang Zhan believes the fundamental value underlying AI is rising faster than in any previous industrial revolution in human history.

If the Industrial Revolution took a century and the information revolution took 20 or 30 years, then the AI intelligence revolution may compress sweeping social change into just a few years. In this era, something may look like a very large bubble last month and a smaller bubble next month, because the underlying value is rapidly filling in those valuations.

As for the computing power market in the second half of this year and further out, Wang Zhan’s judgment can be summed up in four Chinese characters: demand far exceeds supply.

The fundamental constraint on the growth of computing power is not market demand, but production tools. Optical modules cannot be made fast enough, memory has been snapped up and risen tenfold in price, and everyone is fighting for servers. If Seedance 2.0 video generation can be shortened from a four-hour queue to one minute, how many times would usage increase? As long as the bottleneck is opened and the experience improves, demand will explode tenfold or a hundredfold.

For commercialization, Xiwang is targeting the most demanding major internet companies.

Major companies have extremely stringent product requirements, but I tell our team that we must go after the hardest customers to serve and the ones with the highest standards. Only products refined under the greatest pressure can truly establish a solid foundation.

With S3’s massive delivery capacity and the team’s ecosystem strategy, this toughest nut to crack is exactly Xiwang’s next main target.

At the end of the interview, as a witness to and participant in China’s technology development, Wang Zhan said:

In this era, AI is essentially distributing intelligence. It gives humanity an opportunity to close the information gap. As long as you are clear about what you want to do, AI can give you unprecedented support. What we at Xiwang want to do is bring the cost of this extremely powerful thing all the way down.

Know yourself first, then know AI, and only then can you win every battle.

This is not only Wang Zhan’s advice to young people feeling lost in this fast-moving AI era. It may also be a true portrait of Xiwang, a young unicorn that has found a precise path through the crowded computing power market and kept sprinting ahead.