AI agents now have dedicated computing power.

Last Thursday, at the Open Computing Technology Summit (OCTS26) in Beijing, Inspur Information unveiled a series of AI infrastructure products, including what it calls the industry’s first CPU-native liquid-cooled full-rack server and the YuanBrain SD200 Supernode AI Server.

As the industry consensus shifts from “building large models” to “using agents,” Inspur Information has rebuilt the foundation of its computing power architecture for AI workloads whose application patterns are changing.

The Agent Era

Computing Power That Needs Purpose-Built Optimization

This year is a key inflection point for the scaled deployment of agents. IDC forecasts that the global agent market will grow at a compound annual rate of 139% from 2025 to 2030. Gartner expects 40% of enterprise applications to integrate agents this year, and more than 30% of enterprise applications to deeply embed agent capabilities by 2028.

At the software layer, the commercialization trend has reached every level. On one hand, advanced large models such as Kimi, DeepSeek and GLM are moving toward agent-native architectures, with stronger built-in capabilities for task planning, tool use and autonomous execution. On the other, enterprise agent frameworks such as ChatGPT Work and Workbuddy are spreading across office, R&D and operations scenarios. Enterprise AI applications are shifting from “one-off model calls for Q&A” to a “swarm intelligence work model” in which hundreds or thousands of agents collaborate continuously in the background and autonomously complete complex tasks.

Large models of the past were like brains. Applications driven by native agent large models are more like robots fitted with hands and feet, ready to carry out tasks.

At the same time, that evolution places extreme demands on AI computing power. From a single conversation to end-to-end project delivery, token consumption grows exponentially. One user request may trigger hundreds of downstream subtasks and tool calls, which in turn means calls across hundreds of chip cores inside the servers behind it.

Faced with complex workloads, the division of labor in computing power hardware is also being redefined. The traditional split, in which the CPU handles scheduling and the GPU handles computation, is being broken.

At the underlying logic level, each agent is essentially a small CPU sandbox environment, mainly responsible for logic management, process and resource scheduling, and system coordination. That way of working is not the parallel matrix computation at which GPUs excel; it is naturally suited to CPUs. Research shows that in an agent execution pipeline, CPU-related processing can account for as much as 90.6% of end-to-end latency.

That means the CPU will become significantly more important. GPUs determine the upper bound of model capability, while CPU-driven multi-agent collaboration can improve the completeness and reliability of AI outputs through engineering.

In terms of computing power ratios, the agent era will require data centers to add large numbers of standalone, CPU-only computing power clusters. In traditional AI servers, the computing power ratio between CPUs and GPUs is roughly 1:8 to 1:4. In the agent era, data centers will need not only massive GPU capacity for large-model inference, but also CPU servers to carry the load of agent hosts.

Industry-side information shows that nearly all incremental CPU server procurement this year by China’s leading internet companies has gone to agent-related businesses. The corresponding agent infrastructure, or Agent Infra, has also become a direction the industry is exploring collectively.

This paradigm shift is redefining AI infrastructure.

CPU-Native Liquid-Cooled Full-Rack Server

Building a High-Density Carrier for Swarm Intelligence Collaboration

At the 2026 Open Computing Technology Summit held on July 9, Inspur Information released two core offerings for the agent era: the industry’s first CPU-native liquid-cooled full-rack server and the YuanBrain SD200 Supernode AI Server. Together, they offer an open-architecture solution for scaled agent deployment from both ends: the CPU computing power base and the GPU inference engine.

The clearest feature of the agent era is that “swarm intelligence collaboration” becomes the norm. Task completion is no longer a single response from a single model, but a process in which large numbers of agents divide up task planning, tool calling, data retrieval, process execution and result aggregation. Behind that sits CPU computing power on an unprecedented scale.

In public cloud scenarios, leading agent applications are already showing striking CPU resource consumption. Each agent instance typically requires two CPU cores to support sandbox operation, task decomposition and external interaction. Agent products with hundreds of millions of users correspond to enormous pools of continuously running CPU computing power.

On the enterprise side, scaled agent deployment brings a different control problem: agents scattered across endpoints create risks such as confused permissions, missing security audits and inconsistent versions. Enterprises urgently need a unified, controllable and scalable runtime foundation for agents.

Rising power density in data centers is the second driver of architectural reconstruction. The power limit of a traditional air-cooled rack is about 40 to 50 kilowatts. By the end of 2026, the power of a single high-density AI computing power rack will exceed the 300-kilowatt level, a 10- to 50-fold increase. That has already moved beyond the limits of air cooling and even hybrid air-liquid cooling. This is why native liquid cooling has become the inevitable choice for high-density computing power.

The industry’s first CPU-native liquid-cooled full-rack server launched by Inspur Information redefines CPU computing systems with a new native liquid cooling architecture.

Under the native liquid cooling architecture, the system is built on the open OCM architecture. A single rack can support up to 384 heterogeneous CPU processors and more than 40,000 agents running in collaboration at the same time, sharply improving the density and management efficiency of enterprise agent deployment.

The rack’s power consumption can reach the megawatt level, several times that of a traditional general-purpose CPU rack.

Unlike traditional hybrid air-liquid cooling designs, this CPU full-rack server uses a native liquid cooling architecture. Through the co-design of computing and cooling, it decouples and planarizes all heat-generating components, including memory, optical modules and network cards. An integrated cold plate enables an extreme cooling form with zero hoses, zero cables and zero fans, addressing the cooling bottleneck of high-density CPU systems at the hardware foundation while improving operating reliability and energy efficiency.

The rack’s core technical breakthroughs are concentrated in three areas:

CPU computing system reconstruction. Unlike the traditional approach of adding liquid cooling to an air-cooled architecture, this design co-optimizes cooling and computing architecture to create 0.5U ultra-thin computing power nodes, enabling a high-density architecture with 16 CPUs deployed within 2U of space.

Standardized computing power modules. Based on the liquid-cooled OCM open computing module architecture, the design supports seamless compatibility with heterogeneous CPU architectures such as x86 and ARM. It can flexibly derive different forms of computing power nodes according to business needs, while balancing performance stability under high load with the large-memory and high-bandwidth requirements of long-context scenarios.

Full-domain component liquid cooling reconstruction. Breaking through the limitation of traditional liquid cooling that covers only the CPU, the system brings all heat-generating components, including memory, network cards and optical modules, into the liquid cooling system. Its cable-free design supports hot maintenance, ensures zero business interruption and improves full-rack operations and maintenance efficiency by more than 100%.

For the evolution toward future gigawatt-scale AI data centers, high-density liquid-cooled racks have also been adapted to the trend of 800V high-voltage power delivery into the rack. Traditional 380V power supply, in single-rack scenarios at the hundreds-of-kilowatts level, faces problems such as overly thick copper cables and sharply increased construction and maintenance difficulty. 800V high-voltage direct current has become standard for megawatt-class racks. This architectural design reserves room to adapt to the next generation of power supply standards.

YuanBrain SD200 Supernode Upgrade

Token Generation Latency for Trillion-Parameter Models Falls to 4.77ms

If the native liquid-cooled server solves whether agents can run at scale, the upgraded YuanBrain SD200 Supernode AI Server provides agents with a high-quality, low-latency “intelligent output engine.”

This time, Inspur Information introduced the upgraded YuanBrain SD200 Supernode AI Server and said it has been the first to complete high-performance optimization for leading mainstream open-source large models including Kimi K2.6, DeepSeek V4, GLM 5.2 and MiniMax M3. This is an important iteration of the YuanBrain SD200 series: when the product was first released in 2025, YuanBrain SD200 achieved single-token generation speed of 8.9ms, making it China’s first supernode system to break the 10ms threshold.

After a year of architectural optimization, YuanBrain SD200 has made another performance breakthrough. Test data show that on the Kimi K2.6 trillion-parameter model, the server needs only 4.77ms to generate a single token, while first-token latency is down significantly by 35% from before optimization. That is enough to support the low-latency requirements of high-frequency agent calls, multi-turn interaction and parallel multi-agent collaboration.

Inspur Information said the performance gains come from coordinated optimization across the full software and hardware stack. At the hardware level, it has continued to optimize the Fabric interconnect architecture inside the supernode, improving the determinism of inter-card communication routing and reducing communication time. At the software level, it has introduced inference optimization technologies such as multi-token prediction and JIT.

Enterprise agent use cases are quickly becoming more complex, and a single model can no longer cover the full range of needs, including long-text understanding, code generation, logical reasoning, multimodal processing and industry knowledge Q&A. Different models have strengths across different capability dimensions. “Multi-model fusion” is becoming a core path for improving agent intelligence: multiple advanced AI models generate results in parallel, which are then reviewed, compared and fused into a final output. This can effectively break through the capability boundaries and perspective limits of any single model, providing agents with a more stable and trustworthy supply of intelligence.

The effectiveness of this route has already been validated. In the DRACO deep research benchmark, the fusion model achieved a top score of 53.9%, demonstrating its advantages in complex research and multi-step analysis scenarios. In two international benchmarks, AIME 2026 for mathematical reasoning and GPQA for general high-difficulty Q&A, the fusion model led single models with scores of 97.2% and 90.8%, respectively, while improving complex reasoning and professional knowledge Q&A capabilities.

In system capability, a single YuanBrain SD200 machine can carry a large model with up to 4 trillion parameters and also supports parallel deployment of multiple trillion-parameter models, matching the application pattern of multi-model fusion. Combined with the multi-model fusion capability of the YuanBrain Enterprise Intelligence EPAI platform, users need to make only one API call. The system can synchronously distribute the task to multiple large models, let different models generate candidate results, and then complete cross-analysis and comprehensive judgment through a review-and-fusion model to output a more complete and reliable final answer.

Through a system-level combination of open interconnect technology, software-hardware co-optimization and multi-model fusion capabilities, this model breaks through the performance limits of standalone hardware and provides end-to-end inference engine support for complex agent applications.

To meet demand for on-premises enterprise agent deployment, Inspur Information also launched the YuanBrain SD200 Supernode Enterprise Edition. It continues the native memory semantic open interconnect architecture of the standard version and builds a 16-card unified scale-up compute domain based on an open switching architecture, enabling unified addressing and low-latency cross-card communication. It can reduce first-token latency for trillion-parameter models by more than 40%. A single Enterprise Edition machine supports terabyte-class unified GPU memory and can fully host today’s mainstream open-source trillion-parameter models, meeting enterprises’ core business needs for long-context understanding, complex logical reasoning and multi-agent collaboration, while substantially lowering the barrier to deploying high-performance agent computing power.

In the past, enterprises typically deployed only hundred-billion-parameter models, which supported relatively shallow applications such as AI-assisted coding and writing. Trillion-parameter models can directly generate usable complete code and produce complex plans, allowing AI to truly enter production workflows and replace human labor. Now, more enterprises can support deep, production-grade agent applications without buying ultra-large-scale supernodes.

Open Collaboration Accelerates Agent Deployment

Computing Power Upgrades Drive Intelligent Transformation

From large models to agents, innovation in AI infrastructure is moving from simple hardware upgrades to system-level collaborative reconstruction.

The two products released by Inspur Information form a clear division of computing power labor: GPU supernodes are responsible for thinking, continuously generating low-latency, high-quality tokens; CPU-native liquid-cooled full-rack servers are responsible for action, carrying massive agent scheduling and orchestration, tool calls and orderly collaboration. Together with the YuanBrain Enterprise Intelligence EPAI platform for unified control, the three form a complete technical system of “swarm intelligence collaboration plus multi-model fusion,” echoing the core direction of the agent era.

Judging by the pace of industry deployment, the second half of 2026 will be the point at which supernode solutions move into scaled rollout, with leading domestic internet customers already entering batch deployment. In the open computing ecosystem, the continued push from standardized hardware architectures and collaboratively optimized software-hardware systems will keep reducing both per-token generation costs and agent operating costs. That will help AI move toward scaled deployment across entire enterprises and full workflows, becoming a core driver of business process reconstruction and productivity improvement.

The rise of AI agents marks the start of a transformation in computing power infrastructure. Over the long term, token production will become more finely segmented, much like an industrial assembly line: the Prefill and Decode stages will separate, attention computation and feed-forward networks in the Decode stage will also be split, and each stage will be matched with the most suitable chip architecture and system design to optimize efficiency across the full chain.

This direction of evolution also aligns with the underlying logic of Inspur Information’s converged architecture: pooling computing power, memory and storage through high-speed interconnects, so that all resources can be freely connected, combined on demand and maximized in value. Over the next two to three years, gigawatt-scale AI data centers are expected to gradually come online, and new forms of AI computing infrastructure will support AI’s move from technical innovation to large-scale application.