The chip behind Luckin Coffee can no longer stay hidden.
When you walk into a Luckin Coffee store and order a Coconut Latte, every step of the process, from placing the order to preparing the drink, redeeming the coupon and picking it up, is being watched by a pair of “eyes” behind the scenes.
What is it watching?
It has to identify orders in real time, judge production rhythm, verify material status, monitor equipment operations, and sync data back to headquarters for quality control, scheduling and operational decisions.
This is the layer of edge-side AI hidden behind Luckin Coffee’s stores. What they share is a hard requirement: computing power must be deployed close to the action, response times must be fast, stability must be strong, and costs must stay under control.
The chip has become the critical piece of that system.
Today, the chip behind Luckin has finally come into view.
It comes from Iluvatar CoreX, a Chinese general-purpose GPU company that only recently completed its listing.
Just half a month after ringing the bell, Iluvatar CoreX has rolled out four edge computing power products in one move: the Tongyang series.
And this is not just a product launch. The edge computing power Luckin is already using is Tongyang.
So how capable is this product line? Let’s look closer.
Four New Products Launched at Once
Start with the name.
The name Tongyang comes from inscriptions on Shang and Zhou dynasty bronzes. Inside Iluvatar CoreX, however, it carries a more specific meaning:
“Tong” points to high-efficiency computing capability, while “yang” signals its role as the core computing power hub in edge-side scenarios.
In other words, this is a computing power series built for real-world business sites.
The first Tongyang lineup includes four products: TY1000, TY1100, and the computing terminals TY1100_NX and TY1200.
First, the Tongyang TY1000.
It is a module with a standard 699-pin interface, small enough to fit in a pocket, yet Iluvatar CoreX has packed nearly 200T of dense computing power into that compact footprint.
In real-world testing, whether on typical computer vision tasks, NLP inference, inference on the 32B-parameter DeepSeek-R1 model, or scenarios involving embodied intelligence VLA models and world models, the TY1000 delivered performance across multiple metrics that was not weaker than mainstream international solutions.
According to test data disclosed by Iluvatar CoreX, the TY1000’s overall efficiency across multiple workloads exceeded the typical configuration corresponding to NVIDIA’s AGX Orin.
That does not mean it can replace everything across the board. But it does prove one thing: in general-purpose edge inference, Chinese general-purpose GPUs are now capable of direct comparison.
Next is the Tongyang TY1100.
This product has received a further architectural upgrade, using a 12-core ARM v9 CPU and offering more ample system-level computing power.
It targets complex scenarios with higher requirements for both general computing and AI inference, such as multi-sensor fusion, edge data preprocessing and real-time decision-making.
If the TY1000 is closer to a computing power core, the TY1100 looks more like a full edge computing foundation.
Then comes the TY1100_NX, aimed at users more sensitive to memory capacity and price-performance.
Its larger memory configuration gives it greater stability in scenarios such as multi-model parallelism and long-sequence inference, while retaining a plug-and-play deployment model that lowers the barrier to system integration.
Finally, the Tongyang TY1200 is defined by Iluvatar CoreX as a computing power terminal.
Its computing power specification rises to 300 TOPS. More importantly, it is an integrated solution designed for terminal form factors. Its target users are not only algorithm engineers, but also industry customers that want to put AI capabilities directly into devices.
Looking at the product portfolio, the Tongyang series does not chase a single extreme point. Instead, it deliberately spans different levels of computing power, form factors and price bands, covering deployment needs from computing power modules to terminals.
But Iluvatar CoreX has not focused only on chip specifications.
At the ecosystem level, the Tongyang series achieves pin-to-pin compatibility with mainstream products in both interfaces and form factors, sharply reducing the cost for customers to migrate from existing solutions.
For industrial and commercial customers with mature systems already in place, that is almost the dividing line between whether they are willing to try it or not.
More importantly, these products were not launched for the sake of a launch.
In robotics, Tongyang has entered real enterprise application scenarios through a partnership with Gelanruo Robotics. On the industrial side, manufacturers including Biyi Electric are using it to upgrade equipment intelligence. In commercial retail, Luckin Coffee is just one typical case. In transportation, the Tongyang series has also taken part in multiple vehicle-road-cloud integration pilots.
When four completely different industry scenarios begin using the same general-purpose GPU computing power foundation, a bigger question emerges:
What does Iluvatar CoreX really want to do?
Iluvatar CoreX’s Ambition Is Also Showing
Looking only at the Tongyang series, it is easy to read this as a Chinese chip company trying to complete its cloud-edge-device business map ahead of others.
But judging from the architecture roadmap it also disclosed publicly on January 26, the story is clearly not that simple. In cloud scenarios, its core business base, Iluvatar CoreX has even more ambitious goals.
Iluvatar CoreX is not satisfied with the interim goal of domestic substitution. On multiple public occasions, it has made clear that its long-term aim is to benchmark against, and even surpass, industry leaders such as NVIDIA.
To that end, Iluvatar CoreX has laid out an architecture roadmap with specific years attached.
In 2025, Iluvatar CoreX launched its Tianshu architecture, surpassing NVIDIA Hopper. This is understood to be no longer a plan, but a reality: the architecture supports workloads from high-precision scientific computing to AI precision computing, and when AI chips execute attention-related calculations, the actual effective utilization of computing power reaches 90% or higher.
Test data shows that the Tianshu architecture improves efficiency by 60% compared with the current industry average, and delivers about 20% higher performance than the Hopper architecture on average in DeepSeek V3 scenarios.
By 2026, the Tianxuan architecture will add support for ixFP4 precision and benchmark against Blackwell. The Tianji architecture will cover all-scenario AI and accelerated computing, surpassing Blackwell.
In 2027, Iluvatar CoreX’s planned Tianquan architecture points directly to a comprehensive surpassing of the Rubin architecture, with a focus on integrating more precision support and innovative designs.
Behind this roadmap is a full set of underlying technical capabilities.
Technologies including TPC Broadcast, Instruction Co-Exec and Dynamic Warp Scheduling form Iluvatar CoreX’s core advantages in instruction-level parallelism, resource scheduling and computing power utilization.
These capabilities will determine whether it truly has the potential to evolve over the long term in the general-purpose GPU race.
So does Iluvatar CoreX really have that strength?
One straightforward way to judge is generality.
So far, Iluvatar CoreX’s general-purpose GPUs have run more than 400 mainstream models stably, and the company emphasizes Day 0 adaptation capability. Taking DeepSeek as an example, adaptation and inference on the Iluvatar CoreX platform have already become part of customers’ actual deployments.
The second dimension is commercial rollout.
According to its publicly disclosed data, Iluvatar CoreX has delivered more than 52,000 chips in total and served more than 300 customers.
In real-world applications, computing power costs for internet AI customer service have been cut in half while single-machine performance has doubled. In finance, research report generation efficiency has improved by about 70%. In demanding cluster scenarios, its 1,000-card-scale clusters have achieved more than 1,000 days of stable operation.
These figures are eye-catching, and specific enough to matter.
What is especially notable is that Iluvatar CoreX disclosed core metrics such as customer count, mass-production shipment scale and card-level gross margin in relatively complete detail in its prospectus. That level of openness is not common in China’s chip industry today.
This is also where Iluvatar CoreX has begun to separate itself from many big-tech in-house chip efforts and dedicated NPU paths.
It has chosen a harder and slower road: sticking with general-purpose GPUs and building the full stack in-house, from architecture, instruction sets and compilers to the software stack. That means there are no blind spots, and that every step has to be worked through on its own.
Finally, back to that cup of Luckin coffee.
As Chinese computing power truly enters thousands of industries, reaching stores, factories, roads and devices, chips are no longer just specifications on a launch-stage slide. They have become an indispensable part of the business chain.
From this perspective, the significance of this launch may lie not only in the four-generation architecture roadmap and four new edge products, but in the fact that Chinese general-purpose GPUs are both looking up, trying to surpass industry benchmarks and chart their own path into uncharted territory with greater ambition, and looking down, pushing deeper into industrial sites in a way that stays closer to reality.
That may be what Iluvatar CoreX is really trying to prove.
Comments
00No comments yet. Be the first to weigh in.