T-Head's Zhenwu V900 AI Chip Debuts, Mass Production Delayed to 2027

Alibaba's chip unit T-Head unveiled its next-generation training-and-inference AI chip, the Zhenwu V900, at the 2026 Hangzhou Apsara Conference on September 22. The company says it delivers three times the compute of its predecessor, the Zhenwu M890, and that a single cluster can scale up to 500,000 cards. According to reporting from multiple outlets present at the event, the chip will enter mass production in 2027, with the next-generation Panjiu supernode servers built around it launching in the first quarter of 2027.

In other words, this launch marks the next stop on a roadmap — the chip Alibaba Cloud customers can actually rent this year is still the M890.

What specs were disclosed

T-Head disclosed three specifications: 216GB of memory, 1,200GB/s of inter-chip interconnect bandwidth, and support for both FP8 and FP4 precision. The official positioning is that it can meet the training and inference needs of trillion-parameter-class large models.

The Panjiu supernode servers will bundle the Zhenwu V900 with T-Head's own ICN Switch, Panmai, and Zhenyue chips, with switching, networking, and storage control all handled by in-house silicon. T-Head also previewed a further generation, the Zhenwu J900, saying it will adopt a proprietary parallel computing architecture and is planned for 2028; different reports give different quarters for that launch.

What wasn't disclosed makes up a longer list: peak per-card compute, power consumption, process node and foundry partner, and memory type were all absent from the launch materials. The "3x" figure is a multiple relative to the M890, and T-Head has never fully disclosed the M890's absolute compute figures either.

Where the M890 stands now

Alibaba CEO Eddie Wu used his keynote to give a progress update on the chip currently in service: the in-house M890 AI supernode can already support inference for models with 2 trillion parameters, and is beginning to roll out at scale in Alibaba Cloud data centers this quarter. The corresponding cloud product, the Lingjun Zhenwu M890 supernode instance GP9A, is already for sale.

Wu also said that because T-Head's chip product line has matured and customer adoption is growing, annual shipment volumes are expected to rise significantly. Alibaba did not give a specific shipment figure.

The bigger numbers in the keynote landed on models and data centers: Qwen plans to train new models with 5 trillion to 10 trillion parameters, and Alibaba Cloud's global data center capacity is set to exceed 20GW by 2032. Wu called AI models, AI chips, and AI cloud the "three pillars of the machine intelligence era," saying Alibaba would keep investing firmly in all three.

What it means for domestic users

For domestic companies running models on Alibaba Cloud, the most direct takeaway from this launch is the timeline. Before the first quarter of next year, the only domestically developed compute option available is the M890 instance; Alibaba didn't say when the V900 supernode would be offered as a cloud instance, or at what price.

For customers buying hardware for their own data centers, 216GB of memory is a number that can be compared directly — it determines how large a model shard a single card can hold, and how many cards are needed to fit a trillion-parameter model. FP4 support means inference workloads can trade precision for throughput, provided the software stack keeps up. T-Head's chips have so far mainly run Alibaba's own workloads and Alibaba Cloud; there's still no third-party data on how much adaptation cost outside customers would face migrating over.

Mass production in 2027 is still more than a year away. In that time, competitors in China's domestic AI chip market will also be moving to new generations, so where the V900's claimed 3x advantage lands relative to contemporaneous products will have to wait for real per-card compute figures and benchmark results.

Sources: keynote by Eddie Wu at the 2026 Hangzhou Apsara Conference, IT Home, CocoLoop, Sina Finance; chip specifications verified against T-Head's on-stage disclosures, with mass-production and launch timing based on multiple outlets' on-site reporting.