Zhipu GLM-4.5 Runs Full Model on Just 8 H20 GPUs

Zhipu AI's GLM-4.5 has taken a highly symbolic step—it has been specifically optimized for NVIDIA's H20 (the China-specific version of the GPU), allowing the full model to run on just eight cards.

Why This Matters

After U.S. chip export restrictions on China, the best NVIDIA GPU available to Chinese AI companies is the H20—a version with reduced interconnect bandwidth. Many cutting-edge overseas models either cannot run on the H20 or do so with very low efficiency.

Zhipu chose to directly confront this constraint, adapting the model architecture and inference pipeline specifically to the H20's hardware characteristics. The cost and entry barrier of eight H20 cards is far more accessible than solutions requiring dozens or hundreds of A100 GPUs.

Technical Adaptations

Specific optimizations include:

  • Adjusting the attention computation method to match the H20's memory bandwidth characteristics
  • Designing the model sharding strategy specifically for an 8-card configuration
  • Adapting the quantization scheme during inference to the H20's computational precision

Performance

GLM-4.5's performance on Chinese language understanding and generation tasks places it in the top tier of domestic models. It matches or approaches GPT-4-level models on most Chinese benchmarks. There is a gap in English tasks, but it is narrowing.

More importantly, the actual deployment cost—for domestic enterprises, being able to achieve near-frontier results using legally purchasable hardware—holds far greater practical value than benchmark scores.

Industry Impact

Zhipu's approach represents a strategy for the Chinese AI industry to cope with chip restrictions: instead of waiting for the best hardware, make the best possible use of existing hardware.

This aligns with DeepSeek's strategy of using limited compute to drive costs extremely low. When external conditions are constrained, the engineering optimization capabilities developed under pressure may become a long-term competitive advantage.

Sources: CocoLoop, Zhipu AI official release