SGLang speeds up MiniMax-H3 1.95x with zero quality loss
LMSYS benchmarked MiniMax-H3 on 8x H200 GPUs: SGLang's lossless path beats Diffusers by up to 1.95x, while steeper tiers trade SSIM fidelity for up to 6x speed.
5 verified stories covering inference optimization, product updates and industry developments.
LMSYS benchmarked MiniMax-H3 on 8x H200 GPUs: SGLang's lossless path beats Diffusers by up to 1.95x, while steeper tiers trade SSIM fidelity for up to 6x speed.
SGLang and Ant Ling Infra cut single-request TPOT for Ling-3.0-flash from 3.33ms to 0.78ms on 4 Blackwell GPUs by removing host-side stalls.
Ant Group and SGLang keep quantized weights resident in GPU memory, cutting Ling-2.6-1T restart time from 8.8 minutes to about half a minute.
LMSYS's DeepSeek-V4-Pro benchmarks show Nvidia's China-market H20 chip closing the decode-speed gap with the flagship B300 to just 1.42x.
NVIDIA's Nemotron-Labs-Diffusion introduces a tri-mode decoding approach that boosts throughput up to 6× over Qwen3-8B while maintaining accuracy.