The open-source inference framework vLLM, working with Inferact, published a technical blog post on October 7 detailing a suite of inference optimizations built for DeepSeek-V4.1-Flash. In tests simulating agentic workloads, the team measured a 1.9x speedup at low concurrency, and a 5.3x throughput gain when per-user output speed was capped at 150 tokens per second.
The workload was SemiAnalysis's AgentX benchmark, which the team treats as representative of agentic serving scenarios — multi-turn conversations, long context windows, and session affinity that keeps prefix-cache hit rates high. The low-latency setup used 4-way tensor parallelism with FlashInfer; the high-throughput setup used 2-way data parallelism combined with expert parallelism.
What's different about the model itself
According to the blog post, V4.1 Flash uses a causal encoder-decoder architecture with 40 layers total, where layer 20 computes the global KV used by the decoder. Each token activates roughly 16 billion parameters during decoding and about 8 billion during prefill. The global KV cache is compressed and shared across layers, coming to roughly 890 bytes per token at FP4 precision; a separate sliding-window KV cache covers the most recent 128 positions, stored uncompressed at FP8.
That small a cache has a direct consequence: across the full benchmark run, there was no need to offload the KV cache to memory or disk. For long-context agentic tasks, offloading is often one of the main sources of latency.
The standout items among eight optimizations
The team listed eight changes; a few delivered the biggest gains:
- Bounded sliding-window replay: when a prefix-cache hit occurs, the system recomputes just the most recent 128 tokens, paired with CUDA Graph. This cuts time-to-first-token by about 30%, and by nearly 70% on contexts of roughly 100,000 tokens. It's enabled by default on V4.1 and can be disabled with
--no-swa-bounded-replay. - Sparse MQA scoring kernel: 14 to 23x faster than the original implementation at a 512K context, though only 1.2x faster at 8K — translating to roughly a 3% to 6% end-to-end decoding gain.
- NVFP4 attention: shrinks the KV cache by 45% compared with the earlier FP8 approach, and speeds up the corresponding step by about 1.45x.
- Engram lookup: asynchronous prefetching with transparent huge pages speeds up lookups by as much as roughly 10x.
A few more changes came from operator fusion: folding several steps of mixture-of-experts routing into a single kernel gave 1.18x to 1.31x gains at medium batch sizes, while mHC-related kernels ran 1.14x to 1.51x faster than the TileLang version on GB200.
On accuracy, the team ran comparisons on GSM8K and GPQA and said the quality loss was "negligible," within about 1.5 standard errors. The blog post did not share evaluation results on additional tasks.
How much of this applies to deployments in China
All these figures come from Nvidia's Blackwell platform. Due to U.S. export controls, institutions in China have limited access to chips like the GB200 and GB300, and NVFP4 is a Blackwell-exclusive data format — so the related gains can't be directly reproduced on H20 chips or domestic accelerators.
The parts that do transfer sit mainly in the scheduling and caching layers. Bounded sliding-window replay, session affinity, and skipping KV cache offload are only loosely tied to specific hardware, and the code has already landed in vLLM's main branch — teams running the same model can pick it up simply by upgrading. Progress on the various kernel integrations is tracked centrally in vLLM's GitHub issue #57448. No third party has yet measured the actual speedup on domestic chips.
Sources: vLLM's official technical blog, CocoLoop, SemiAnalysis's AgentX benchmark documentation; the blog's figures were checked against the stated speedup multiples, test hardware, and parallelism configurations.