Inference platform Baseten published an engineering blog post on October 2: using Claude Code to drive Fable 5, they had an AI agent write a single-model inference engine from scratch, and it beat vLLM across every traffic pattern they tested.
The engine is called VibeQwen, built specifically to serve Alibaba's Qwen-3.6-35B-A3B. In the post, author Shawn Rushefsky wrote:
"the generated engine, called VibeQwen, outperformed vLLM across all tested traffic patterns, and by very large margins in some cases"
How the Numbers Were Measured
The test hardware was Nvidia's B200, the baseline was vLLM 0.25.1, and the load-testing tool was AIPerf. The results:
- Single-stream decoding reached 1,792 TPS versus vLLM's 943 TPS, about 90% faster;
- Time to first token dropped from 28ms to 12ms, roughly 2.3x faster;
- At 32 concurrent streams, throughput was 71% higher.
On cost, building VibeQwen took about a week end to end, consumed roughly 1.7 billion tokens (most of them cache hits) and about 200 B200-hours, for a total bill in the low thousands of dollars.
Baseten tried the same approach once more on image segmentation. The second result, called Sammie, serves Meta's SAM 3.1 and runs on H100s. It took a few days and cost a few hundred dollars and about 200 million tokens. At 32 concurrent streams, its throughput was 50% higher than Meta's own reference server, hitting 91 images per second.
What the Agent Actually Did
Baseten extended an in-house framework called MetaInfer into a full serving stack and let the agent fill in the rest. The agent could consult existing open-source implementations for reference, and after every change it ran an AIPerf load test and used the numbers to decide what to try next.
One constraint was written very specifically: any change that might affect output precision required the agent to stop and get human approval. Small numerical drift is one of the hardest things to catch in an inference engine, and if a speedup comes at the cost of precision, a benchmark leaderboard won't show it. That's exactly the failure mode this gate was built to block.
The blog post draws its own boundaries, too. Both VibeQwen and Sammie are still experimental — neither has handled production traffic. The Sammie comparison is rougher, and the author admits those numbers only amount to "suggestive evidence": the two setups use different architectures and different accelerators, and no controlled test was run.
What This Means for Chinese AI Teams
VibeQwen targets an open-source Qwen model, and a large share of production services in China happen to run on Qwen paired with vLLM or SGLang, so these numbers speak directly to what many teams deal with day to day.
It's worth comparing the economics of the two approaches. A general-purpose engine has to serve hundreds of model architectures, which rules out many aggressive optimizations tailored to just one; a dedicated engine focuses on a single model and can shape attention kernels, MoE routing and scheduling entirely around the 35B-A3B architecture. Building a dedicated engine used to take a team months; Baseten compressed that into a week and a few thousand dollars — by rough math, cheaper than a month of one inference engineer's salary.
- The engine is locked to the model. Once Qwen is upgraded with a different architecture, the engine will most likely need to be regenerated from scratch.
- Testing only covered the B200. The accelerators available in China vary widely, and the blog post has no data on whether the same process reproduces on other hardware.
- Load testing doesn't necessarily cover what production needs most: long-tail stability, memory fragmentation, and handling malformed requests.
Even so, the idea of "a custom, auto-generated engine for every flagship model" now has one verifiable example. Whether vLLM's role shifts from "production workhorse" to "baseline and fallback" will depend on whether any team actually puts an engine like this into production and publishes months of real-world data.
Sources: Baseten engineering blog, CocoLoop; TPS, time to first token, concurrent throughput, token usage and GPU-hours verified against the Baseten blog post.