DeepInfra said on May 4 that it had raised $107 million in Series B funding. The amount is not the largest round in AI infrastructure, but the investor list says a lot.
500 Global led the round, with Nvidia Corp., Samsung Next, Supermicro, A.Capital, Felicis, Peak6, Upper90 and early Google engineer Georges Harik also participating.
Nvidia investing in an independent "inference cloud" startup would have been unusual a couple of years ago, when most inference demand was still absorbed inside the large cloud platforms. By 2026, Nvidia is putting capital into an independent infrastructure layer, a sign that inference has become a market of its own.
Where DeepInfra stands now
DeepInfra CEO Nikola Borisov put the shift plainly in the announcement:
"Inference is no longer a thin layer - it's the system constraint that will define the majority of workloads."
In other words, inference is no longer a thin API wrapper. It is becoming the system bottleneck for much of the next generation of compute workloads.
- More than 190 open AI models are supported, including Llama, Qwen, DeepSeek and Mistral.
- More than 30% of token throughput comes from agent workloads, meaning autonomous agents rather than ordinary chat.
- Eight U.S. data centers run company-owned GPU infrastructure on Nvidia's Dynamo distributed inference platform.
- The GPU mix includes Blackwell and Vera Rubin. Vera Rubin only entered volume production this year, which makes the Nvidia relationship hard to ignore.
- The company also says it processes 5 trillion tokens a week, has grown token volume 25 times since its Series A, and has tripled revenue since the start of 2026.
The 30% figure is the one to watch. Agent workflows behave very differently from chatbots: a single task may call a hundred tools and run dozens of API requests. They are always-on systems, not sessions that begin only when a user types a prompt. Traditional clouds were not designed around that profile; their billing models, cold-start assumptions and scheduling logic were built for human interaction loops.
Why owning GPUs suddenly matters
The core story behind the round is simple: DeepInfra does not rent GPUs from AWS; it buys and operates its own capacity.
For several years, the default way to run AI inference was to rent GPU services from AWS, GCP or Azure. The problem is that those same cloud companies also sell model APIs, keep prices high and often carry cold-start latency. Together, Fireworks, Groq and DeepInfra rose together over the past two years because they attacked that mismatch.
DeepInfra claims a 20-fold cost-efficiency improvement. That number needs context, because benchmarks depend heavily on the comparison. But for teams running open models at high volume, such as products with million-scale monthly active users, moving to a dedicated inference cloud can commonly cut the bill by more than half.
How the inference cloud field stacks up
| Company | Positioning |
|---|---|
| Together AI | Full-stack platform with its own CUDA alternative |
| Fireworks AI | Closely tied to the Hugging Face ecosystem |
| Groq | LPU specialist that has drawn attention from Nvidia and AMD |
| Cerebras | Wafer-scale chips and a path toward public markets |
| DeepInfra | Eight self-operated data centers, 190-plus open models, and a focus on agent workloads |
DeepInfra does not have the flashiest valuation in the group, but its positioning is unusually clear: inference infrastructure for open models. It is not trying to broker closed models or compete at the model layer. The closest comparison is Together, with one important difference: DeepInfra keeps tighter control of the stack, from GPUs to scheduling to the API surface.
Agents are the real bet
Borisov's line that inference is the system constraint is not just a marketing phrase.
As agent workflows become mainstream, each user may trigger dozens of inference calls per minute in the background. Cursor writing code, Claude Code editing files, or Salesforce Agentforce handling service tickets all have a different load profile from a ChatGPT-style exchange where a user asks once and the model answers once.
The vendor that can keep p99 latency low, control costs and offer zero data retention in that setting will be the one enterprise customers sign with.
DeepInfra's latest fundraising language mentions "agentic workloads" more often than in earlier rounds. Over the next year, this category will be less about who has the broadest model catalog and more about whose infrastructure can survive a world where thousands of agents hit inference at the same time.
Nvidia's decision to invest suggests it thinks that moment may actually be arriving.
Sources: DeepInfra Closes $107M Series B to Power Production-Scale AI Inference (GlobeNewswire); Deepinfra lands $107M in funding to build out its dedicated inference cloud for open-source models (SiliconANGLE); CocoLoop; DeepInfra Raises $107M Series B to Scale Inference Infrastructure (DeepInfra Blog)