Fireworks Turns Kimi K3 Into Ember-1, Cuts Tokens 40%

Inference provider Fireworks AI has released Ember-1, built on Moonshot AI's open-weight model Kimi K3. Fireworks made no changes to the model's architecture — only post-training — with a single goal: get K3 to think less while answering just as well. The company says it cuts generated tokens by roughly 40% with essentially no quality loss.

Ember-1 is currently available as a research preview on Fireworks Serverless, API-only. No weights or training code have been released, so it cannot be self-hosted. Pricing matches K3 on Fireworks: $3 per million input tokens, $0.30 for cached input, and $15 per million output tokens.

What's actually being saved is "thinking"

Fireworks explained its motivation in a blog post: users like what K3 can do for coding, but its reasoning chains are so long that costs stay high in automated coding pipelines. By the company's own numbers, reasoning models like K3 sometimes spend more than 90% of their output on internal reasoning before ever producing an answer — and when an agent calls tools step by step, the model re-reasons through everything it already thought at every single step.

The most direct fix would be lowering K3's reasoning-effort tier, which Fireworks says it tried — quality dropped too much at the lower setting. Ember-1 takes a different route. In the company's own words:

"Ember-1 is Kimi K3 post-trained to reason in fewer tokens, not run at lower effort."

Training covered math, coding, instruction-following, dialogue, search, tool calling and software engineering, spanning more than 50 training runs and over 200 evaluations, all on Fireworks' own Serverless Training infrastructure using proprietary data. Fireworks says the underlying method is new and unpublished, so it can't currently be reproduced by outsiders.

Benchmarks: up in some places, down in others

Fireworks compared Ember-1 against three tiers of K3:

BenchmarkK3 LowK3 HighK3 MaxEmber-1
Terminal Bench 2.176.4%77.6%80.9%82.0%
SWE-bench Verified80.4%86.0%93.2%92.2%
DeepSWE 1.155.8%62.8%66.4%75.2%

On DeepSWE, Ember-1 leads K3 Max by nearly 9 points; on SWE-bench Verified, it trails by about 1 point. Third-party outlets citing separate figures also put Ember-1 slightly below K3 Max on SWE-Interact. All of this is Fireworks' own testing — there's been no independent evaluation yet.

The more persuasive numbers come from production traffic. Fireworks ran A/B tests with two customers on real coding workloads and found roughly 35% fewer tokens per task overall. The most complete data set showed reasoning tokens down 71.3%, total tokens down 39%, with task scores of 0.753 versus 0.751. One customer has already moved Ember-1 into production; the company hasn't been named.

A rough cost estimate

Since the per-token price is identical, any savings come entirely from generating less. Based on the A/B figures above, K3 Max's output cost works out to roughly $0.74 per task, versus about $0.45 for Ember-1. For a team running 10,000 coding tasks a day, that's a rough drop from about $7,400 to $4,500 in daily output costs — over $80,000 a month. This estimate only covers output tokens, not input or caching, so actual savings will depend on task length.

For developers in China, Ember-1's significance is mostly methodological. K3 is open-weight, so anyone can post-train it, and Fireworks has shown that "compressing reasoning length" can be sold as a standalone product on its own. Domestic inference providers hold similarly open models — K3, DeepSeek, Qwen — so it's worth watching whether comparable "token-saving" versions appear. Using Ember-1 itself from within China means routing through an overseas API, so latency and compliance are considerations users will need to weigh themselves.

Sources: Fireworks AI official blog, MarkTechPost, CocoLoop, Fireworks model page; verified against the three benchmark scores, A/B test token and score data, and per-million-token input/output pricing.