Ember-1 Cuts Kimi K3 Output Tokens 39% at Identical Price

September 28, 2026 • news
Inference EfficiencyReasoning ModelsAI Agents

Fireworks AI has shipped Ember-1, a model produced by post-training Moonshot AI's open-weight Kimi K3 to generate shorter reasoning traces without reducing declared reasoning effort. Fireworks reports roughly 40% fewer tokens per task compared with K3 at maximum effort, backed by internal benchmarks and live production A/B tests with two enterprise customers. For developers running agentic coding pipelines, that compression translates directly into lower output-token billing and reduced per-turn latency — without touching the inference-time effort dial.

The distinction from effort-level tuning matters technically. Fireworks states that dialing down K3's reasoning effort setting gave up too much quality to be a viable cost lever. Ember-1 instead changes what the model has learned to generate: it retains useful self-correction and assumption-revisiting while eliminating redundant restatement and unproductive reasoning loops.

Why Reasoning Token Bloat Compounds in Agent Workloads

Fireworks notes that reasoning models like K3 can spend more than 90% of generated tokens on internal chain-of-thought. That proportion becomes structurally expensive in multi-turn agentic settings because prior reasoning is replayed into context on each subsequent call. Context size therefore grows roughly quadratically with turn count, meaning long traces from early turns are re-read — and re-billed — on every later invocation. This is precisely the infrastructure-level cost pattern that makes agentic workloads expensive to scale, and token-count reduction at the model level addresses it more cleanly than inference-time controls.

Training Methodology

Fireworks Research ran more than 50 training experiments and over 200 evaluations to build Ember-1, with all training executed on Fireworks Serverless Training infrastructure. The training corpus spans mathematics, coding, instruction following, conversation, search, tool use, and software engineering, covering both single-turn problems and extended multi-step interactions. Task and environment feedback guides on-policy planning, rewarding reasoning that actually corrects mistakes while training out loops that do not. Fireworks states it used its own data exclusively, with no customer data involved. The training algorithms themselves have not been published, and Fireworks has not released Ember-1's weights or training code. The model is available only as a Research Preview through the Fireworks serverless API.

Benchmark Results and Production A/B Data

Fireworks compared Ember-1 against K3 at three effort levels across five benchmarks. All figures are Fireworks-reported, computed using public Kimi K3 API pricing.

Benchmark N K3 Low K3 High K3 Max Ember-1 Cost Δ vs K3 Max
Terminal Bench 2.1 89 76.4% 77.6% 80.9% 82.0% −51.9% / −$23.1
SWE-bench Verified 500 80.4% 86.0% 93.2% 92.2% −15.5% / −$68.1
SWE-Interact 75 6.7% 13.3% 21.3% 20.0% −32.5% / −$60.8
DeepSWE 1.1 113 55.8% 62.8% 66.4% 75.2% −23.7% / −$126.9
τ-2 Bench Airline 50 64% 64% 64% 66% −5.9% / −$0.3

In Fireworks' own evaluations, Ember-1 outperforms K3 Max on Terminal Bench 2.1 and DeepSWE 1.1, trails by 1 percentage point on SWE-bench Verified, and lags by 1.3 points on SWE-Interact. Fireworks also reports that on Doximity's Bedside Bench — a physician-validated set of 500 clinical cases — Ember-1 established a new cost-per-task Pareto frontier, a result Fireworks attributes to its own Specialized Intelligence Index.

The production A/B data is the most operationally concrete figure in the release. Fireworks ran live traffic tests with two customers on production coding workloads; in the published run, output tokens per task fell from 49,300 to 29,900 — a 39% reduction. Reasoning tokens dropped 71.3%. Task score moved from 0.751 to 0.753, and average steps per task fell from 23.8 to 21.4. MarkTechPost reports one of the two customers has since moved Ember-1 into full production.

Pricing is identical to standard K3 on Fireworks: $3.00 per million input tokens, $0.30 per million cached input tokens, and $15.00 per million output tokens, per AI Mastery's pricing data. The savings are entirely a function of generating fewer output tokens. At that output rate, Ember-1 is more expensive per token than DeepSeek V4 Pro ($0.87 output) but cheaper than Claude Sonnet 5 ($10.00 output), per AI Mastery's pricing data — making the economic argument strongest for teams already committed to K3's capability profile.

AI Mastery Analysis

The 40% headline figure deserves careful reading. The production A/B result — 49,300 versus 29,900 output tokens per task, a 39% reduction — is the strongest evidence Fireworks provides, but it covers a single published run on coding traffic from two customers. Whether that compression holds across task types with meaningfully different reasoning structures, such as long-horizon planning versus short synthesis tasks, is not established.

The SWE-bench Verified gap is worth flagging: Ember-1 scores 92.2% against K3 Max's 93.2% on a 500-task sample. Whether a 1-point quality concession is acceptable depends entirely on cost sensitivity and task criticality; for high-stakes software changes, that gap may not be negligible.

The withholding of weights, training code, and algorithm details limits how much practitioners can reason about edge cases, fine-tuning potential, or failure modes. As benchmark framing increasingly functions as a product signal rather than a neutral measurement, the absence of reproducibility infrastructure means independent verification of the core claims remains impossible until Fireworks opens more of the stack.

Post-training for reasoning compression is a reproducible strategy that generalises to any sufficiently capable base model, and Fireworks has demonstrated it can be applied without Moonshot's involvement using only the open weights. If the technique proves robust across more diverse workloads and reaches independent verification, it represents a meaningful addition to the cost-reduction toolkit that sits between model selection and infrastructure-layer optimisation. The Research Preview label signals that Fireworks itself is not yet treating this as a finished product — which is the appropriate level of epistemic caution given where the evidence currently stands.

Sources

Frequently asked questions

How much does Ember-1 cost per million tokens?

According to AI Mastery's pricing data, Ember-1 is priced identically to base Kimi K3 on Fireworks: $3.00 per million input tokens, $0.30 per million cached input tokens, and $15.00 per million output tokens. The cost savings come entirely from generating fewer output tokens, not from a lower per-token rate.

How many fewer tokens does Ember-1 use compared to Kimi K3?

In Fireworks' production A/B test on coding workloads, output tokens fell from 49,300 to 29,900 per task — a 39% reduction — while the task score moved only from 0.751 to 0.753. Reasoning tokens specifically dropped 71.3%. Fireworks reports the headline figure as roughly 40% fewer tokens.

Does Ember-1 outperform Kimi K3 Max on benchmarks?

According to Fireworks' own evaluations, Ember-1 outperforms K3 Max on two of five benchmarks: Terminal Bench 2.1 (82.0% vs 80.9%) and DeepSWE 1.1 (75.2% vs 66.4%). It trails K3 Max by 1 percentage point on SWE-bench Verified (92.2% vs 93.2%) and by 1.3 points on SWE-Interact (20.0% vs 21.3%).

Can I self-host Ember-1 or access its weights?

No. Fireworks has not released Ember-1's weights, training code, or training algorithms. The model is available only as a Research Preview through the Fireworks serverless API, meaning self-hosting is not currently an option.

What is different about Ember-1 versus simply lowering Kimi K3's reasoning effort setting?

Fireworks states that dialing down K3's reasoning effort setting gave up too much quality to be a viable cost lever. Ember-1 instead changes what the model has learned to generate: it retains useful self-correction and assumption-revisiting while eliminating redundant restatement and unproductive reasoning loops — a training-time change rather than an inference-time control.

Free interactive tools for the decisions this piece raises.

Related Reading