The Spectrum Dispatch News

technology

Fireworks Launches Ember-1, a Token‑Efficient Kimi K3 Variant

The new model matches Kimi K3 quality while using 40 % fewer tokens, according to internal benchmarks and live A/B tests.

Fireworks Launches Ember-1, a Token‑Efficient Kimi K3 Variant

Fireworks Research has introduced Ember-1, a specialized model built on Kimi K3 that delivers the same quality with 40 % fewer tokens, according to the company’s blog. The model was created after users reported that Kimi K3’s long reasoning traces made automated coding expensive at scale, and that simply lowering reasoning effort sacrificed too much quality. To solve this, Fireworks ran more than 50 training experiments and over 200 evaluations, developing new training algorithms on its Serverless Training platform, which eliminated the need for GPU provisioning and reduced experimentation time and cost. Ember-1 was trained on a broad set of tasks using only internal data, with no customer data involved. Evaluation on the Specialized Intelligence Index, public benchmarks, and live production traffic showed that Ember-1 uses fewer tokens without a drop in quality. The company notes that reasoning models like Kimi K3 can spend more than 90 % of generated tokens on internal reasoning, a burden that grows quadratically in multi‑turn agentic workloads. Experiments indicated that Kimi K3’s reasoning could be shortened by 35–50 % without losing accuracy across seven benchmarks and two customers’ production traffic. On the Doximity Bedside Bench—a physician‑validated set of 500 clinical cases across ten categories—Ember-1 set a new Pareto frontier for cost/task among open and closed models including GPT-5.6 Sol, GPT-6 Astra, and Claude Opus 5. Across other industry benchmarks with more than 50 test samples, Ember-1 sits on or near the quality‑vs‑cost Pareto frontier, matching Kimi K3’s maximum quality at a fraction of the cost and strictly dominating K3‑low effort settings. Analyses also positioned Ember-1 as a leader on the frontier when compared with GPT-6 Astra, Claude Opus‑5, and GLM 5.3. Detailed results from head‑to‑head comparisons show improvements such as a rise from 77.6 % to 80.9 % on Terminal Bench 2.1, a gain from 93.2 % to 92.2 % on SWE‑bench Verified, and an increase from 66.4 % to 75.2 % on DeepSWE 1.1, accompanied by token‑cost reductions ranging from‑0.3 USD to‑126.9 USD per test. In live A/B tests with two customers on production coding workloads, Ember-1 delivered approximately 35 % fewer tokens per task at comparable quality, while downstream metrics like task completion and success scores held or improved; one customer has already deployed Ember-1 in production and plans to scale it to replace the base model. Internally, Fireworks observed a reasoning token drop from 49.3K to 29.9K and a total token reduction of 39 % when switching from Kimi K3 to Ember-1, with developers reporting no noticeable change in workflow. Ember-1 is now offered as a serving option alongside Kimi K3 as a Research Preview on Serverless, with two‑week research releases intended to gather community feedback before potential permanent availability.

Fireworks Launches Ember-1, a Token‑Efficient Kimi K3 Variant

Key facts

  • Ember-1 delivers Kimi K3 quality with 40 % fewer tokens (source).
  • Training involved >50 experiments, >200 evaluations, new algorithms on Serverless Training (source).
  • Live A/B tests showed ~35 % fewer tokens per task at comparable quality (source).
  • Ember-1 set a new Pareto frontier on Doximity Bedside Bench (source).

Sources

← All posts