Fireworks Ember-1 Cuts Kimi K3 Tokens by 5.9% to 51.9% at the Same Price
Fireworks released Ember-1, a post-trained Kimi K3, on September 23 and shipped it to OpenRouter and Vercel AI Gateway on September 27. Per-token pricing matches K3 exactly at $3 in and $15 out, while token savings across five benchmarks range from 5.9% to 51.9%.
- Fireworks post-trained Kimi K3 to think less and priced it identically.
- Token savings swing from 5.9% to 51.9% depending on the benchmark.
- Weights stay closed, so it only runs on the
fireworks/ember-1endpoint.
A reasoning model's bill does not come from the answer. It accumulates while the model thinks to itself before answering. Ember-1, released by Fireworks Research on September 23, is a model built entirely by post-training that thinking to be shorter. The base is Moonshot AI's Kimi K3, and when it landed on OpenRouter and Vercel AI Gateway on September 27, anyone already routing through those gateways could switch by changing the model name.
This is not a model with new capabilities. Given the same task, it makes the model say less on the way there, which shrinks the output-token line on the invoice. That is the whole product.
The only thing that changed from Kimi K3 is reasoning length
Ember-1 is the same 2.78T-parameter MoE model as Kimi K3, with the same 1,040K-token context window. The price is the same too: $3 per million input tokens, $0.30 for cached input, $15 for output. Since the unit price is identical, every dollar saved comes from token count alone.
Fireworks names Kimi K3's token split as the starting point. "Reasoning models like Kimi K3 spend the majority of their generated tokens, sometimes more than 90%, on internal reasoning." The answer the user actually reads is the remaining 10% or so.
Turning the dial down did not work, they write. "Turning down K3's reasoning effort didn't solve this. Lower effort settings gave up too much quality." Kimi K3 accepts reasoning_effort at low, high, and max, defaulting to max. Instead of that dial, Fireworks retrained the model itself. They ran more than 50 experiments and 200 evaluations across math, coding, instruction following, conversation, search, tool use, and software engineering, training the model to keep self-reflection while dropping the reasoning loops that circle the same spot.
Here are the five benchmarks from the announcement. The comparison column labeled "K3 Max" is Kimi K3 at its default reasoning_effort=max.
| Benchmark | Samples | Kimi K3 | Ember-1 | Token reduction |
|---|---|---|---|---|
| Terminal Bench 2.1 | 89 | 80.9% | 82.0% | 51.9% |
| SWE-bench Verified | 500 | 93.2% | 92.2% | 15.5% |
| SWE-Interact | 75 | 21.3% | 20.0% | 32.5% |
| DeepSWE 1.1 | 113 | 66.4% | 75.2% | 23.7% |
| τ-2 Bench Airline | 50 | 64% | 66% | 5.9% |
On score alone, three of five went up and two went down. DeepSWE 1.1 moved the most, from 66.4% to 75.2%, a gain of 8.8 points. The drops are much smaller: 1.3 points on SWE-Interact and 1.0 point on SWE-bench Verified. With sample sizes between 50 and 500, though, a difference of roughly a point is hard to read as a real improvement either way.
"40% fewer tokens" was a different number on every benchmark
The headline claim in the announcement is "40% fewer tokens." Not one of the benchmarks actually lands near 40%. Sorted by size, the reductions in the table above look like this.
Source: the benchmark table in the official Fireworks announcement. Bar length is the reduction in generated tokens against Kimi K3.
Top to bottom, that is almost a nine-fold spread. Converted to dollars, the gap widens further. Running all 50 τ-2 Bench Airline cases saved $0.30, while 113 DeepSWE 1.1 cases saved $126.90.
The spread comes from the shape of the task. The longer a task made the model think in the first place, the more there is to cut. Terminal Bench 2.1 stretches reasoning out because the model works through many steps in a terminal. τ-2 Bench Airline, at the other end, is a few turns of airline-booking dialogue, so reasoning was already short and there is nothing to trim. Depending on which end your workload sits closer to, the saving is either 50% or stops at 6%.
What two customers measured in production
Separately from the benchmarks, Fireworks published results from A/B tests on coding workloads at two customers.
Output tokens fell 39% and the task score went from 0.751 to 0.753, which is effectively flat. Fireworks calls the result "no news is good news." "Developers carried on without noticing the switch, consuming substantially fewer tokens." These are numbers Fireworks measured itself at two customers, and it has not said what the work was.
Cheaper, but it did not take the whole frontier
The other metric the announcement leads with is Doximity's Bedside Bench. On a set of 500 clinical cases, Fireworks says Ember-1 set a new Pareto frontier against open and closed models including GPT-5.6 Sol, GPT-6 Astra, Claude Opus 5, and GLM 5.3. Put cost per task on the x-axis and Ember-1 lands almost the same score as Kimi K3 from further left.
Plot the same benchmark against time instead, and Ember-1 drops off the frontier.

Ember-1 does sit left of Kimi K3. Fewer generated tokens means it finishes sooner. But the models tracing the frontier line are Gemini 3.8 Flash, GPT-5.6 Sol, GPT-6 Astra, and Claude Opus 5, and Ember-1 sits below it. GPT-5.6 Sol scores higher in less time than Ember-1 takes. If latency rather than billing is your bottleneck, switching to Ember-1 does not buy you much.
The objections raised on Hacker News
The Hacker News thread posted on September 27 drew 584 points and 249 comments. The pushback ran along five lines.
- The table does not say what was lost.
tomrodwrote that the benchmarks show only half the picture: nothing separates out which capabilities were shaved off in exchange for the tokens. Commenters cited a case of regressed performance on golang work. - It does not dominate the way the claim suggests.
7777777philcountered that Opus 5.5 and GPT-6 Sol lead Ember-1 on specific metrics. - A cheaper option already exists. According to
netvarun, GLM 5.3 delivers performance close to Kimi K3 at less than half the price. Cutting tokens 40% buys nothing if the unit price is twice as high. - The method description is vague. On phrases like "task and environment feedback" and "on-policy planning,"
kingstnapwrote that they sound scientific but carry no information. Neither the training algorithm nor the data has been published. - What "open weights" covers. As
intothemildpointed out, Kimi K3 itself, the base model, carries enough license restriction that calling those weights freely usable is a stretch.
Ember-1 itself ships with neither weights nor training code. Fine-tuning is blocked and self-hosting is not possible. Verifying any of this independently means calling the API and measuring the results.
What it takes to turn it on
These conditions are confirmed against primary sources.
| Item | Detail |
|---|---|
| Who can use it | Anyone. No waitlist and no application step |
| Pricing | $3 per million input tokens, $0.30 cached input, $15 output. Identical to Kimi K3 |
| Context | 1,040K tokens. Text and image input, function calling, and implicit prompt caching |
| Status | Research Preview. Starts as two weeks of serverless access, made permanent based on demand |
| Regional availability | Fireworks, OpenRouter, and Vercel AI Gateway all publish no country restrictions, so there is no geographic gate on calling it from APAC or anywhere else. Billing is a USD card |
| Data residency | Serverless only, and Fireworks names no regions, so you cannot pin inference to ap-southeast-1 or any other region. The controls on offer are Zero Data Retention and No Prompt Training on the Fireworks endpoint. If Singapore PDPA obligations or a customer contract require a named processing location, that is not something this preview answers |
| Constraints | Fireworks serverless only. Closed weights, no fine-tuning, no self-hosting |
The model ID is fireworks/ember-1 on all three routes. The endpoint is OpenAI compatible, so switching existing code is a one-line change.
curl https://openrouter.ai/api/v1/chat/completions \
-H "Authorization: Bearer $OPENROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "fireworks/ember-1",
"messages": [{"role": "user", "content": "Find out why the tests in this repo are failing"}]
}'
If you already run agents on Kimi K3, checking this takes half a day. The unit price is the same, so all you need is to split traffic across the two models and compare usage.completion_tokens on the responses. Look at the number from your own tasks, not the 40% benchmark average. As the table above shows, the same model varies from 5.9% to 51.9% by task, and multi-step coding work is where the reduction runs largest. The two-week preview window counts from September 22, so it may close in early October, and whether the model becomes permanent depends on demand gathered by then. Measuring once before that leaves you a basis for picking the next model even if the window shuts. If the saving comes in below expectations, the harness is the next thing to touch: the three defaults in devlery's earlier write-up on cutting token cost 28% with the AWS Strands harness work on the other side of the same invoice.