Ask
Launch

Fireworks launches Kimi K3 flagship model and a pay-per-token Serverless Training API

Fireworks AI pricing

Fireworks added Kimi K3 (plus Fast and a 10%-premium US-only variant) to its serverless rate card, shipped a new Serverless Training API for per-token LoRA training, and switched Fire Pass's free model to Kimi K3 Fast with a 1M-token context window.

Before

Headline serverless models topped out at Kimi K2.7/K2.6 and GLM 5.x; no serverless training product existed; Fire Pass granted free access to GLM 5.2 Fast (256k context).

After

Kimi K3 leads the rate card at $3.00 / $0.30 / $15.00 per 1M input/cached/output tokens (Standard), with Fast ($4.50/$0.45/$22.50) and a US-only variant priced at a flat 10% premium ($3.30/$0.33/$16.50). A new Serverless Training API prices LoRA training on Qwen 3.5 9B, Qwen 3.6 27B and Kimi K3 by prefill/cached-prefill/sample/train tokens ($0.66-$32.55 per 1M depending on stage and model). Fire Pass now grants Kimi K3 Fast (1M context) instead of GLM 5.2 Fast.

Fireworks used its site banner to promote Kimi K3 as a new flagship model on 2026-07-29, adding it to the serverless price card alongside a Fast variant and a US-only-routing variant. The US variant is the first model to carry its own priced row for a broader mechanic newly documented in the docs: US-only Serverless endpoints are billed at a flat 10% premium over the base model’s serverless price on any model, not just Kimi K3.

The bigger structural change is a new product: the Serverless Training API, a Tinker-compatible offering that attaches to a shared, always-on trainer pool for LoRA training with no provisioning step and no idle cost. It bills per token across four dimensions — prefill, cached prefill, sample, and train — currently covering three models (Qwen 3.5 9B, Qwen 3.6 27B, Kimi K3), with more “coming soon” per the docs.

Fire Pass, the promo-code pass launched a week earlier with zero per-token pricing on one open-weight model, swapped its included model from GLM 5.2 Fast (256k context) to Kimi K3 Fast (1M context) and added FireConnect auto-detection plus new supported harnesses (Codex, Pi, LangChain Deep Agents). The rest of the rate card is unchanged: H100/H200 dedicated at $7.00/hr, B200 at $10.00/hr, B300 at $12.00/hr, fine-tuning from $0.50 per 1M training tokens, embeddings from $0.008 per 1M, and batch inference at 50% of serverless.

From Fireworks AI's pricing timeline
Kimi K3 flagship, US-only Serverless premium, and Serverless Training API

Fireworks added Kimi K3 (plus a Fast variant and a US-only variant) to its serverless rate card, publishing the first priced row for a new US-only Serverless mechanic — a flat 10% premium over base serverless pricing that the docs describe as applying to any model. Fireworks also shipped the Serverless Training API, a Tinker-compatible product that meters LoRA training per token (prefill, cached prefill, sample, train) on a shared trainer pool with no provisioning or idle cost, covering Qwen 3.5 9B, Qwen 3.6 27B, and Kimi K3. Fire Pass's included free model swapped from GLM 5.2 Fast (256k context) to Kimi K3 Fast (1M context).

About Fireworks AI
fireworks.ai ↗

Fireworks AI runs a pure-usage inference platform: serverless per-token endpoints led by Kimi K3 (its 2026-07-29 flagship model, from $3.00/1M input on Standard) alongside other open-weight models (DeepSeek, GLM, Qwen), plus on-demand dedicated GPU deployments at $7/hr H100/H200, $10/hr B200, $12/hr B300 — rates that Fireworks has announced will rise to $8/$13/$15 respectively (GB300 $18 to $20) effective September 1, 2026.

Free tier
Yes
Commits
Available
Transparency
public

Fireworks AI pricing history

  1. Aug 2026
    DeepSeek V4 Flash (0731) repriced; Qwen 3.8 Max added
  2. Aug 2026
    Two new serverless models added — Muse Glimmer 30B and Nemotron 3.5 Lightning
  3. Aug 2026
    On-demand GPU price increase announced, effective September 1, 2026
  4. Aug 2026
    GB300 GPU tier + region-restricted deployment premium added
  5. Jul 2026
    Kimi K3 flagship, US-only Serverless premium, and Serverless Training API
Full Fireworks AI timeline

More Fireworks AI activity

All pricing activity