Ask
Price change

DeepSeek V4 Flash (0731) quietly repriced on Fireworks' serverless card

Fireworks AI pricing

Fireworks repriced DeepSeek V4 Flash (0731): Standard input rose from $0.14 to $0.22 per 1M tokens (+57%), cached input fell from $0.028 to $0.007 (-75%), and output rose from $0.28 to $0.66 (+136%), with no forward-notice framing on the page.

Before

DeepSeek V4 Flash (Standard): $0.14 / $0.028 / $0.28 per 1M tokens (input / cached input / output), published as two identical rows -- an undated "DeepSeek V4 Flash" and a dated "DeepSeek V4 Flash (0731)".

After

DeepSeek V4 Flash (0731) (Standard): $0.22 / $0.007 / $0.66 per 1M tokens; Priority moved to $0.275 / $0.00875 / $0.825. The undated duplicate row has been removed, leaving the dated (0731) row as the sole DeepSeek V4 Flash entry.

Fireworks repriced an already-published serverless model rather than adding a new tier — a departure from its usual pattern of expanding the catalog with new rows at new price points. DeepSeek V4 Flash (0731) now costs 57% more on input, 136% more on output, and 75% less on cached input than the identical rate it shared with an undated “DeepSeek V4 Flash” row as recently as August 14, 2026; that undated row has since been removed from the pricing docs, leaving the dated row as the sole entry for the model.

Unlike Fireworks’ August 12, 2026 on-demand GPU increase, which was announced with a forward effective date and a two-column current/future table, this repricing carries no such notice on the page — it simply reads as already in effect for anyone calling the model today. Separately, in the same capture window, Fireworks added a new row, Qwen 3.8 Max, to the serverless card at $2.00 / $0.25 / $6.00 per 1M tokens (Standard-only).

From Fireworks AI's pricing timeline
DeepSeek V4 Flash (0731) repriced; Qwen 3.8 Max added

Fireworks silently repriced DeepSeek V4 Flash (0731) on the serverless rate card — Standard input rose from $0.14 to $0.22 per 1M tokens (+57%), cached input fell from $0.028 to $0.007 (−75%), and output rose from $0.28 to $0.66 (+136%); Priority moved from $0.21 / $0.042 / $0.42 to $0.275 / $0.00875 / $0.825 on the same dimensions. The separate undated "DeepSeek V4 Flash" row (previously identical to the (0731) row) has been removed, leaving the dated row as the sole DeepSeek V4 Flash entry. Unlike the forward-dated, announced on-demand GPU increase from 2026-08-12, this repricing carries no effective-date framing on the page — it reads as already in effect. Separately, a new row, Qwen 3.8 Max, was added to the serverless card at $2.00 / $0.25 / $6.00 (Standard-only, no Priority path).

About Fireworks AI
fireworks.ai ↗

Fireworks AI runs a pure-usage inference platform: serverless per-token endpoints led by Kimi K3 (its 2026-07-29 flagship model, from $3.00/1M input on Standard) alongside other open-weight models (DeepSeek, GLM, Qwen), plus on-demand dedicated GPU deployments at $7/hr H100/H200, $10/hr B200, $12/hr B300 — rates that Fireworks has announced will rise to $8/$13/$15 respectively (GB300 $18 to $20) effective September 1, 2026.

Free tier
Yes
Commits
Available
Transparency
public

Fireworks AI pricing history

  1. Aug 2026
    DeepSeek V4 Flash (0731) repriced; Qwen 3.8 Max added
  2. Aug 2026
    Two new serverless models added — Muse Glimmer 30B and Nemotron 3.5 Lightning
  3. Aug 2026
    On-demand GPU price increase announced, effective September 1, 2026
  4. Aug 2026
    GB300 GPU tier + region-restricted deployment premium added
  5. Jul 2026
    Kimi K3 flagship, US-only Serverless premium, and Serverless Training API
Full Fireworks AI timeline

More Fireworks AI activity

All pricing activity