Ask
Packaging

Fireworks adds a GB300 GPU tier and a region-restricted deployment premium

Fireworks AI pricing

Fireworks added GB300 288 GB to its on-demand GPU card at $18.00/hr and introduced a 1.5x region-restricted deployment premium for US/Europe-pinned dedicated GPUs, mirroring its existing 10% US-only Serverless token premium.

Before

On-demand dedicated GPU card topped out at B300 288 GB ($12.00/hr); no published geographic-routing premium existed for dedicated deployments.

After

GB300 288 GB is now listed at $18.00/hr, the highest published on-demand rate. Region-restricted deployments (US, Europe) are priced at a flat 1.5x the standard on-demand rate and require a Contact Sales request.

Fireworks’ on-demand dedicated GPU pricing page picked up a new hardware tier and a new pricing axis in the same capture. GB300 288 GB joins H100/H200 ($7.00/hr), B200 ($10.00/hr) and B300 ($12.00/hr) at $18.00/hr — the largest published VRAM-per-GPU option and the highest on-demand rate on the card.

Alongside it, the page now documents region-restricted deployments: GPUs pinned to US-only or Europe-only infrastructure are priced at a flat 1.5x the standard on-demand rate, gated behind a Contact Sales request rather than self-serve checkout. This is the first geographic-routing premium Fireworks has published for dedicated GPU deployments, and it parallels the 10% US-only Serverless premium the company added to its token-based serverless card in July 2026 — extending the same geography-as-a-meter logic from per-token inference to per-GPU-hour deployments.

Separately, the docs clarified that the long-published $50/$500/$5,000/$50,000 monthly spend-tier ceilings apply only to legacy self-serve postpaid accounts; today’s default prepaid accounts set an independent monthly spend limit via firectl quota update monthly-spend-usd, while the tier itself continues to gate serverless TPM ceilings and training-GPU allocation regardless of billing mode.

From Fireworks AI's pricing timeline
GB300 GPU tier + region-restricted deployment premium added

Fireworks added GB300 288 GB to the on-demand dedicated GPU card at $18.00/hr, above B300's $12.00/hr — the highest published on-demand rate on the card. The on-demand pricing page also now documents a region-restricted deployments option (US, Europe) priced at a flat 1.5x the standard on-demand rate, requiring a Contact Sales request — the first geographic-routing premium on the dedicated-GPU side of the product, mirroring the existing 10% US-only Serverless premium on the token side. Separately, the docs now note GLM 5.2 Fast US is exempt from that 10% token-side premium (priced identically to global GLM 5.2 Fast), and clarify that the long-published $50/$500/$5,000/$50,000 spend-tier ceilings apply only to legacy self-serve postpaid accounts, not today's default prepaid accounts.

About Fireworks AI
fireworks.ai ↗

Fireworks AI runs a pure-usage inference platform: serverless per-token endpoints led by Kimi K3 (its 2026-07-29 flagship model, from $3.00/1M input on Standard) alongside other open-weight models (DeepSeek, GLM, Qwen), plus on-demand dedicated GPU deployments at $7/hr H100/H200, $10/hr B200, $12/hr B300 — rates that Fireworks has announced will rise to $8/$13/$15 respectively (GB300 $18 to $20) effective September 1, 2026.

Free tier
Yes
Commits
Available
Transparency
public

Fireworks AI pricing history

  1. Aug 2026
    DeepSeek V4 Flash (0731) repriced; Qwen 3.8 Max added
  2. Aug 2026
    Two new serverless models added — Muse Glimmer 30B and Nemotron 3.5 Lightning
  3. Aug 2026
    On-demand GPU price increase announced, effective September 1, 2026
  4. Aug 2026
    GB300 GPU tier + region-restricted deployment premium added
  5. Jul 2026
    Kimi K3 flagship, US-only Serverless premium, and Serverless Training API
Full Fireworks AI timeline

More Fireworks AI activity

All pricing activity