Ask
Packaging

Embeddings priced by parameter size

Fireworks AI pricing

Fireworks introduces parameter-size-based pricing for its embeddings models.

From Fireworks AI's pricing timeline
Embeddings Pricing by Parameter Size

Fireworks published differential embeddings pricing by base model parameter count: <150M params at $0.008/1M tokens, 150–350M at $0.016/1M, Qwen3 8B at $0.10/1M. The schedule undercuts OpenAI text-embedding-3-small ($0.02/1M) by 60% for the smallest tier and creates a granular cost ladder for retrieval-pipeline cost optimization.

About Fireworks AI
fireworks.ai ↗

Fireworks AI runs a pure-usage inference platform: serverless per-token endpoints led by Kimi K3 (its 2026-07-29 flagship model, from $3.00/1M input on Standard) alongside other open-weight models (DeepSeek, GLM, Qwen), plus on-demand dedicated GPU deployments at $7/hr H100/H200, $10/hr B200, $12/hr B300 — rates that Fireworks has announced will rise to $8/$13/$15 respectively (GB300 $18 to $20) effective September 1, 2026.

Free tier
Yes
Commits
Available
Transparency
public

Fireworks AI pricing history

  1. Aug 2026
    DeepSeek V4 Flash (0731) repriced; Qwen 3.8 Max added
  2. Aug 2026
    Two new serverless models added — Muse Glimmer 30B and Nemotron 3.5 Lightning
  3. Aug 2026
    On-demand GPU price increase announced, effective September 1, 2026
  4. Aug 2026
    GB300 GPU tier + region-restricted deployment premium added
  5. Jul 2026
    Kimi K3 flagship, US-only Serverless premium, and Serverless Training API
Full Fireworks AI timeline

More Fireworks AI activity

All pricing activity