Ask
Packaging

DeepInfra's self-serve Claude catalog narrows from seven models to one

DeepInfra pricing

DeepInfra's /pricing page now lists only claude-haiku-4-5 ($1.00/$5.00 per 1M); six other Claude models, including the $10.00/$50.00 claude-fable-5 row, are gone, down from seven models shown as of 2026-08-26.

Before

Seven Claude models on /pricing: claude-opus-5, claude-fable-5 ($10.00/$50.00), claude-sonnet-5, claude-haiku-4-5, claude-sonnet-4-6, claude-opus-4-7, claude-opus-4-8.

After

One Claude model on /pricing: claude-haiku-4-5 ($1.00 in / $5.00 out per 1M tokens).

As of the 2026-08-28 capture, DeepInfra’s public pricing page shows only one Anthropic model — claude-haiku-4-5 at $1.00 in / $5.00 out per 1M tokens — in its Claude section, down from seven models captured just two days earlier on 2026-08-26 (claude-opus-5, claude-fable-5, claude-sonnet-5, claude-haiku-4-5, claude-sonnet-4-6, claude-opus-4-7, claude-opus-4-8). The removed models included claude-fable-5 at $10.00/$50.00 per 1M, which had been the single highest per-token rate on DeepInfra’s entire price list; Kimi-K3 ($2.85/$14.25 per 1M) is now the highest-priced model shown on the page. No announcement or dated changelog entry accompanies the change, consistent with DeepInfra’s practice of publishing pricing moves as silent page edits rather than notices — it isn’t possible to tell from the pricing page alone whether this reflects a licensing change with Anthropic, a DeepInfra catalog decision, or a temporary listing gap.

From DeepInfra's pricing timeline
Six of seven Claude models pulled from the /pricing page; GLM-5.2 gets a new 35% promo

DeepInfra's /pricing page Claude section drops from seven models (as of 2026-08-26) to one: claude-haiku-4-5 ($1.00/$5.00) is now the only Claude model listed, with claude-opus-5, claude-fable-5 ($10.00/$50.00, the former highest per-token rate on the page), claude-sonnet-5, claude-sonnet-4-6, claude-opus-4-7, and claude-opus-4-8 no longer shown. Kimi-K3 ($2.85 in / $14.25 out per 1M) is now the highest per-token rate on the page. Separately, GLM-5.2 — already only listed on the /models and /deepstart catalogs, not /pricing — picks up a new "35% off" promotional tag cutting its post-2026-08-04-cut rate of $0.75/$2.40 ($0.14 cached) to $0.488/$1.56 ($0.091 cached) per 1M (source: deepinfra.com/pricing, /models, /deepstart 2026-08-28).

About DeepInfra
deepinfra.com ↗

DeepInfra is a serverless inference cloud that bills per-token for language and embedding models and per-inference-execution-time for most other models, with no contracts or upfront costs. Representative per-1M-token rates: DeepSeek-V3.1 $0.25 in / $0.95 out, DeepSeek-V4-Pro $1.30 / $2.60, Llama-3.3-70B-Turbo $0.10 / $0.32, Llama-3.1-8B $0.02 / $0.04. Llama-3.1-8B-Instruct-Turbo's output rate rose 33% on 2026-07-29 (from $0.03), DeepInfra's second token-price increase since its 2026-07-14 reversal, while gemma-4-31B-it-turbo was cut to $0.09 in / $0.34 out per 1M in the same update. GLM-5.2, Z-AI's flagship long-horizon model featured on DeepInfra's /models catalog, was then cut about 20% on 2026-08-04 (from $0.93 in / $3.00 out to $0.75 / $2.40 per 1M), the first outright cut since the July reversal began, though its price never appears on the main /pricing page. GLM-5.2 was cut again on 2026-08-28, this time via a new 35%-off promotional tag taking it from $0.75 / $2.40 to $0.488 / $1.56 per 1M — still only visible on /models and /deepstart.

Free tier
No
Commits
Available
Transparency
public

DeepInfra pricing history

  1. Aug 2026
    Six of seven Claude models pulled from the /pricing page; GLM-5.2 gets a new 35% promo
  2. Aug 2026
    GLM-5.2 cut ~20% — but only on the /models catalog
  3. Jul 2026
    Second rate rise: Llama-3.1-8B-Instruct-Turbo output up 33%
  4. Jul 2026
    Flex service tier added at 0.8× base price
  5. Jul 2026
    GPU-hour rates raised; DeepSeek-V3.1 token price up
Full DeepInfra timeline

More DeepInfra activity

All pricing activity