Ask
Price change

DeepInfra hikes its cheapest small model output price 33%

DeepInfra pricing

DeepInfra raised Llama-3.1-8B-Instruct-Turbo output price 33% to $0.04 per 1M tokens (from $0.03) and cut gemma-4-31B-it-turbo to $0.09/$0.34 per 1M.

Before

Meta-Llama-3.1-8B-Instruct-Turbo $0.02 in / $0.03 out per 1M tokens; gemma-4-31B-it-turbo $0.12 in / $0.37 out per 1M tokens.

After

Meta-Llama-3.1-8B-Instruct-Turbo $0.02 in / $0.04 out per 1M tokens (plus 33% on output); gemma-4-31B-it-turbo $0.09 in / $0.34 out per 1M tokens (minus 25% input, minus 8% output).

DeepInfra’s price list moved on two individual model SKUs between the 2026-07-21 and 2026-07-29 captures. Meta-Llama-3.1-8B-Instruct-Turbo — the model this page and DeepInfra’s own catalog repeatedly cite as the cheapest small-model reference — rose from $0.02 in / $0.03 out to $0.02 in / $0.04 out per 1M tokens, a 33 percent increase on the output rate. In the same capture, gemma-4-31B-it-turbo was cut from $0.12/$0.37 to $0.09/$0.34 per 1M tokens. Headline per-GPU-hour rates, DeepCluster reserved pricing, usage tiers, and the three-step Flex/Standard/Priority service tier were all unchanged. This is the second rate increase DeepInfra has published since its 2026-07-14 GPU-hour hike broke an eighteen-month streak of price cuts, suggesting the July reversal was not a one-off.

From DeepInfra's pricing timeline
Second rate rise: Llama-3.1-8B-Instruct-Turbo output up 33%

DeepInfra raises Meta-Llama-3.1-8B-Instruct-Turbo — the model this page and DeepInfra's own catalog cite as the cheapest small-model reference — from $0.02 in / $0.03 out to $0.02 in / $0.04 out per 1M tokens (+33% on output), while cutting gemma-4-31B-it-turbo from $0.12/$0.37 to $0.09/$0.34 per 1M in the same update. GPU-hour, DeepCluster, usage-tier, and service-tier rates are unchanged; this is the second rate increase since the 2026-07-14 reversal, suggesting it was not a one-off (source: deepinfra.com/pricing 2026-07-29).

About DeepInfra
deepinfra.com ↗

DeepInfra is a serverless inference cloud that bills per-token for language and embedding models and per-inference-execution-time for most other models, with no contracts or upfront costs. Representative per-1M-token rates: DeepSeek-V3.1 $0.25 in / $0.95 out, DeepSeek-V4-Pro $1.30 / $2.60, Llama-3.3-70B-Turbo $0.10 / $0.32, Llama-3.1-8B $0.02 / $0.04. Llama-3.1-8B-Instruct-Turbo's output rate rose 33% on 2026-07-29 (from $0.03), DeepInfra's second token-price increase since its 2026-07-14 reversal, while gemma-4-31B-it-turbo was cut to $0.09 in / $0.34 out per 1M in the same update. GLM-5.2, Z-AI's flagship long-horizon model featured on DeepInfra's /models catalog, was then cut about 20% on 2026-08-04 (from $0.93 in / $3.00 out to $0.75 / $2.40 per 1M), the first outright cut since the July reversal began, though its price never appears on the main /pricing page. GLM-5.2 was cut again on 2026-08-28, this time via a new 35%-off promotional tag taking it from $0.75 / $2.40 to $0.488 / $1.56 per 1M — still only visible on /models and /deepstart.

Free tier
No
Commits
Available
Transparency
public

DeepInfra pricing history

  1. Aug 2026
    Six of seven Claude models pulled from the /pricing page; GLM-5.2 gets a new 35% promo
  2. Aug 2026
    GLM-5.2 cut ~20% — but only on the /models catalog
  3. Jul 2026
    Second rate rise: Llama-3.1-8B-Instruct-Turbo output up 33%
  4. Jul 2026
    Flex service tier added at 0.8× base price
  5. Jul 2026
    GPU-hour rates raised; DeepSeek-V3.1 token price up
Full DeepInfra timeline

More DeepInfra activity

All pricing activity