Ask
All companies
technology

DeepInfra pricing

deepinfra.com facts checked analysis reviewed
Estimate your DeepInfra cost — model your usage, see overages, and find the cheapest plan. Open calculator →
Quick summary
Product
Serverless inference cloud — per-token LLM/embedding APIs, per-image and per-minute media models, per-hour on-demand GPU containers, and reserved DeepCluster GPU clusters
Industry
technology
Commits
Available (annual)
In this page
AI Summary
  • DeepInfra is a serverless inference cloud that bills per-token for language and embedding models and per-inference-execution-time for most other models, with no contracts or upfront costs. Representative per-1M-token rates: DeepSeek-V3.1 $0.25 in / $0.95 out, DeepSeek-V4-Pro $1.30 / $2.60, Llama-3.3-70B-Turbo $0.10 / $0.32, Llama-3.1-8B $0.02 / $0.04. Llama-3.1-8B-Instruct-Turbo's output rate rose 33% on 2026-07-29 (from $0.03), DeepInfra's second token-price increase since its 2026-07-14 reversal, while gemma-4-31B-it-turbo was cut to $0.09 in / $0.34 out per 1M in the same update. GLM-5.2, Z-AI's flagship long-horizon model featured on DeepInfra's /models catalog, was then cut about 20% on 2026-08-04 (from $0.93 in / $3.00 out to $0.75 / $2.40 per 1M), the first outright cut since the July reversal began, though its price never appears on the main /pricing page. GLM-5.2 was cut again on 2026-08-28, this time via a new 35%-off promotional tag taking it from $0.75 / $2.40 to $0.488 / $1.56 per 1M — still only visible on /models and /deepstart.
  • Hosted closed models are also priced per-token: Gemini 2.5 Pro $1.25 / $10.00, Kimi-K3 $2.85 / $14.25 per 1M tokens. As of 2026-08-28, DeepInfra's /pricing page lists only one Claude model — claude-haiku-4-5 at $1.00 / $5.00 — down from seven as recently as 2026-08-26; Claude Sonnet, Claude Opus, and claude-fable-5 no longer appear on the public pricing catalog. Embeddings run $0.005–$0.01 per 1M input tokens; Voxtral audio is billed per minute ($0.00100–$0.00300); Flux image generation is priced per image scaled by resolution and iteration count; and Qwen3-TTS text-to-speech is billed at $20.00 per 1M characters of input text, a fifth metering unit on the same DeepInfra account.
  • On-demand GPU containers bill per minute (shown as per-hour): A100 $0.89, H100 $2.20, H200 $2.69, B200 $3.69, B300 $4.89 per GPU-hour for custom-LLM deployments; the dedicated GPU-instances surface lists 1×B200 $3.69/hr, 2×B200 $7.38/hr, 4×B200 $14.76/hr, 8×B200 $29.52/hr with no egress fees. These GPU-hour rates rose 16% to 32% on 2026-07-14 — the first price increase in DeepInfra's tracked pricing history, with only the A100 held flat at $0.89.
  • DeepCluster is a reserved, customer-owned NVIDIA B300 GPU cluster (256–5,000 GPUs) that DeepInfra procures and operates under 3-to-5-year terms: all-in $2.99/GPU-hr on a 3-year term and $1.98/GPU-hr on a 5-year term, marketed as up to 70% cheaper than a $6.50/GPU-hr public-cloud reference.
  • Billing requires a card on file or pre-payment, with automatic usage tiers (Tier 1 $20 invoicing threshold up to Tier 5 $10,000) that move accounts up as cumulative spend grows; invoices generate monthly and whenever the tier threshold is hit. Accounts are capped at 200 concurrent requests by default, and customers can set a spending limit to avoid surprises.
  • DeepInfra has no published free tier (card or pre-pay required), but runs the DeepStart program granting qualifying early-stage startups 1 billion free tokens. It raised a $107M Series B and positions on transparent open-model inference pricing against Together, Fireworks, Replicate, and RunPod. It also offers a three-step per-request Service Tier: Flex bills 0.8× the base per-token price for non-production and asynchronous work, the default Standard tier bills 1×, and Priority bills 1.5× for faster time-to-first-token during peak demand (set via service_tier "priority").
Pricing summary
DeepInfra 2026 — pure-usage inference cloud across five metering units
Per-token LLM/embedding APIs + per-image, per-audio-minute and per-character media + per-GPU-hour containers + reserved DeepCluster — pay only for what you use.
Serverless inference (per token)
$0.02–$2.85 / 1M in
Developers calling open + hosted LLMs by API
256–5,000 GPUs
DeepCluster (reserved)
$1.98–$2.99 / GPU-hr
Teams needing dedicated B300 capacity at scale
DeepStart credits
1B tokens
Startups raised 250K–10M USD, founded ≤2 years ago
All prices read from deepinfra.com/pricing, /gpu-instances and /deepcluster (USD, accessed 2026-08-28). Per-token rates vary by model; representative rows shown. A per-request service tier multiplies the base per-token rate: Flex 0.8×, Standard 1×, Priority 1.5×.

About

DeepInfra is a serverless inference cloud that runs open-weight and hosted AI models on its own GPU fleet and bills customers only for what they consume. Its core promise — stated on the pricing page itself — is “you only pay for what you use… no long-term contracts or upfront costs,” with language models billed per token and most other models billed for inference execution time. The catalog spans hundreds of open models (DeepSeek, Qwen, Llama 3/4, Gemma, Mistral, Nemotron, Phi) alongside hosted closed models (Claude, Gemini), embedding models, Flux image generation, and Voxtral audio transcription, all reachable from one account and one bill.

The company sells to developers, individual builders, and engineering teams that want frontier-model inference without standing up their own GPU infrastructure, and it competes directly with Together AI, Fireworks, Replicate, and RunPod on open-model price and breadth. Founded in 2022 by the team behind the imo messenger (200M+ users), DeepInfra closed a $107M Series B on 2026-05-04 — co-led by 500 Global and Georges Harik, with NVIDIA, Samsung Next, and Supermicro participating — and reports processing roughly five trillion tokens per week, 25× the volume at its Series A. It runs models on H100 and A100 GPUs optimized for inference, with automatic scaling and a default cap of 200 concurrent requests per account.

Beyond the per-token API, DeepInfra layers two heavier compute products on top: on-demand GPU containers (per-GPU-hour A100 through B300 cards, billed by the minute) for custom-model and training workloads, and DeepCluster — a reserved, customer-owned NVIDIA B300 cluster of 256 to 5,000 GPUs that DeepInfra procures, deploys, and operates under three-to-five-year terms. The result is a single vendor spanning the full spectrum from a single API call to a multi-million-dollar dedicated cluster.


Pricing summary : How DeepInfra’s per-token, per-GPU-hour, and reserved-cluster pricing works

DeepInfra runs a pure-usage model with no seats and no base platform fee, metered across five distinct billing units (tokens, images, audio minutes, TTS characters, GPU-hours) grouped into four self-serve dimensions plus a reserved-commitment tier:

  1. Per-token inference (LLMs + embeddings): Language models are priced per 1M input and output tokens, varying by model across a roughly 142× span — Llama-3.1-8B at $0.02 / $0.04, DeepSeek-V3.1 at $0.25 / $0.95, up to Kimi-K3 at $2.85 / $14.25, the highest per-token rate now shown on the pricing page. As of 2026-08-28 the page’s Claude section lists only claude-haiku-4-5 ($1.00 / $5.00) — the six other Claude models this page previously tracked (claude-opus-5, claude-fable-5, claude-sonnet-5, claude-sonnet-4-6, claude-opus-4-7, claude-opus-4-8, including the $10.00/$50.00 claude-fable-5 row that had been the single highest rate on the list) no longer appear on deepinfra.com/pricing. Many models show a discounted cached-prompt input rate inline. Newly launched models can also carry an introductory percentage-off promotion — MiMo-V2.5-Pro debuted with a “61% off” tag cutting $1.00/$3.00 ($0.20 cached) per 1M list price to $0.39/$1.17 ($0.078 cached), shown as a struck-through list price beside the promotional rate on the /models and /deepstart pages; GLM-5.2 picked up a similar “35% off” tag on 2026-08-28, cutting its post-cut $0.75/$2.40 ($0.14 cached) rate to $0.488/$1.56 ($0.091 cached) — both promos live only on /models and /deepstart, never on the /pricing page itself. Embeddings run $0.005–$0.01 per 1M input tokens. A per-request Service Tier multiplier sits on top of these rates and now has three steps: Flex at 0.8× base price (“lower cost for non-production and asynchronous work, in exchange for slower responses and occasional unavailability”), the default Standard at 1× (best-effort during peak demand), and Priority at 1.5×, which schedules requests ahead of standard traffic for faster time-to-first-token (enabled per request via service_tier: "priority"; availability varies by model — the model directory carries Flex and Priority as filters and per-model badges). See the broader token-based billing pattern across the corpus.
  2. Per-output-unit media: Flux image generation is priced per image, scaled by resolution and iteration count (e.g. FLUX-2-pro is listed at $0.015/image on the pricing page; its model page details $0.03 for the first output megapixel and $0.015 per megapixel after that, so a multi-megapixel render costs more than the headline). Voxtral audio transcription is billed per minute of audio input ($0.00100–$0.00300/min). Text-to-speech introduces a separate unit again — Qwen3-TTS and Qwen3-TTS-VoiceDesign are quoted at $20.00 per 1M characters of input text.
  3. Per-GPU-hour containers: On-demand dedicated GPUs bill in minute granularity (invoiced weekly): A100 $0.89, H100 $2.20, H200 $2.69, B200 $3.69, B300 $4.89 per GPU-hour. The GPU-instances product lists 1×B200 $3.69/hr up to 8×B200 $29.52/hr with no egress fees.
  4. Reserved DeepCluster commitment: Customer-owned B300 clusters (256–5,000 GPUs) at $2.99/GPU-hr (3-year) or $1.98/GPU-hr (5-year), all-in — the only commitment-based tier and the only sales-led surface.

What makes this different: DeepInfra collapses five different metering units (tokens, GPU-hours, images, audio minutes, TTS characters) and a multi-year reserved cluster into one transparent, contract-free account — and on DeepCluster it inverts the cloud model so the customer owns the hardware while DeepInfra operates it.


Pricing by product

Serverless LLM inference (per 1M tokens — representative rows)

ModelContextPrice (in / out per 1M)Key mechanics
DeepSeek-V4-Pro1024k$1.30 / $2.60 ($0.10 cached)Flagship DeepSeek; per-token
DeepSeek-V4-Flash1024k$0.09 / $0.18 ($0.018 cached)Efficiency-focused MoE
DeepSeek-V3.2160k$0.26 / $0.38 ($0.13 cached)Same input as V3.1 at ~40% of its output rate
DeepSeek-V3.1160k$0.25 / $0.95 ($0.13 cached)Popular open reasoning model
Llama-3.3-70B-Instruct-Turbo128k$0.10 / $0.32Turbo throughput tier
Meta-Llama-3.1-8B-Instruct-Turbo128k$0.02 / $0.04Lowest-cost small model
Qwen3-235B-A22B-Instruct-2507256k$0.09 / $0.55Large MoE, low input rate
Kimi-K3 (Moonshot AI)1024k$2.85 / $14.25 ($0.285 cached)Highest per-token rate now shown on the /pricing page
GLM-5.2 (Z-AI)1024k$0.488 / $1.56 ($0.091 cached) — 35% offFeatured flagship on the /models catalog (not shown in the per-provider /pricing table); list price $0.75 / $2.40 ($0.14 cached, the post-2026-08-04-cut rate) struck through beside a new promotional rate tagged “35% off” as of 2026-08-28
DeepSeek-V4-Flash-07311024k$0.08 / $0.18 ($0.016 cached)Official release superseding the preview build; listed only on the /models and /deepstart catalogs (not the per-provider /pricing table)
MiMo-V2.5-Pro (XiaomiMiMo)1024k$0.39 / $1.17 ($0.078 cached) — 61% offLaunch promo shown only on /models and /deepstart (not the per-provider /pricing table): list price $1.00 / $3.00 ($0.20 cached) struck through beside the promotional rate, tagged “61% off”
Gemini 2.5 Pro976k$1.25 / $10.00Hosted Google model
claude-haiku-4-5195k$1.00 / $5.00Sole Claude model left on the /pricing page as of 2026-08-28 — six other Claude models (claude-opus-5, claude-fable-5, claude-sonnet-5, claude-sonnet-4-6, claude-opus-4-7, claude-opus-4-8) that were listed as recently as 2026-08-26 no longer appear

Service Tiers (per-request scheduling multiplier)

TierSchedulingPriceKey mechanics
Flex”Lower cost for non-production and asynchronous work, in exchange for slower responses and occasional unavailability”0.8× base priceAdded 2026-07-21; batch / dev / async workloads trade latency and availability for a 20% discount
StandardDefault scheduling, best-effort during peak demand1× base priceApplies to every per-token request unless overridden
PriorityScheduled ahead of standard traffic for faster time-to-first-token during peak demand1.5× base priceEnable per request via service_tier: "priority"; availability varies by model

Service-tier availability is per model — the model directory exposes Flex and Priority as filters, and each model card shows the tiers it supports.

Embeddings, audio, image & speech (other metering units)

ProductUnitPriceKey mechanics
Embeddings (bge / gte / e5 family)per 1M input tokens$0.005–$0.01Per-token, by model
Voxtral audio (speech-to-text)per minute of audio$0.00100–$0.00300Per-minute, Mini vs Small
Flux image generationper imagefrom $0.0005/image (FLUX-2-pro listed at $0.015)Scaled by resolution × iterations; the FLUX-2-pro model page bills the first output megapixel at $0.03 and each subsequent megapixel at $0.015
Qwen3-TTS / Qwen3-TTS-VoiceDesign (text-to-speech)per 1M characters$20.00Priced on characters of text in, not audio out

On-demand GPU containers (per GPU-hour)

GPUMemoryPriceKey mechanics
A10080GB$0.89 / GPU-hourCustom-LLM deploy; minute granularity, invoiced weekly
H10080GB$2.20 / GPU-hourSXM-connected multi-GPU
H200141GB$2.69 / GPU-hourAuto-scaling on load
B200180GB$3.69 / GPU-hourAlso on-demand: 8×B200 $29.52/hr, no egress fees
B300270GB$4.89 / GPU-hourTop single-card rate

DeepCluster — reserved B300 (sales-led)

ConfigurationPricePublic-cloud referenceKey mechanics
256–5,000 GPUs · 3-year term$2.99 / GPU-hr$6.50 / GPU-hr54% cheaper; customer owns hardware
256–5,000 GPUs · 5-year term$1.98 / GPU-hr$6.50 / GPU-hr70% cheaper; DeepInfra operates it

Sales motions across products: PLG / self-serve for per-token APIs, on-demand GPU instances, and DeepStart credits; sales-led for DeepCluster and enterprise (contact [email protected]).


Hidden costs : What a real DeepInfra inference bill actually adds up to

DeepInfra’s headline per-token rates look tiny, but production traffic and dedicated GPU uptime are where the bill is built. Two representative archetypes:

A mid-size app on DeepSeek-V3.1 inference

Line itemMonthly cost
800M input tokens @ $0.25 / 1M$200
250M output tokens @ $0.95 / 1M$238
Embeddings: 200M tokens @ $0.01 / 1M$2
Total$440

Per-token economics stay cheap at app scale — but note the account would cross into Tier 3 ($500 paid) over a few months, changing the invoicing cadence rather than the rate.

A team renting two B200 GPUs full-time for a custom model

Line itemMonthly cost
1×B200 @ $3.69/hr × 730 hrs$2,694
1×B200 @ $3.69/hr × 730 hrs$2,694
Total$5,388

Once a workload justifies always-on dedicated GPUs, the bill jumps two orders of magnitude versus per-token calls — the point at which DeepCluster’s $1.98–$2.99/GPU-hr reserved economics start to matter.

Want to estimate your own DeepInfra bill? Use the DeepInfra pricing calculator to model your monthly cost based on token volume, GPU-hours, and reserved-cluster terms.


Pricing evolution : From per-token inference to reserved customer-owned clusters

DeepInfra’s pricing has moved through four distinct phases: a 2023 execution-time model (pay per second of inference), a 2023–2024 shift to per-token language pricing, a 2025–2026 expansion into multi-unit metering plus reserved capacity, and — from mid-2026 — a turn toward packaging the rate instead of only cutting it. For most of that span the headline trend was relentless downward pressure on unit rates: a GPU-hour rate that fell roughly 2.5× in eighteen months and a small-model token cut with every model generation. That trend broke on 2026-07-14, when dedicated GPU rates rose 16–32% and DeepSeek-V3.1 rose about 20% — the first increases in the tracked span — and one week later, on 2026-07-21, a Flex service tier arrived at 0.8× base price. The reversal reached the per-token side of the card on 2026-07-29, when Meta-Llama-3.1-8B-Instruct-Turbo — the exact model this page and DeepInfra’s own catalog cite as the cheapest small-model reference — took a 33% output-rate increase while a neighboring Gemma model was cut in the same update, confirming the July reversal was not confined to GPU-hour tenants or a single flagship model. Less than a week later, on 2026-08-04, the direction flipped back down for one model — GLM-5.2 fell roughly 20% — but only on the /models catalog, not on the /pricing page most buyers actually check. Three-and-a-half weeks after that, on 2026-08-28, both surfaces moved again: the /pricing page itself shrank, dropping six of its seven listed Claude models and handing the page’s highest per-token rate to Kimi-K3, while GLM-5.2 took a further 35% promotional cut off-page. The cheapest way to buy DeepInfra is now a scheduling choice as much as a lower rate card, and tracking the full rate card now means watching two different pages — one of which can just as easily narrow as it can quietly discount.

Cadence

QuarterPrice changesProduct / SKU additionsNotes
2023 Q101Launch model: pure execution-time billing — $0.0005/second ($0.03/min) on A100, 1 hour free GPU, $0.04/GB-hr memory reservation.
2023 Q4112023-12 per-token LLM pricing introduced (Llama-2-70b $0.70/$0.90 in/out) alongside execution time; marketed “50% less than ChatGPT-3.5 Turbo.”
2024 Q2022024-04 Embeddings list ($0.005–$0.01/1M) and Custom-LLM GPU rental added (A100 $2.00, H100 $4.00 /GPU-hr); $1.80 signup credit live.
2024 Q3112024-09 Llama-3.1 cuts (8B $0.055, 70B $0.35/$0.40); automatic Usage Tiers ($20–$5,000) and DeepStart program enter the page.
2025 Q1212025-02 GPU cut (A100 $1.50, H100 $2.40, H200 added $3.00); LoRA pricing added; $1.80 signup credit removed; Tier 5 raised to $10,000.
2025 Q2012025-05 Execution-time pricing block retired (per-token + per-GPU-hour become the core meters); Llama 4 Scout & Maverick launch.
2025 Q3122025-08 aggressive GPU cut (A100 $0.89, H100 $1.69, H200 $1.99); per-provider page redesign; 2025-09 Voxtral per-minute audio added.
2025 Q4022025-12 inline cached-input rates appear; B200 self-serve GPU row ($2.49/GPU-hr); FLUX.2 image models launch.
2026 Q2122026-05 DeepCluster (customer-owned B300, $1.98–$2.99/GPU-hr) launches; B300 self-serve row added ($4.20); $107M Series B closes 2026-05-04. 2026-06 Priority Service Tier (1.5× base price) added; model catalog expands.
2026 Q3512026-07-14 first increase in the tracked span — custom-LLM H100 $1.79→$2.20, H200 $2.19→$2.69, B200 $2.79→$3.69, B300 $4.20→$4.89/GPU-hr and DeepSeek-V3.1 $0.21/$0.79→$0.25/$0.95; 2026-07-21 Flex service tier added at 0.8× base price, Mistral-Nemo-Instruct-2407 cut to $0.019/$0.03; 2026-07-29 Meta-Llama-3.1-8B-Instruct-Turbo output up 33% to $0.04, gemma-4-31B-it-turbo cut to $0.09/$0.34; 2026-08-04 GLM-5.2 (Z-AI flagship, /models-catalog-only listing) cut ~20% — $0.93/$3.00/$0.18 → $0.75/$2.40/$0.14 in/out/cached; 2026-08-28 the /pricing page’s Claude section narrows from seven models to one (claude-haiku-4-5 only — claude-fable-5 at $10/$50, formerly the page’s highest rate, and five other Claude models drop off, leaving Kimi-K3 $2.85/$14.25 as the new highest), and GLM-5.2 (still /models-only) takes a further 35%-off promo to $0.488/$1.56.

Tracked range: 2023-02–2026-07. Quarters not listed above were verified stable (0 price changes, 0 SKU additions).

Notable changes

  • 2023-02 — Earliest archived pricing page bills purely by inference execution time at $0.0005/second with 1 hour free GPU; no per-token rates exist yet (source: Wayback deepinfra.com/pricing 2023-02).
  • 2023-12 — Per-token LLM pricing introduced (Llama-2-70b $0.70 in / $0.90 out per 1M), framed as “50% less than ChatGPT-3.5 Turbo” and “55% less than Replicate” on execution time (source: Wayback 2023-12).
  • 2024-09 — Llama-3.1 cuts and the first automatic Usage Tiers; DeepStart startup-credit program appears in the nav (source: Wayback 2024-09).
  • 2025-05 — Execution-time pricing is retired entirely, simplifying the meter set to per-token + per-GPU-hour (source: Wayback 2025-05).
  • 2025-08 — Custom-LLM GPU rates cut about 40% (A100 $1.50→$0.89, H100 $2.40→$1.69, H200 $3.00→$1.99); pricing page redesigned per-provider (source: Wayback 2025-08).
  • 2026-05-04 — DeepCluster launches and DeepInfra closes a $107M Series B co-led by 500 Global and Georges Harik, with NVIDIA, Samsung Next, and Supermicro participating; the company reports ~5 trillion tokens/week and 25× token growth since Series A (source: deepinfra.com/series-b).
  • 2026-06-30 — A per-request Priority Service Tier appears on the pricing page (1.5× base price for faster time-to-first-token during peak demand, set via service_tier: "priority"), DeepInfra’s first explicit latency-vs-cost packaging knob; headline rates are otherwise unchanged (source: Wayback deepinfra.com/pricing 2026-06).
  • 2026-07-14 — DeepInfra raises prices for the first time in the tracked span: custom-LLM GPU-hour rates up 16–32% (H100 $1.79→$2.20, H200 $2.19→$2.69, B200 $2.79→$3.69, B300 $4.20→$4.89, with the A100 held flat at $0.89), on-demand B200 instances move in lockstep ($2.79→$3.69/hr for one card, $22.32→$29.52/hr for eight), and DeepSeek-V3.1 rises to $0.25 in / $0.95 out per 1M (source: deepinfra.com/pricing and /gpu-instances 2026-07-14).
  • 2026-07-21 — A third service tier, Flex at 0.8× base price (“lower cost for non-production and asynchronous work, in exchange for slower responses and occasional unavailability”), joins Standard (1×) and Priority (1.5×), and appears as a browsable filter plus a per-model badge in the model directory. The only rate move is Mistral-Nemo-Instruct-2407, cut from $0.02/$0.04 to $0.019/$0.03 per 1M (source: deepinfra.com/pricing 2026-07-21).
  • 2026-07-29 — DeepInfra raises Meta-Llama-3.1-8B-Instruct-Turbo — the model this page and DeepInfra’s own catalog cite as the cheapest small-model reference — from $0.02 in / $0.03 out to $0.02 in / $0.04 out per 1M tokens (+33% on output), while cutting gemma-4-31B-it-turbo from $0.12/$0.37 to $0.09/$0.34 per 1M in the same update; GPU-hour, DeepCluster, usage-tier, and service-tier rates are unchanged, and this is the second rate increase since the 2026-07-14 reversal (source: deepinfra.com/pricing 2026-07-29).
  • 2026-08-04 — GLM-5.2, Z-AI’s flagship long-horizon model, is cut roughly 20% across all three published rates ($0.93 in / $3.00 out / $0.18 cached → $0.75 / $2.40 / $0.14 per 1M), while every other Featured model checked in the same comparison (DeepSeek-V4-Pro, DeepSeek-V4-Flash, Kimi-K2.6/K2.7-Code, Nemotron-3-Ultra, Qwen3-Max, Qwen3.6-35B-A3B, GLM-5.1, MiMo-V2.5-Pro) holds its rate exactly; GLM-5.2’s price appears only on the /models Featured carousel, never on the /pricing page itself (source: deepinfra.com/models 2026-08-04).
  • 2026-08-28 — DeepInfra’s /pricing page Claude section narrows from seven models (as of 2026-08-26) to one: claude-haiku-4-5 ($1.00 / $5.00) is now the only Claude model listed, with claude-opus-5, claude-fable-5 ($10.00/$50.00, formerly the page’s single highest per-token rate), claude-sonnet-5, claude-sonnet-4-6, claude-opus-4-7, and claude-opus-4-8 no longer shown; Kimi-K3 ($2.85 in / $14.25 out per 1M) becomes the new highest rate on the page (source: deepinfra.com/pricing 2026-08-28).
  • 2026-08-28 — GLM-5.2 — already listed only on /models and /deepstart, never /pricing — picks up a new “35% off” promotional tag, cutting its post-2026-08-04-cut rate of $0.75 / $2.40 ($0.14 cached) to $0.488 / $1.56 ($0.091 cached) per 1M (source: deepinfra.com/models, /deepstart 2026-08-28).

The GPU-rate descent and the July 2026 reversal in detail

The custom-LLM GPU rate is the clearest single thread of DeepInfra’s price-cutting reputation. The A100 GPU-hour rate fell from $2.00 (2024-04) to $1.50 (2025-02) to $0.89 (2025-08) — roughly a 2.25× reduction in eighteen months — and the H100 fell even harder, from $4.00 to $1.69 over the same window. These cuts tracked falling wholesale GPU economics and intensifying competition with Together AI, Fireworks, and RunPod, and they reset the per-GPU-hour floor for the whole open-model inference market — a live case study in the token-cost deflation paradox where per-unit prices fall even as total inference spend climbs. The 2026 DeepCluster launch extends the same logic to multi-year buyers: rather than cut the on-demand rate further, DeepInfra offers customer-owned B300 capacity at $1.98/GPU-hr all-in, undercutting its own on-demand rate for anyone who can commit five years.

The reversal. On 2026-07-14 that descent stopped and inverted. Custom-LLM rates went H100 $1.79→$2.20, H200 $2.19→$2.69, B200 $2.79→$3.69 and B300 $4.20→$4.89 per GPU-hour — increases of 16% to 32% — while only the A100 held flat at $0.89, and DeepSeek-V3.1 rose about 20% to $0.25/$0.95 per 1M. The shape of the move is the tell: every Hopper-and-newer card went up and the oldest card did not, which prices scarcity of current-generation silicon rather than a general margin grab, and keeps intact the $0.89 A100 number that cost-sensitive developers quote as DeepInfra’s floor. For a buyer, the practical consequence is that this price list can no longer be modelled as a one-way ratchet down: a team that budgeted a year of B200 uptime on the June rate is now roughly a third over, and dedicated-GPU tenants — who commit to uptime rather than per-call volume — absorb the full increase with no packaging change to opt into.

Why Flex reads as the answer. One week later, on 2026-07-21, DeepInfra shipped Flex at 0.8× base price instead of walking the increase back. That sequencing turns a rate rise into a rate choice: price-sensitive traffic that can tolerate “slower responses and occasional unavailability” opts down per request, while latency-sensitive traffic can still opt up to Priority at 1.5×. It is also capacity management wearing a discount’s clothes — Flex work is explicitly deferrable, so it can be pushed into demand troughs while peak capacity is sold at Standard or Priority. Raising the base rate on scarce Blackwell-class hardware and simultaneously offering 20% off to anyone willing to wait is, in substance, congestion pricing. The one group it does not reach is exactly the group the increase hit hardest: service tiers apply to per-token inference, not to the per-GPU-hour containers whose rates moved most.

A second data point. On 2026-07-29, the reversal reached the per-token side of the price list for the first time since 2026-07-14: Meta-Llama-3.1-8B-Instruct-Turbo — the exact model this page and DeepInfra’s own catalog repeatedly cite as the cheapest small-model reference — rose 33% on output, from $0.03 to $0.04 per 1M tokens, while gemma-4-31B-it-turbo was cut from $0.12/$0.37 to $0.09/$0.34 in the same update. The pairing matters: DeepInfra did not raise every small model’s rate, it raised the specific number developers actually quote, while cutting a neighboring Gemma model enough to keep the page’s overall story mixed rather than uniformly upward — the same asymmetric-move shape as the 2026-07-14 GPU increase, where only the A100 held flat. Two increases in fifteen days is no longer a single data point; it is the start of a pattern, and the first sign that the reversal is not confined to dedicated-GPU tenants.

A cut, on the page nobody watches (2026-08-04). Six days later the pattern reversed again — this time back down, and on a flagship rather than a small model. GLM-5.2, Z-AI’s newest long-horizon model and a Featured entry in DeepInfra’s own catalog, was cut roughly 20% across every published number (input, output, and cached rate alike), while every other Featured model checked alongside it — DeepSeek-V4-Pro, DeepSeek-V4-Flash, the Kimi-K2.6/K2.7-Code pair, Nemotron-3-Ultra, Qwen3-Max, Qwen3.6-35B-A3B, GLM-5.1, and MiMo-V2.5-Pro — held its rate exactly, isolating the move to one SKU. Two things follow. First, this is a genuine counter-example to reading July as “prices only go up now”: three weeks into the reversal, DeepInfra still cut a flagship rate by a fifth, so the 07-14 and 07-29 increases look like selective repricing of specific high-traffic SKUs rather than a blanket policy change. Second, GLM-5.2’s price has never appeared on the /pricing page at all — it lives only in the /models Featured carousel — so this is the first tracked price move that a buyer monitoring only the headline pricing page would miss entirely, both the old number and the cut. The per-model transparency the page is known for now spans two surfaces, and only one of them is called “pricing.”

The Claude catalog narrows, and GLM-5.2 gets a second cut (2026-08-28). Three and a half weeks later the /pricing page itself moved, not just the /models catalog. The Claude section — seven models as recently as 2026-08-26 — dropped to one: claude-haiku-4-5 at $1.00/$5.00. Gone are claude-opus-5, claude-sonnet-5, claude-sonnet-4-6, claude-opus-4-7, claude-opus-4-8, and claude-fable-5, which at $10.00/$50.00 per 1M had been the single most expensive row on the entire page. With claude-fable-5 gone, Kimi-K3 ($2.85/$14.25) becomes the new most expensive listed rate — a materially cheaper ceiling for anyone scanning the page for a worst-case ballpark. No changelog entry accompanies the removal, so a developer who had claude-fable-5 or claude-opus-5 wired into a cost model has no way to tell, from the page alone, whether the model was deprecated, repriced elsewhere, or simply de-listed from this table. The same day, GLM-5.2 — still absent from /pricing — took a second cut in three and a half weeks, this time framed as a 35%-off promotion rather than a rate change: $0.75/$2.40 down to $0.488/$1.56 per 1M. Read together, the two moves cut in opposite directions on the two surfaces DeepInfra now needs to be read across: the page most buyers check got simpler and cheaper at the top end by subtraction, while the page fewer buyers check got a deeper discount than its 2026-08-04 cut already implied.


What’s unique : Five metering units, a customer-owned cluster, and a two-sided price ladder

1. Five metering units under one account. DeepInfra meters tokens (LLMs/embeddings), images (Flux, scaled by resolution × iterations), audio minutes (Voxtral speech-to-text), characters of input text (Qwen3-TTS at $20.00 per 1M), and GPU-hours (on-demand containers) on the same account — a breadth of billing primitives few inference clouds expose in one transparent price list. The text-to-speech unit is the quietest of the five and the easiest to mis-model: it bills on characters of text going in, not seconds of audio coming out.

2. Customer-owned reserved hardware. DeepCluster inverts the cloud rental model: the customer owns the NVIDIA B300 hardware (balance-sheet asset, depreciation-eligible) while DeepInfra procures, deploys, and operates it — pricing it all-in per GPU-hour rather than as a lease.

3. Inline cached-prompt rates. Many per-token rows show a discounted cached-input price next to the standard rate, surfacing prompt-cache economics directly in the public price list rather than burying them in docs.

4. A price list that moves in public — now in both directions. DeepInfra’s pricing page is a moving target by design: it cut GPU-hour rates roughly 2.5× between 2024 and 2025, dropped small-model token rates with every model generation, and retired its original per-second execution-time meter entirely along the way. What makes 2026 interesting is that the same mechanism ran in reverse on 2026-07-14 — GPU rates up 16–32%, DeepSeek-V3.1 up ~20% — published on the same public page, with the same detail, as the cuts had been. The transparency survived the increase; the implicit promise that the number only ever falls did not. On 2026-07-29 the reversal reached the token side too: Llama-3.1-8B-Instruct-Turbo’s output rate — the number the whole page uses as its cheapest-model reference — rose 33%, even as gemma-4-31B-it-turbo was cut in the same update, so the direction of travel is now genuinely mixed rather than a one-time GPU correction. Then on 2026-08-04 it swung back down: GLM-5.2, a Featured flagship model, was cut roughly 20% across input, output, and cached rates — the first outright cut since the reversal began — but that price move happened entirely on the /models catalog, since GLM-5.2 has never been listed on the /pricing page itself. On 2026-08-28 the “moving in public” idea extended to the catalog itself: six of seven Claude models quietly dropped off /pricing (leaving only claude-haiku-4-5 and handing the page’s highest-rate title to Kimi-K3), while GLM-5.2 took a second off-page cut via a new 35%-off promo — a page that can now change shape, not just price.

5. A two-sided per-request price ladder. Since 2026-07-21 the per-request Service Tier has three steps rather than two: Flex at 0.8× base price for non-production and asynchronous work (“slower responses and occasional unavailability”), the default Standard at 1×, and Priority at 1.5× for faster time-to-first-token during peak demand. That is a 1.9× spread between the cheapest and dearest way to run the identical model on identical hardware, selected per request by an API flag — no plan, no seat, no contract, and no negotiation. Most vendors express this as tiers you subscribe to; DeepInfra expresses it as a dial each individual call turns, which means a single application can run its interactive path at Priority and its nightly batch at Flex on the same key.


Strengths & weaknesses

StrengthsWeaknesses
Fully transparent per-model price list, published publiclyNo always-free tier — card or pre-pay required to start
Five metering units (tokens, images, audio minutes, TTS characters, GPU-hours) on one billPer-token rates vary model-by-model, so forecasting requires per-model math
Reserved DeepCluster from $1.98/GPU-hr, now a wider saving since on-demand B200 rose to $3.69DeepCluster is sales-led with multi-year terms and no self-serve path
No contracts or upfront costs for the usage products200 concurrent-request cap by default may throttle high-traffic apps
Flex (0.8×) and Priority (1.5×) let each request pick its own cost/latency point without a plan or contractFour numbers per model to forecast — Flex, base, cached-input, Priority — on top of an already per-model rate card
Rate changes land on the public page rather than inside a quoteThe 2026-07-14 and 2026-07-29 increases (GPU +16–32%, DeepSeek-V3.1 +~20%, Llama-3.1-8B-Instruct-Turbo output +33%) show the card can move up on both GPU-hour and per-token rates, and the tenants hit have no advance notice or cheaper tier to opt into
Every model’s rate is published somewhere on the site, not gated behind a quoteNot every model’s rate lives on the /pricing page itself — GLM-5.2’s ~20% cut on 2026-08-04 (and its further 35%-off promo on 2026-08-28) happened entirely on the /models catalog, so a buyer who only bookmarks the headline pricing page can miss both a model’s current rate and any change to it
Discontinued or de-featured models simply disappear from the /pricing page rather than lingering at a stale rateNo changelog marks the disappearance — six of seven Claude models, including claude-fable-5 (formerly the page’s highest rate), vanished from /pricing on 2026-08-28 with no note on whether they were deprecated or just de-listed

Billing UX : Usage tiers, spending limits, and threshold-based invoicing

  • Card-on-file or pre-pay requirement — you must add a card or pre-pay before you can use any service; there is no always-free entry path.
  • Automatic usage tiers — every account sits in a usage tier (Tier 1 $20, Tier 2 $100, Tier 3 $500, Tier 4 $2,000, Tier 5 $10,000), and DeepInfra moves accounts up automatically as cumulative spend grows.
  • Threshold-based invoicing — an invoice generates at the start of each month and again whenever the account hits its tier’s invoicing threshold, so heavier accounts are billed more frequently.
  • Spending limit — accounts can set a spending limit “to avoid surprises,” capping run-away usage cost.
  • Concurrency cap — each account is limited to 200 concurrent requests by default (raisable on request), a built-in guardrail against unbounded fan-out.
  • GPU billing granularity — dedicated GPU containers are billed in minute granularity and invoiced weekly, distinct from the monthly per-token invoicing cycle.
  • Per-request Service Tier control — each inference request can be sent at Flex (0.8× base price, for non-production and asynchronous work, with slower responses and occasional unavailability), the default Standard tier (1× base price), or Priority (1.5× base price for faster time-to-first-token during peak demand) by setting service_tier: "priority", letting callers trade cost against latency at the request level.
  • Per-model tier availability filters — the model directory exposes Flex and Priority as browsable filters and stamps each model card with the service tiers it supports, so a caller can confirm a discount or speed-up applies before wiring it into code.

Strategic wins : Why DeepInfra’s transparent multi-unit pricing works

1. Radical price transparency as a developer-acquisition wedge

By publishing a per-1M-token rate for nearly every open model — including discounted cached-input rates inline — DeepInfra lets a developer estimate cost before signing up, lowering the friction that often gates usage-based pricing adoption. Transparency is itself the marketing, and it scales across hundreds of token-billed models without a sales call.

2. Compounding price cuts as a moat — and reputation banked for the day rates rise

DeepInfra has cut unit rates repeatedly and publicly — the A100 GPU-hour fell from $2.00 (2024-04) to $0.89 (2025-08), and small-model token rates fell with each generation. Aggressive, legible price-cutting earns word-of-mouth in cost-sensitive communities like r/LocalLLaMA and makes the company the reflexive “cheap inference” reference, a position that compounds as the trillion-token economy drives ever-larger volumes through the cheapest provider. The 2026-07-14 increase is the first real test of that equity, and the mechanics of the move suggest it was designed with the reputation in mind: every Hopper-and-newer card went up while the A100 stayed at $0.89, preserving the single cheapest number on the page — the one developers actually quote — even as current-generation capacity repriced. A vendor with no history of cuts raising GPU rates a third invites churn; one that has cut ~2.5× in eighteen months is spending goodwill it genuinely banked. A second increase followed on 2026-07-29, this time on the small-model side of the card: Llama-3.1-8B-Instruct-Turbo’s output rate rose 33% to $0.04, while the $0.02 input rate and the $0.89 A100 GPU-hour floor — the two numbers most often quoted as DeepInfra’s cheapest — both held, so the pattern of protecting the specific figure that anchors the company’s reputation while raising the rate around it continued for a second event running.

3. One account spanning five metering units

Offering tokens, images, audio minutes, TTS characters, and GPU-hours on a single bill captures a customer’s full inference footprint rather than just the LLM slice, increasing account stickiness as workloads diversify. This mirrors the multi-meter direction seen at peers like Replicate but with a more transparent rate card.

4. Customer-owned reserved capacity

DeepCluster’s “you own the hardware, we operate it” framing converts a pure-opex cloud spend into a balance-sheet asset for large buyers — a differentiated commitment-based pitch versus standard reserved-instance leases, and a way to win the largest accounts without discounting the self-serve rate card.

5. Selling speed and patience — a two-sided ladder around the base rate

The 2026-06-30 Priority Service Tier (1.5× base for faster time-to-first-token) was the first time DeepInfra monetized quality of service rather than raw compute; the 2026-07-21 Flex tier (0.8× base for work that tolerates delay) completed the pair. Together they let the company price the same model three ways without touching the headline rate that wins cost-sensitive developers: latency-sensitive workloads such as agents and real-time UX self-select up, batch and dev traffic self-selects down, and everyone else pays the advertised number. That is textbook self-selecting price discrimination with none of the usual overhead — no plans, no seats, no negotiation, one API flag. The timing matters too: shipping Flex a week after a 16–32% GPU increase gives price-sensitive customers something to do about the rise other than leave. And because Flex work is explicitly deferrable, the discount doubles as demand-shaping — cheap traffic fills troughs, scarce peak capacity is sold at Standard or Priority, which is a better margin lever than cutting the floor rate again.


Areas to improve : Closing the free-tier and forecasting gaps

1. No self-serve free tier raises the trial barrier

Requiring a card or pre-pay before any usage is a higher bar than peers that offer trial credits. A small always-free monthly token allowance (separate from the gated DeepStart program) would lower first-call friction.

2. Per-model rate sprawl, now multiplied by service tiers

With hundreds of models each carrying its own input, output, and cached-input rate, finance teams already struggled to forecast spend — and the 2026 service tiers multiply the problem rather than add to it, since Flex (0.8×), Standard (1×), and Priority (1.5×) each apply against every one of those numbers. A team that mixes tiers by workload is effectively forecasting six price points per model, and availability varies model by model. The 2026-08-04 GLM-5.2 cut adds a further wrinkle: some model rates, including a Featured flagship, live only on the /models catalog and never appear on the /pricing page’s per-provider tables, so a team monitoring the headline page alone can miss both the number and any change to it. A first-party cost estimator or budget-projection tool tied to the usage tiers — one that takes a traffic mix and a tier split and returns a range, pulled from a single canonical rate list rather than two separate pages — would reduce bill-shock risk far more now than it would have a quarter ago.

3. DeepCluster has no self-serve on-ramp

The reserved-cluster product is entirely sales-led with multi-year terms, so a team that knows it wants a 256-GPU cluster still has to email [email protected]. A published configurator that quotes indicative GPU-hour pricing for a given GPU count and term — even gated behind a short form — would shorten the path from interest to contract and let buyers self-qualify before sales engages.

4. Price increases deserve the same publishing ritual as the cuts

The 2026-07-14 GPU-hour increase (H100, H200, B200 and B300 up 16–32%) landed the way every previous change has: the number on the page simply became a different number, with no dated rate-change log, no effective date, and no advance notice. That asymmetry is fine while rates only fall — nobody complains about a surprise discount — but it is a real problem for the customers most exposed to a rise, since dedicated-GPU tenants pay for uptime rather than per-call volume and cannot mitigate by shifting traffic to a cheaper service tier. The concrete fix is a public pricing changelog with effective dates plus a notice period for increases on the per-GPU-hour products, so a team budgeting a year of B200 uptime learns about a 32% move before the invoice does. The pattern repeated on 2026-07-29, when Llama-3.1-8B-Instruct-Turbo’s output rate rose 33% with the same silence — no changelog entry, no notice — showing the gap isn’t limited to GPU-hour tenants: any developer who hard-coded a per-token rate into a cost model is exposed to the identical blind spot. Given that DeepInfra’s whole reputation rests on price transparency, publishing rises as deliberately as it publishes cuts costs nothing and protects the thing that makes the transparent rate card persuasive in the first place — and the fix should now cover per-token rate moves on general-availability models, not just per-GPU-hour products. The same silence now applies to removals: on 2026-08-28, six of seven Claude models — including claude-fable-5, the page’s former highest rate — dropped off /pricing with no changelog note explaining whether they were deprecated or simply de-listed. A team that hard-coded claude-fable-5 or claude-opus-5 into a cost model has no way to tell from the page alone what happened; the same effective-dated changelog that would fix silent rises would fix this too.


Monetization stack & signals : how DeepInfra builds & buys its revenue engine

Buys 1 Builds 2

The read — where the monetization investment is going

DeepInfra builds the meter behind its own usage pricing — its first-party Usage API records per-token/per-second units, rates and cost itself — and buys only the invoicing edge (Stripe Invoice IDs, a hosted billing portal). The margin under its relentless public price cuts is defended by in-house inference-cost engineering (Blackwell + NVFP4 → 5¢/M tokens), not by a bought FinOps tool.

Stack — build vs buy
Builds in-house · 2
  • In-house usage meter In-house build Docs Jun 2026

    “DeepInfra's first-party Usage API returns per-model UsageItem records keyed on `units` ("billed seconds or tokens"), `rate` ("rate in cents/sec or cents per token"), `cost` ("model cost in cents") and `pricing_type` — the company meters its own per-token / per-second consumption rather than buying a Metronome/Orb metering layer.”

  • In-house inference cost engine Cost & FinOps Press Feb 2026

    “NVIDIA: "DeepInfra reduced the cost per million tokens from 20 cents on the NVIDIA Hopper platform to 10 cents on Blackwell … further cut that cost to just 5 cents — for a total 4x improvement in cost per token" (on a large MoE model) — cost-per-token engineering DeepInfra does in-house to defend the margin on its publicly-cut rate card.”

Buys (vendor) · 1
  • Stripe Billing Docs 1 Docs 2 Jun 2026

    “The Usage API's UsageMonth object carries an `invoice_id` field described verbatim as "Stripe Invoice ID, or EMPTY|NOT_FINAL" — DeepInfra's own meter rolls usage into Stripe invoices.”

Signals reviewed · derived from press & filings, product docs

Key takeaways

  1. Publish the full price list. DeepInfra’s per-model transparency turns the pricing page itself into a developer-acquisition asset — buyers can model cost before they sign up, the opposite of a gated quote-only motion, and a clean example of the usage-based pricing models playbook.
  2. Make price moves a public ritual — then price the schedule, not just the unit. DeepInfra cut GPU-hour rates ~2.5× in eighteen months and let the market see every step, and when it raised them 16–32% on 2026-07-14 it published that just as plainly; visible, repeated moves buy mindshare a quiet discount never would, but they also mean an increase is unmissable. The hedge is packaging: with Flex at 0.8× and Priority at 1.5× (2026-06-30 and 2026-07-21), the same base rate now serves patient batch traffic and latency-sensitive traffic at different prices, so willingness-to-pay is captured on the scheduling axis instead of by moving the headline number again. A second per-token increase on 2026-07-29 (Llama-3.1-8B-Instruct-Turbo output +33%) confirms the reversal wasn’t a one-off, so teams pricing off this vendor should now model rate risk on the token side too, not just dedicated-GPU uptime. By 2026-08-28 the ritual had a third form: catalog narrowing. Six Claude models dropped off /pricing without a changelog entry, so “watch the page” now means watching for models disappearing, not just numbers moving.
  3. One account, many meters. Spanning tokens, images, audio minutes, TTS characters, and GPU-hours on a single bill captures the whole inference footprint instead of just the LLM slice.
  4. Surface cache economics inline. Showing discounted cached-input rates next to standard rates makes prompt-cache savings legible without docs spelunking — a transparency edge over peers that bury caching in API docs.
  5. Invert the reserved model when you can. DeepCluster’s customer-owned hardware framing reframes a multi-year commitment as a balance-sheet asset, not just a discount, and protects the self-serve rate card from being undercut by the enterprise deal.

UBP implications

  1. Multi-unit metering is table stakes — and scheduling is becoming a second axis of the same meter. Charging per token, per image, per audio minute, per character, and per GPU-hour on one account shows usage-based pricing fragmenting into product-specific value metrics. DeepInfra’s 2026 service tiers extend that fragmentation sideways rather than adding a sixth unit: the same token, on the same model, costs 0.8× (Flex), 1× (Standard) or 1.5× (Priority) depending on when the buyer will accept it being served. As per-unit costs deflate, expect more usage-based vendors to price when alongside how much — congestion pricing arriving in software the way it did in electricity and freight.
  2. Transparency lowers UBP adoption friction — and the first increase is where it gets tested. A fully public per-model rate card counters the “unpredictable bill” objection that slows usage-based pricing, and visibility is a feature. But it cuts both ways: DeepInfra’s 2026-07-14 rise was as visible as every cut before it, which is the trade a transparent vendor accepts — publish everything and be ready to defend a rise as openly as you advertised the reductions. A second rise on 2026-07-29, this time on a per-token small-model rate rather than GPU-hours, suggests the test isn’t a one-time event but an ongoing discipline the vendor will need to sustain.
  3. Reserved commitments still anchor the top of a usage funnel. Even a pure-usage vendor needs a commitment tier (DeepCluster) to serve the largest, most cost-sensitive buyers.

Sources


Bottom line

DeepInfra is one of the most transparent open-model inference clouds in the market: pure-usage per-token APIs that publish a rate for nearly every model, four more metering units (images, audio minutes, TTS characters, GPU-hours) on the same account, and a customer-owned DeepCluster reserved tier from $1.98/GPU-hr for buyers who outgrow on-demand. July 2026 changed the story in three ways worth pricing into a decision: on 2026-07-14 dedicated GPU rates rose 16–32% — the first increase in the tracked span, so this rate card can no longer be assumed to only fall — on 2026-07-21 a Flex tier at 0.8× base price completed a three-point service ladder (Flex 0.8× / Standard 1× / Priority 1.5×) that lets each request choose its own cost-versus-latency point, and on 2026-07-29 the reversal reached the per-token side too, with Meta-Llama-3.1-8B-Instruct-Turbo’s output rate — the model DeepInfra’s own catalog cites as its cheapest — up 33% while a neighboring Gemma model was cut in the same update. A fourth beat followed on 2026-08-04: GLM-5.2, a Featured flagship model, was cut roughly 20% — the first outright cut since the reversal began — but the move happened entirely on the /models catalog, never touching the /pricing page, so a complete read of DeepInfra’s rate card now means checking two pages, not one. A fifth beat landed on 2026-08-28: the /pricing page’s Claude section itself narrowed from seven models to one (claude-haiku-4-5 only), dropping claude-fable-5 — the page’s former highest per-token rate at $10.00/$50.00 — with no changelog note, while Kimi-K3 ($2.85/$14.25) became the new highest rate shown, and GLM-5.2 took a further 35%-off promotional cut off-page. The absence of a self-serve free tier is still the main on-ramp gap, and dedicated-GPU tenants, small-model callers, and anyone who had a now-delisted Claude model wired into a cost model alike now have no advance-notice mechanism to plan around a change, but for teams that already know they’ll pay for inference, the price clarity is hard to beat.

Want to compare DeepInfra against other inference-cloud pricing? Browse the pricing blueprint.

Pricing timeline : Major events on a vertical axis

Each milestone below corresponds to a public pricing change, product launch, or material adjustment. Major events use a filled marker; minor adjustments use a faded one.

Six of seven Claude models pulled from the /pricing page; GLM-5.2 gets a new 35% promo

DeepInfra's /pricing page Claude section drops from seven models (as of 2026-08-26) to one: claude-haiku-4-5 ($1.00/$5.00) is now the only Claude model listed, with claude-opus-5, claude-fable-5 ($10.00/$50.00, the former highest per-token rate on the page), claude-sonnet-5, claude-sonnet-4-6, claude-opus-4-7, and claude-opus-4-8 no longer shown. Kimi-K3 ($2.85 in / $14.25 out per 1M) is now the highest per-token rate on the page. Separately, GLM-5.2 — already only listed on the /models and /deepstart catalogs, not /pricing — picks up a new "35% off" promotional tag cutting its post-2026-08-04-cut rate of $0.75/$2.40 ($0.14 cached) to $0.488/$1.56 ($0.091 cached) per 1M (source: deepinfra.com/pricing, /models, /deepstart 2026-08-28).

Six of seven Claude models pulled from the /pricing page; GLM-5.2 gets a new 35% promo - DeepInfra's /pricing page Claude section drops from seven models (as of 2026-08-
captured

GLM-5.2 cut ~20% — but only on the /models catalog

GLM-5.2 (Z-AI's flagship long-horizon agentic model, 1M-token context) is cut roughly 20% across all three published rates — $0.93 in / $3.00 out / $0.18 cached to $0.75 / $2.40 / $0.14 per 1M tokens — while every other Featured model checked in the same comparison (DeepSeek-V4-Pro, DeepSeek-V4-Flash, Kimi-K2.6/K2.7-Code, Nemotron-3-Ultra, Qwen3-Max, Qwen3.6-35B-A3B, GLM-5.1, MiMo-V2.5-Pro) holds its rate exactly. GLM-5.2 is not shown on the /pricing page at all — it surfaces only in the /models Featured carousel — making this the first tracked price move that lives entirely outside DeepInfra's main pricing surface (source: deepinfra.com/models 2026-08-04).

GLM-5.2 cut ~20% — but only on the /models catalog - GLM-5.2 (Z-AI's flagship long-horizon agentic model, 1M-token context) is cut ro
captured

Second rate rise: Llama-3.1-8B-Instruct-Turbo output up 33%

DeepInfra raises Meta-Llama-3.1-8B-Instruct-Turbo — the model this page and DeepInfra's own catalog cite as the cheapest small-model reference — from $0.02 in / $0.03 out to $0.02 in / $0.04 out per 1M tokens (+33% on output), while cutting gemma-4-31B-it-turbo from $0.12/$0.37 to $0.09/$0.34 per 1M in the same update. GPU-hour, DeepCluster, usage-tier, and service-tier rates are unchanged; this is the second rate increase since the 2026-07-14 reversal, suggesting it was not a one-off (source: deepinfra.com/pricing 2026-07-29).

Second rate rise: Llama-3.1-8B-Instruct-Turbo output up 33% - DeepInfra raises Meta-Llama-3.1-8B-Instruct-Turbo — the model this page and Deep
captured

Flex service tier added at 0.8× base price

DeepInfra adds a third per-request Service Tier: Flex, priced at 0.8× base price and described as "lower cost for non-production and asynchronous work, in exchange for slower responses and occasional unavailability" — joining Standard (1×) and Priority (1.5×) and turning the service-tier control into a three-point cost/latency ladder. Flex also appears as a new capability filter in the model directory and as a per-model badge on model cards. Headline rates are otherwise unchanged (DeepSeek-V3.1 $0.25/$0.95, A100 $0.89 / H100 $2.20 / H200 $2.69 / B200 $3.69 / B300 $4.89 per GPU-hour, 8×B200 $29.52/hr, DeepCluster $1.98–$2.99/GPU-hr, usage tiers $20–$10,000); the only rate move on the model list is Mistral-Nemo-Instruct-2407, cut from $0.02 in / $0.04 out to $0.019 in / $0.03 out per 1M (source: deepinfra.com/pricing 2026-07-21).

Flex service tier added at 0.8× base price - DeepInfra adds a third per-request Service Tier: Flex, priced at 0.8× base price
captured

GPU-hour rates raised; DeepSeek-V3.1 token price up

DeepInfra raises its dedicated GPU-hour rates for the first time in the tracked span: custom-LLM H100 $1.79→$2.20, H200 $2.19→$2.69, B200 $2.79→$3.69, B300 $4.20→$4.89 per GPU-hour (A100 holds at $0.89); on-demand B200 instances move in lockstep (1×B200 $2.79→$3.69/hr, 8×B200 $22.32→$29.52/hr). DeepSeek-V3.1 per-token rises to $0.25 in / $0.95 out (from $0.21 / $0.79) and Llama-3.1-8B-Turbo output drops to $0.03. The model catalog expands (DeepSeek-V4-Flash $0.09/$0.18, gemini-3.x, gemma-4, claude-fable-5 $10/$50, claude-sonnet-5). DeepCluster ($1.98–$2.99/GPU-hr), usage tiers ($20–$10,000), and the Priority service tier (1.5×) are unchanged (source: deepinfra.com/pricing & /gpu-instances 2026-07-14).

GPU-hour rates raised; DeepSeek-V3.1 token price up - DeepInfra raises its dedicated GPU-hour rates for the first time in the tracked
captured

Priority Service Tier added (1.5× per-token multiplier)

DeepInfra adds a per-request Service Tier control to the pricing page: the default Standard tier bills at 1× base price, while a new Priority tier schedules requests ahead of standard traffic for faster time-to-first-token at 1.5× base price (set via service_tier: "priority"). Headline per-token, per-GPU-hour, and DeepCluster rates are unchanged; the model catalog expands (DeepSeek-V4-Flash, Gemini 3.x, Gemma 4, Nemotron-3, Claude Haiku/Sonnet/Opus 4.x), and Qwen3-235B-A22B-Instruct-2507 input edges to $0.09 (source: deepinfra.com/pricing 2026-06-30).

Priority Service Tier added (1.5× per-token multiplier) - DeepInfra adds a per-request Service Tier control to the pricing page: the defau
captured

Per-token + per-GPU-hour + DeepCluster reserved capacity

Current pricing: per-1M-token LLM rates (DeepSeek-V3.1 $0.21/$0.79, Llama-3.1-8B $0.02/$0.05, Claude Opus $5.00/$25.00), embeddings $0.005–$0.01/1M, Voxtral audio $0.00100–$0.00300/min, Flux per-image, on-demand B200 GPU $2.79/hr, custom-LLM A100 $0.89/H100 $1.79/H200 $2.19/B200 $2.79/B300 $4.20 per GPU-hour, and DeepCluster reserved B300 from $1.98/GPU-hr (5-yr). Usage tiers $20–$10,000 invoicing thresholds.

Per-token + per-GPU-hour + DeepCluster reserved capacity - Current pricing: per-1M-token LLM rates (DeepSeek-V3.1 $0.21/$0.79, Llama-3.1-8B
captured

DeepCluster (customer-owned B300) and $107M Series B

DeepCluster launches: a customer-owned NVIDIA B300 cluster (256–5,000 GPUs, 99.982% uptime SLA) that DeepInfra procures and operates, all-in at $2.99/GPU-hr (3-yr, 54% cheaper than a $6.50 cloud reference) or $1.98/GPU-hr (5-yr, 70% cheaper). DeepInfra closes a $107M Series B on 2026-05-04 (source: Wayback deepcluster 2026-05; deepinfra.com/series-b).

DeepCluster (customer-owned B300) and $107M Series B - DeepCluster launches: a customer-owned NVIDIA B300 cluster (256–5,000 GPUs, 99.9
captured

Inline cached-input rates and self-serve B200

Per-token rows begin showing discounted cached-input prices inline (e.g. DeepSeek-V3.1 $0.21 / $0.168 cached). B200 appears as a self-serve custom-LLM GPU row at $2.49/GPU-hour; FLUX.2 image models launch (source: Wayback 2025-12).

Inline cached-input rates and self-serve B200 - Per-token rows begin showing discounted cached-input prices inline (e.g. DeepSee
captured

Voxtral per-minute audio transcription added

Voxtral speech-to-text models are added as a fourth metering unit, billed per minute of audio input ($0.00100/min Mini, $0.00300/min Small) — joining per-token, per-image, and per-GPU-hour on one bill (source: Wayback 2025-09).

Voxtral per-minute audio transcription added - Voxtral speech-to-text models are added as a fourth metering unit, billed per mi
captured

Aggressive GPU price cut and per-provider page redesign

Custom-LLM GPU rates cut hard: A100 $1.50→$0.89, H100 $2.40→$1.69, H200 $3.00→$1.99 per GPU-hour. The pricing page is redesigned into per-provider model sections (DeepSeek, Qwen, Llama 4, Gemma, Phi); "Contact Sales" enters the nav; B200 clusters referenced for dedicated buyers (source: Wayback 2025-08).

Aggressive GPU price cut and per-provider page redesign - Custom-LLM GPU rates cut hard: A100 $1.50→$0.89, H100 $2.40→$1.69, H200 $3.00→$1
captured

Execution-time pricing retired; Llama 4 launched

The per-second "Execution Time Pricing" block disappears from the pricing page, leaving per-token (LLM/embeddings) and per-GPU-hour as the core meters. Llama 4 Scout & Maverick go live; GPU rates hold at A100 $1.50 / H100 $2.40 / H200 $3.00 (source: Wayback 2025-05).

Execution-time pricing retired; Llama 4 launched - The per-second "Execution Time Pricing" block disappears from the pricing page,
captured

GPU price cut, H200 added, LoRA pricing, signup credit removed

Custom-LLM GPU rates cut: A100 $2.00→$1.50, H100 $4.00→$2.40, H200 added at $3.00/GPU-hr. LoRA-tuned model pricing appears. Llama-3.1-8B cut to $0.03/$0.05; Tier 5 threshold raised to $10,000; the $1.80 signup credit is removed (card or pre-pay now required) (source: Wayback 2025-02).

GPU price cut, H200 added, LoRA pricing, signup credit removed - Custom-LLM GPU rates cut: A100 $2.00→$1.50, H100 $4.00→$2.40, H200 added at $3.0
captured

Llama-3.1 price cuts and automatic Usage Tiers

Llama-3.1 rates fall sharply vs Llama-2: 8B to $0.055/$0.055, 70B to $0.35/$0.40, 405B at $1.79 in. Automatic five-step Usage Tiers appear (Tier 1 $20 → Tier 5 $5,000 threshold) and the DeepStart startup program enters the nav (source: Wayback 2024-09).

Llama-3.1 price cuts and automatic Usage Tiers - Llama-3.1 rates fall sharply vs Llama-2: 8B to $0.055/$0.055, 70B to $0.35/$0.40
captured

Embeddings, custom-LLM GPU rental, and $1.80 signup credit

Page adds an Embeddings price list ($0.005–$0.01 per 1M tokens) and Custom-LLM dedicated-GPU rental (A100 $2.00, H100 $4.00 per GPU-hour, billed by the minute, invoiced weekly). Billing copy notes "$1.80 when you sign up" as starter credit (source: Wayback 2024-04).

Embeddings, custom-LLM GPU rental, and $1.80 signup credit - Page adds an Embeddings price list ($0.005–$0.01 per 1M tokens) and Custom-LLM d
captured

Per-token pricing introduced alongside execution time

DeepInfra adds per-token LLM pricing — Llama-2-70b-chat $0.70 in / $0.90 out per 1M, Mistral-7B $0.13/$0.13 — marketed as "50% less than ChatGPT-3.5 Turbo," while execution-time billing ($0.0005/sec, "55% less than Replicate") remains for image/audio models (source: Wayback 2023-12).

Per-token pricing introduced alongside execution time - DeepInfra adds per-token LLM pricing — Llama-2-70b-chat $0.70 in / $0.90 out per
captured

Launch: per-second execution-time billing only

Earliest archived pricing page bills purely by inference execution time — $0.0005/second ($0.03/minute), billed per millisecond on A100 GPUs, with 1 hour of GPU free and reservable GPU memory at $0.04 per GB/hour. No per-token pricing exists yet (source: Wayback deepinfra.com/pricing 2023-02).

Launch: per-second execution-time billing only - Earliest archived pricing page bills purely by inference execution time — $0.000
captured
Trivia
  • · DeepInfra publishes a per-million-token rate for nearly every open model it hosts — from Llama-3.1-8B at $0.02 in / $0.04 out to flagship DeepSeek-V4-Pro at $1.30 in / $2.60 out — making it one of the most transparent open-model inference price lists in the market, with prompt-cache rates shown inline.
  • · The same DeepInfra account spans four billing primitives at once: per-token LLM and embedding APIs, per-image Flux generation priced by resolution and step count, per-minute Voxtral audio transcription, and per-GPU-hour on-demand B200/H200 containers — a single bill across four metering units.
  • · DeepInfra raised a $107M Series B to scale its inference cloud, and runs a DeepStart program granting qualifying startups 1,000,000,000 free tokens (valued at DeepSeek-V3.1 prices) for companies that have raised $250K–$10M and were founded within the last two years.

Questions & answers

How much does DeepInfra cost per token?
Per-1M-token rates vary by model. Examples: DeepSeek-V3.1 $0.25 in / $0.95 out, DeepSeek-V4-Pro $1.30 / $2.60, Llama-3.3-70B-Turbo $0.10 / $0.32, Llama-3.1-8B $0.02 / $0.04, Gemini 2.5 Pro $1.25 / $10.00, Kimi-K3 $2.85 / $14.25. Rates span from $0.02 per 1M input tokens (Llama-3.1-8B) to $2.85 (Kimi-K3, $14.25 out), the most expensive row now on the /pricing page. As of 2026-08-28 the page's Claude section lists only claude-haiku-4-5 ($1.00 / $5.00) — six other Claude models, including the $10.00/$50.00 claude-fable-5 row that had been the single highest rate on the list, no longer appear (they were still listed as of 2026-08-26). Many models also show a discounted cached-prompt input rate inline. These are the default Standard-tier (1×) rates; a per-request service tier multiplies them — Flex bills 0.8× for non-production and asynchronous work, Priority bills 1.5× for faster time-to-first-token.
Does DeepInfra charge per GPU-hour for dedicated hardware?
Yes. Custom-LLM deployments on dedicated GPUs bill in minute granularity (invoiced weekly): A100 $0.89, H100 $2.20, H200 $2.69, B200 $3.69, B300 $4.89 per GPU-hour. The on-demand GPU-instances product lists 1×B200 $3.69/hr, 2×B200 $7.38/hr, 4×B200 $14.76/hr, 8×B200 $29.52/hr with no egress fees.
What is DeepInfra DeepCluster pricing?
DeepCluster is a reserved, customer-owned NVIDIA B300 GPU cluster (256–5,000 GPUs) that DeepInfra procures and operates. All-in pricing is $2.99/GPU-hr on a 3-year term and $1.98/GPU-hr on a 5-year term, marketed as up to 54%–70% cheaper than a $6.50/GPU-hr public-cloud reference. Terms are 3 to 5 years and are sales-led (contact [email protected]).
Does DeepInfra have a free tier?
No published free tier — a card on file or pre-payment is required before you can use the service. However, the DeepStart program grants qualifying startups 1,000,000,000 free tokens (valued at DeepSeek-V3.1 prices) for companies that raised $250K–$10M and were founded within the last 2 years.
How does DeepInfra billing work?
DeepInfra bills pay-as-you-go with no contracts or upfront costs. Every account sits in a usage tier (Tier 1 $20 threshold up to Tier 5 $10,000); an invoice generates at the start of each month and whenever the tier's invoicing threshold is reached. You can set a spending limit, and accounts are capped at 200 concurrent requests by default.
How are DeepInfra image and audio models priced?
Image and audio models are billed per media unit rather than per token. Flux image generation is priced per image scaled by resolution and iteration count (e.g. FLUX-1-schnell $0.0005 × (w/1024) × (h/1024) × iters; FLUX-2-pro listed at $0.015/image, whose model page bills $0.03 for the first output megapixel then $0.015 per additional megapixel). Voxtral audio transcription is billed per minute of audio input ($0.00100/min Mini, $0.00300/min Small). Text-to-speech uses a different unit again: Qwen3-TTS and Qwen3-TTS-VoiceDesign cost $20.00 per 1M characters, so speech-synthesis cost scales with the text sent in rather than the audio length produced.