Ask
Sharpens 22 companies · First observed November 2023 · Updated August 2026 Explore in the graph

Token prices fall as access tiers rise

Quick answer

AI API token prices fall with almost every model generation — frontier rates dropped roughly 10× from the GPT-4 era to the GPT-5 era — while the price of human-facing access climbs toward a $200/mo ceiling. Compute deflates; access inflates.

~10× cheaper per token, GPT-4 → GPT-5 era

What's happening — and why

What's happening: the cost of calling an AI model through its API keeps dropping. Each new model generation tends to launch cheaper than the one it replaces, and vendors add further discounts — prompt caching, batch processing — in between releases.

Why: raw inference is becoming a commodity. Open-weight models set aggressive price floors, serving efficiency improves, and competition is intense, so vendors pass the savings to developers to win volume. They keep their pricing power for human-facing subscriptions instead, where willingness-to-pay is far higher than for a raw token.

How it works

$ / 1M TOKENS TIME → GPT-4 4o GPT-5 $200 access tier
Per-token API price falls each model generation (blue); human-facing access climbs (amber).

Evidence over time

37 supporting · 9 counter — hover or tap a point for detail, click to jump to the row.

supports ↑ challenges ↓ 2023 2024 2025 2026
supporting evidence counterexample

Evidence

Company Date What happened
Anthropic May 2026 Opus rebased from the legacy $15/$75 band to $5/$25 per 1M and Haiku to $1/$5 — a ~3× cut on the flagship tier within the same product line.
Novita AI Apr 2026 Kimi K2.6 launched at $0.95/$4.00 then trimmed to $0.8/$3.4 by the 2026-06-02 capture — a downward revision within ~6 weeks, mid-generation.
OpenAI Nov 2023 GPT-4 Turbo launched ~10× cheaper than the original GPT-4.
OpenAI May 2024 GPT-4o launched ~50% cheaper than GPT-4 Turbo; 4o-mini (2024-07) pushed frontier-ish quality to commodity price.
Google Nov 2024 Gemini 1.5 token pricing cut ~50–65%.
DeepSeek May 2024 V2 at $0.14/1M input — a market-shock price that reset expectations.
DeepSeek Dec 2024 V3 at $0.27/1M — frontier-class performance at commodity rates.
Anthropic Aug 2024 Prompt caching cut input cost by up to 80% without a model change.
Hume AI Jun 2026 EVI 1 launched at $0.102/min (March 2024); EVI 2 cut to $0.072/min (Sept 2024); EVI 3 now $0.07/min; EVI 4 MINI $0.04/min — ~60% per-minute deflation across two years of model generations. A clear per-minute token-deflation example beyond text tokens.
Together AI Jun 2026 Cut managed GPU Cluster rates twice in one week — on-demand HGX H100 $4.79→$3.99/hr and reserved 91–180d H100 to a $3.09/hr floor — after a 2026-06-24 round that also cut serverless model rates and began publishing cached-input discounts. Weekly-cadence deflation at the managed-cluster layer.
Vercel Jun 2026 Cut its v0 Max Fast model ~3× ($30/$150 → $10/$50 per 1M input/output; cache read $1/1M) and retired the legacy Ultra plan — generational token deflation reaching the application layer's own model lineup.
Hume AI Jun 2026 Added an in-table model toggle: EVI 4 MINI runs ~half the EVI 3 per-minute overage ($0.035–$0.02 vs $0.07–$0.04) and Octave 2 ~halves Octave 1's per-character overage while doubling minutes per character — a quiet per-unit price cut at every paid tier without moving headline prices.
Anthropic Jul 2026 Claude Sonnet 5 launched at introductory $2/$10 per 1M through Aug 31, 2026 (batch $1/$5), BELOW the $3/$15 band Sonnet held across four generations since Claude 3 — the standard rate reverts to $3/$15 on Sep 1. The first downward move on the production-standard Sonnet anchor: generational deflation reaching the tier that had been the corpus's most durable price constant.
Speechmatics Jul 2026 Cut real-time enhanced STT -23% ($0.56→$0.43/hr) — the SKU powering live voice agents — for the third self-serve rate cut since 2023, and added a multilingual Batch Melia 1 floor at $0.129/hr (the new 'from' headline). Per-minute audio-model deflation beyond text tokens.
Lambda Labs Jun 2026 Raised on-demand GPU rates significantly: H100 SXM from $2.99 to $3.99/hr (+33%), A100 SXM 80GB from $1.79 to $2.79/hr (+56%). GPU on-demand cloud access inflating as demand surges — the opposite of token-API deflation.
You.com Sep 2025 Top consumer tier repriced ~6.7× — $30 Team replaced by $200 Max.
Character.ai Feb 2026 Introduced mid-chat ads for free users — monetisation rising, not falling.
Anthropic Jun 2026 Opened a new top API band — Claude Fable 5 / Mythos 5 at $10/$50 per 1M, double Opus 4.8's $5/$25 — while leaving the existing Opus/Sonnet/Haiku bands unchanged. The frontier ceiling re-inflated 2× even as the rest of the ladder held; deflation is per-tier, and the most-capable tier resets upward.
Mistral AI Jul 2026 The corpus's first frontier-lab token *increase*: Mistral Small 4 rose 50% on input / 100% on output ($0.1/$0.3 → $0.15/$0.6 per 1M) and OCR doubled to $4/1K pages — a cost-sensitive model repriced UP mid-life. Flagship Medium 3.5 ($1.5/$7.5) and Large 3 ($0.5/$1.5) held. Deflation is not monotonic per SKU: a vendor can raise a specific under-priced model even as the generational trend falls.
xAI Jul 2026 Shipped grok-4.5 at $2 in / $6 out per 1M — ABOVE the still-live grok-4.3 ($1.25/$2.50), the first time xAI priced a new flagship higher than its predecessor rather than continuing its multi-year per-token markdown. It keeps grok-4.3 available and reserves the premium for higher intelligence — the frontier ceiling re-inflating while the prior tier holds.
DeepInfra Jul 2026 The deflation poster child reversed: raised dedicated GPU-hour rates 16-32% (H100 +23%, H200 +23%, B200 +32%, B300 +16%) and lifted DeepSeek-V3.1 tokens from $0.21/$0.79 to $0.25/$0.95 per 1M — the first across-the-board increase from a vendor whose reputation in cost-sensitive communities was built on relentless public price cuts. Even a commodity-inference vendor can move its whole card up when Blackwell supply and demand tighten.
Google Jul 2026 Added the AlphaEvolve agent to Vertex AI priced as the base Gemini model rate PLUS a 2x agent surcharge (3x all-in) — Gemini 3.1 Pro Preview $6/$36 per 1M vs the $2/$12 base. Headline Gemini token rates were unchanged; the increase is an explicit agent-premium multiplier layered on top, the frontier re-inflating via an agent surcharge rather than a token hike.
Together AI Jul 2026 Deflation continues at the infra layer: cut on-demand Dedicated Inference — HGX H100 $6.49→$5.49/hr (-15%) and HGX B200 $11.95→$8.99/hr (-25%) — while launching a Provisioned Throughput (PTU) SKU at $0.05/PTU-min. Managed-GPU access keeps falling even as the model-intelligence ceiling rises elsewhere.
RunPod Jul 2026 Trimmed two Secure Cloud Pod rates — H100 SXM (80GB) $3.29→$2.99/hr (-9%) and RTX Pro 6000 (96GB) — pressing its cost-leader positioning as Blackwell-generation supply expands. Bare GPU-hour deflation persists at the discount-cloud layer.
Baseten Jul 2026 The catalog-prune mechanic in its purest form: the Model APIs rate card went from 11 SKUs to 8, retiring GLM 5.1 ($1.30 in / $4.30 out), GLM 5 ($0.95/$3.15), Kimi K2.5 ($0.60/$3.00) and NVIDIA Nemotron 3 Super ($0.30/$0.75), while adding Inkling at $1.00/$0.17 cache/$4.05. With Nemotron 3 Super gone the lowest published input rate outside GPT OSS 120B rose to $0.60, and the cheapest published cache-input rate rose from $0.06 to $0.12 — a floor increase delivered without repricing a single surviving model. Eight days later (2026-07-29) it re-expanded to 10 SKUs, added Kimi K3 at $3.00/$0.30/$15.00 and GLM-5.2 Fast at $2.10/$0.21/$6.60, and CUT GLM-5.2 cache input 46% ($0.26 → $0.14).
Groq Jul 2026 Shrank rather than repriced: the published serverless catalog fell from 8 priced LLMs to 6 between the 2026-07-14 and 2026-07-21 captures, removing Llama 4 Scout 17Bx16E ($0.11/$0.34) and Qwen3 32B ($0.29/$0.59) from both the rate card and the docs Supported Models page, with every retained model price stable. The enterprise-only column thinned too (Qwen3-VL 32B gone, Minimax M2.5 replaced by M2.7). Removing the two cheapest models lifts the effective floor while leaving no price diff to detect.
Novita AI Jul 2026 The arc completed on one platform: catalog pruned 221 → 196 (2026-07-21) → 172 (2026-07-29), a delisting of 49 models in eight days, alongside a straight 10x raise on Flux.1 Kontext Pro from $0.036 to $0.36 per image — which now costs more than the ostensibly higher-tier Kontext Max at $0.072 (unchanged). The prior $0.036 rate was independently confirmed across three earlier captures, so this is a genuine vendor price change. The same sweep moved Tencent Hy3 off $0/M to $0.14/M input and $0.58/M output, and added Kimi K3 at $3/$15 as the most expensive LLM on a card whose positioning is cheapest-place-to-run-open-weights.
DeepInfra Jul 2026 Second raise in 15 days, and it landed on the entry model the vendor's own catalog cites as its cheapest reference: Meta-Llama-3.1-8B-Instruct-Turbo output rose 33%, $0.03 → $0.04 per 1M (input held at $0.02). gemma-4-31B-it-turbo was cut in the same capture ($0.12/$0.37 → $0.09/$0.34, -25% input / -8% output). Headline per-GPU-hour rates, DeepCluster reserved pricing, usage tiers and the Flex/Standard/Priority ladder were unchanged. Raising the cheapest SKU is the highest-leverage way to lift a rate card's perceived floor.
Weights & Biases Jul 2026 The largest single-SKU raise in the corpus, on the cheapest model on the card: DeepSeek V4-Flash went from a flat $0.01/$0.01 per 1M to $0.14 input / $0.07 cached / $0.28 output — 14x on input, 28x on output, plus a new cached tier the old two-part rate lacked. It lifted the whole catalog floor: the cheapest published input rate on W&B Inference rose from $0.01 to $0.05 (GPT OSS 20B, IBM Granite 4.1 8B, JetBrains Mellum2 12B). Eight days later W&B cut GLM 5.2 ~45% ($1.39/$4.40 → $0.76/$2.42), Kimi K2.7 Code and K2.6 by 25-32%, and added MiniMax M3 at $0.23/$0.96.
Together AI Jul 2026 First reserved-GPU raise since the aggressive June cuts, isolated to the reserved-H100 tenors: $3.59 → $3.69/hr (7-30 days), $3.29 → $3.45 (31-90), $3.09 → $3.19 (91-180). On-demand H100 and every published H200/B200 on-demand and reserved rate held flat. In the same capture all four serverless speech-to-text models were re-metered onto a unified $0.0015 per audio minute, cutting Nemotron 3.5 ASR 67% (from $0.0045/min) — a raise and a 67% cut on one rate card in one cycle.
Moonshot AI Jul 2026 The frontier ceiling re-inflating at a vendor built on undercutting: Kimi K3 (2.8T parameters, 1,048,576-token context) launched at $3.00 input / $0.30 cache-hit / $15.00 output per 1M — roughly 3.2x the input and 3.75x the output of Kimi K2.6 ($0.95/$4.00, cache $0.16) — opening Moonshot's first genuine 5x ladder between frontier and mainstream after two years of aggressive per-token share-taking. A Kimi K2.7 Code SKU landed at $0.95/$4.00 with a HighSpeed variant at exactly double ($1.90/$8.00). Separately (2026-07-21) the legacy moonshot-v1 family — which priced identical weights differently by requested context window, making context length literally the meter — was dated for full sunset on 2026-08-31.
OpenAI Jul 2026 The everyday tier repriced UP while the top two held: the GPT-5.6 line launched as GPT-5.6 Sol ($5.00/$0.50 cached/$30.00 per 1M), Terra ($2.50/$0.25/$15.00) and Luna ($1.00/$0.10/$6.00). Sol and Terra match the outgoing GPT-5.5/GPT-5.4 price points exactly, but Luna at $1/$6 sits ABOVE the GPT-5.4 mini it replaces ($0.75/$4.50) — and the GPT-5.4 nano tier ($0.20/$1.25) has no successor in the headline lineup. Generational deflation held at the flagship and mid tiers and reversed at the cheap end.
Google Jul 2026 Deflation and inflation in one release: Gemini 3.6 Flash becomes the default fast model and CUTS output 17% ($9.00 → $7.50 per 1M) while holding $1.50 input and the same 5,000-free-grounding-queries allowance; but Gemini 3.5 Flash-Lite lands at $0.30/$2.50, priced level with Gemini 2.5 Flash rather than undercutting Gemini 3.1 Flash-Lite at $0.25/$1.50 — so the new cost-efficient option is more expensive than the one it sits beside. Consumer ceilings confirmed: Google AI Plus $4.99/mo, Pro $19.99, Ultra $99.99 (5x) and $199.99 (20x).
Anthropic Jul 2026 Deflation holding at fixed capability, stated as clearly as the corpus gets: Claude Opus 5 replaced Opus 4.8 as the recommended default for complex agentic coding at the IDENTICAL $5/$25 per 1M — same 1M context, same 128k max output, same prompt-caching multipliers (cache write $6.25/$10, cache read $0.50), same Batch API rates ($2.50/$12.50) and Fast-mode eligibility ($10/$50). Opus 4.8, 4.7, 4.6 and 4.5 all remain live at $5/$25; Opus 4.1 retires 2026-08-05. A capability refresh at a flat price is the cleanest form of deflation — more model for the same money.
Hugging Face Jul 2026 Cross-cloud arbitrage cutting GPU access by more than half: Inference Endpoints added AWS-hosted NVIDIA H100 at $4.50/hr against the existing GCP-hosted H100 at $10.00/hr, plus its first Blackwell-generation options — B200 at $9.25/hr and RTX PRO 6000 at $2.75/hr, all scaling linearly at 2x/4x/8x. ZeroGPU Spaces now disclose they run on RTX PRO 6000 Blackwell hardware. Per-hour GPU deflation arriving through cloud choice rather than a rate cut.
OpenAI Aug 2026 Deflation at fixed capability got FASTER, and the pass-through is now near-instant. GPT-5.6 Luna launched on 2026-07-23; within twelve days its per-token rate was cut roughly 80% across every dimension — input $1.00 to $0.20, cache write $1.25 to $0.25, cache read $0.10 to $0.02, output $6.00 to $1.20 per 1M — with Terra down 20% ($2.50/$15.00 to $2.00/$12.00) and Sol unchanged at $5/$30. The cut then propagated identically through five corpus resellers in eight days: Glean's Model Hub (2026-08-04, which also moved Luna from Premium to Standard so it falls inside the included 100/user/week allowance), Cursor's Other Models pool (08-06 and re-confirmed 08-11), Augment Code (08-11, its largest single-model move since it began publishing per-model pricing), Vercel/v0 (08-11) and LiveKit Inference (08-11, per-minute $0.0040 to $0.0008). A frontier-adjacent model losing 80% of its price inside a fortnight, with zero-markup resellers transmitting it at once, is the fastest documented deflation event in this corpus.
Fireworks AI Aug 2026 The published floor keeps falling even as the same vendor's GPU card rises. Fireworks added NVIDIA Nemotron 3.5 Lightning 30B A3B at $0.05 input / $0.01 cached / $0.20 output per 1M — Standard-only, no Priority path — undercutting the prior cheapest published input rate on the entire card, OpenAI GPT OSS 20B at $0.07. Muse Glimmer 30B joined at $0.35/$0.04/$1.50 Standard with a Priority path at exactly 1.5x. Perplexity lists the same Nemotron model on its Gateway API at $0.0115 input / $0.17 output per 1M, an order of magnitude below anything else in that catalog. Two days earlier the same vendor had published a 2026-09-01 increase across every on-demand GPU tier.

Counterexamples

  • DeepSeek · Aug 2026 — THE FLOOR-SETTER RAISED, AND INVENTED A FIFTH INFLATION ROUTE. DeepSeek's V2 launch at $0.14 per 1M input in May 2024 is the event this trend credits with resetting the industry's price-floor expectations. On 2026-08-11 its docs added a footnote warning of an unspecified 'significant' increase; on 2026-08-14 it published the card, effective 16:00 UTC on 2026-08-16. Pricing splits into peak (01:00-04:00 and 06:00-10:00 UTC) and off-peak windows, off-peak at exactly half of peak. V4-Flash cache-miss input goes from a flat $0.14/1M to $0.22 off-peak and $0.44 peak; output from $0.28 to $0.66/$1.32; V4-Pro cache-miss input from $0.435 to $0.66/$1.32. The steepest move is cache-hit input, DeepSeek's most-marketed rate: $0.0028 to $0.007/$0.014 on Flash (2.5x/5x) and $0.003625 to $0.022/$0.044 on Pro (over 12x). No grace period, no legacy-rate opt-out. This is the first time-of-use token pricing in the corpus: the other four inflation routes price a different product; this prices the same product differently depending on the hour.
  • Fireworks AI · Aug 2026 — Infra-layer inflation on a published schedule, not a supply feed. Fireworks posted a forward-dated increase across its entire on-demand GPU card effective 2026-09-01 — H100/H200 $7.00 to $8.00/hr (+14%), B200 $10.00 to $13.00 (+30%), B300 $12.00 to $15.00 (+25%), GB300 $18.00 to $20.00 (+11%) — and it is, by the company's own framing, the FIRST repricing of an already-published on-demand SKU since the product launched in January 2024; every prior event had been the addition of a tier. The page shows a two-column current-versus-September table rather than an immediate change. It landed three weeks after a $1.5B Series D at a $17.5B valuation and an announced $1B ARR, and one day after the same vendor introduced a flat 1.5x premium for US- or Europe-pinned dedicated deployments.
  • SambaNova · Aug 2026 — The prune-then-raise arc's raise stage executed as a straight REVERSAL of a published cut. Less than three weeks after cutting gemma-4-31B-it to $0.22 input / $0.59 output per 1M, SambaNova Cloud restored it to $0.38/$1.15 — its pre-cut June rate — a 73% input and 95% output increase on an unchanged SKU, leaving gpt-oss-120b as the sole cheapest model on the six-model card. Every other model held (DeepSeek-V3.1/V3.2 $3.00/$4.50, Llama-3.3-70B $0.60/$1.20, MiniMax-M2.7's $0.06 cached input). Three days later it also dropped the no-card $5 free-credit grant. A cut that reverts inside a quarter is invisible in any year-over-year comparison, because the discount leaves the record with it.
  • Sarvam AI · Aug 2026 — A ~7x raise that only one of the vendor's two first-party surfaces admits to. Sarvam's API reference docs, which showed Sarvam-105B at Rs 4 input / Rs 2.5 cached / Rs 16 output per 1M as recently as 2026-07-30, now price it at Rs 29.28 / Rs 10.98 / Rs 73.2 across two model IDs (`sarvam-105b` and `sarvam-105b-conversations`) — while sarvam.ai/api-pricing, captured the same day, still advertises Rs 4/2.5/16 alongside Sarvam-30B at Rs 2.5/1.5/10, byte-for-byte unchanged. The docs page also added two beta-priced third-party models never listed on any Sarvam surface before: Gemma-4 31B (Rs 36.6/13.73/91.5) and GLM 5.2 (Rs 128.1/23.79/402.6). A buyer reading the marketing page would underestimate their bill by roughly 7x.
  • Novita AI · Aug 2026 — The prune stage run twice in four days, ending with an empty free shelf. On 2026-08-11 three models tagged Time Limited Free converted to paid — Macaron V1 Venti to $1.5/M in ($0.3 cache read) / $4.5/M out, Macaron V1 Tall to $0.45/$2.6, Ling 3.0 Flash to $0.06/$0.18 — with Ling 3.0 Tiny added at $0/M as the replacement free listing. On 2026-08-14 Ling 3.0 Tiny was delisted outright rather than repriced, the shortest lifespan of any free listing tracked on this card, leaving no $0 language model at all and the catalog easing from 177 to 174. Weights & Biases ran the same play on 2026-08-12, delisting five models (32 to 28 SKUs, including Qwen3 235B A22B-2507 at $0.10/$0.10 and Kimi K2.5 at $0.60/$3.00) and adding one, Nemotron 3.5 Lightning at $0.10/$0.05/$0.25 — pure roster churn with no surviving model repriced.
  • Hyperbolic · Jul 2026 — The sharpest infra-layer counterexample of 2026: published on-demand starting rates nearly DOUBLED in about five weeks — H100 SXM $1.50 → $2.89/hr (+93%), H200 $2.40 → $3.49 (+45%), B200 $3.50 → $5.99 (+71%) — while the per-hour rates previously shown for RTX 4090 ($0.30) and RTX 3070 ($0.16) were pulled entirely, replaced by a 'from $0.20/GPU/hr' catalog claim. Serverless per-million-token rates were untouched, so the raise is isolated to GPU access. The page notes pricing is 'refreshed weekly based on the best available rates from suppliers', which makes this aggregated third-party supply repricing rather than a list-price decision — but it is the clearest evidence that GPU-hour access does not deflate monotonically.
  • Zhipu AI · Jul 2026 — A frontier lab raising a human-facing AI subscription 80-140% while its token rates stayed competitive: the GLM Coding Plan — which launched September 2025 with a $3 first month and settled near $10/mo for Lite after the February 2026 promo withdrawal — now lists at $18 / $72 / $160 per month for Lite / Pro / Max, with a billing-term ladder bottoming at $12.60/mo on annual (renewing at $151.20/yr from year two). That is roughly 26% above the previous effective Lite rate and 40% at Max. Quota language also became less legible, replacing 'about 80 prompts per 5 hours' with relative multiples (Pro = 5x Lite, Max = 20x Lite). The scissors in miniature inside one vendor: cheap tokens, dearer access.
  • Poe · Jul 2026 — Consumer access inflating by shrinkflation rather than by a price rise — the cleanest measurement of the scissors' human-facing blade. The $19.99/mo Premium plan held its sticker ($199.99/yr) while its bundled allowance fell from 1,000,000 to 660,000 compute points per month, a 34% cut that works out to roughly a 51% increase in price per point, confirmed in the plan's own product data on poe.com/subscription_plans. Because points map to model cost, heavy users of frontier models absorb the entire increase — so the same buyer pays more per token of intelligence in the same month that the underlying token rates fell at Baseten, W&B and SambaNova.
  • Composio · Jul 2026 — Deflation does not reach the agent-tool layer: from 2026-08-15 Composio takes tool-call overage from $0.299 to $4.00 per 1K (a 13x increase) while cutting Pro's included tool calls from 200,000 to 50,000 at an unchanged $29, and splits a single meter into six dimensions (tool calls, trigger events, LLM tokens, premium credit, sandbox GB-hr, filesystem GB). The LLM-token dimension rides the deflating market; the orchestration dimensions above it do not. Existing customers are grandfathered only through 2026-12-31.

Trivia

  • DeepSeek V2's May 2024 launch at $0.14 per million input tokens reset the entire industry's price-floor expectations for coding-capable models — within 18 months every major frontier vendor had cut per-token prices at least once. No prior software category has experienced a 10-100× cost compression within two years of reaching mainstream adoption, making token deflation structurally unlike any historical software pricing precedent.

  • Anthropic's August 2024 prompt-caching launch produced the largest single-day effective price reduction in the corpus (up to 80% off cached input) without any model change — demonstrating that token deflation arrives through infrastructure optimisation (caching, batching) as well as through model-generation cuts, and that the two mechanisms compound. A buyer using both batch (50% off) and caching (up to 80% off cached portions) can reduce effective input cost by 90%+ on workloads with stable system prompts.

  • The scissors dynamic — token API prices falling while consumer subscription ceilings rise — is the defining pricing paradox of the 2024-2026 AI market. You.com's 6.7× consumer repricing (from $30 Team to $200 Max) and Character.ai's ad insertion for free users both happened while the underlying token rates kept deflating. The divergence shows that AI products have successfully decoupled human-facing access pricing from the compute costs beneath them.

  • Anthropic's June 2026 Fable 5 / Mythos 5 launch is the corpus's clearest proof that token deflation is per-tier, not per-vendor: the company opened a new $10/$50 flagship band — exactly double Opus 4.8's $5/$25 — without cutting or raising any existing tier. Three weeks earlier (2026-05-30) it had rebased Opus itself ~3× downward from the legacy $15/$75. So inside one product line and one month, the established tiers deflated while the 'best-available' ceiling re-inflated 2×, confirming that what falls is the price of a fixed capability level, not the price of the frontier.

  • Together AI cut its on-demand H100 rate twice inside a single week ending 2026-06-30 ($4.79 → $3.99/hr), and its $3.09/hr reserved floor now sits BELOW Lambda Labs' $3.99 on-demand rate logged just three weeks earlier — the first time in the corpus that a managed-cluster vendor's reserved price has undercut a peer's on-demand price, evidence that the 'GPU scarcity inflates access' counter-current is breaking down under managed-cluster competition.

  • The $3/$15 Claude Sonnet band was the single most durable price constant in the corpus — held flat across four model generations from Claude 3 through Sonnet 4.6 while capability climbed. On 2026-07-06 Sonnet 5 finally bent it, launching at an introductory $2/$10 through Aug 31 before reverting to $3/$15. The anchor didn't break permanently — it dipped — but even a temporary intro cut on the corpus's most stable price is the clearest sign that generational deflation now reaches the tiers vendors most want to keep stable.

  • The same 2026-07-06 batch that pushed Sonnet's anchor down also produced the corpus's first frontier-lab token *increase*: Mistral raised Small 4 by 50% on input and 100% on output. Two frontier labs moved their token prices in opposite directions on the same date — the cleanest proof that token deflation is a dominant trend, not a law: a specific under-priced SKU can still be repriced up mid-life.

  • DeepInfra broke its own brand on 2026-07-14: the vendor whose reputation in cost-sensitive communities was built on relentless public price CUTS raised dedicated GPU-hour rates 16-32% across the board (H100 +23%, B200 +32%) and lifted its popular DeepSeek-V3.1 tokens from $0.21/$0.79 to $0.25/$0.95 — the corpus's first case of a commodity-inference vendor raising its WHOLE card rather than one under-priced SKU, evidence that even the deflation poster child can be forced up when Blackwell-generation supply and demand tighten.

  • xAI priced a new flagship ABOVE its predecessor for the first time on 2026-07-14: grok-4.5 launched at $2/$6 per 1M versus the still-live grok-4.3 at $1.25/$2.50 — reversing xAI's multi-year per-token markdown and keeping the cheaper model live so the premium is reserved for the higher-intelligence tier. Combined with Google's AlphaEvolve 3×-all-in agent surcharge the same day, the frontier ceiling re-inflated via BOTH a higher flagship and an agent multiplier while Together (−15%/−25%) and RunPod kept cutting GPU beneath them.

  • The corpus's most under-reported inflation mechanic moves no price at all. Baseten cut its Model APIs card from 11 SKUs to 8 on 2026-07-21, retiring GLM 5.1 ($1.30/$4.30), GLM 5 ($0.95/$3.15), Kimi K2.5 ($0.60/$3.00) and Nemotron 3 Super ($0.30/$0.75) — which lifted the cheapest published cache-input rate on the card from $0.06 to $0.12 without repricing a single surviving model. Groq did the same thing the same day, dropping Llama 4 Scout ($0.11/$0.34) and Qwen3 32B ($0.29/$0.59) with every retained price stable. Delisting the cheap end is a price increase that leaves no diff on any rate.

  • Novita delisted 49 of 221 models in eight days — 221 → 196 on 2026-07-21, then 196 → 172 on 2026-07-29 — and in the same window took Flux.1 Kontext Pro from $0.036 to $0.36 per image, a 10x rise that makes it pricier than the ostensibly higher-tier Kontext Max at $0.072. The old rate was confirmed across three independent prior captures, so this is a genuine vendor move on a platform whose entire positioning is "cheapest place to run open-weight models."

  • Two vendors ran cuts and raises on the same rate card within eight days. Baseten pruned to 8 SKUs (lifting its floor) on 2026-07-21, then re-expanded to 10 and CUT GLM-5.2 cache input 46% ($0.26 → $0.14) on 2026-07-29. Weights & Biases raised DeepSeek V4-Flash 14x on input and 28x on output ($0.01/$0.01 → $0.14/$0.28) on 2026-07-21, lifting its catalog floor from $0.01 to $0.05 — then cut GLM 5.2 roughly 45% and Kimi K2.7/K2.6 by 25-32% on 2026-07-29. Directional summaries of "the infra layer" no longer survive a week.

  • Moonshot ended two years of undercutting in one launch: Kimi K3 arrived on 2026-07-16 at $3.00/$15.00 per 1M — 3.2x the input and 3.75x the output of Kimi K2.6 ($0.95/$4.00) — opening the company's first genuine 5x ladder between its frontier model and its mainstream line. Nine days earlier its most distinctive mechanic was dated for retirement: the moonshot-v1 family, which charged different per-token rates for the SAME weights depending on the context window requested, sunsets 2026-08-31, retiring context length as a billing axis.

  • The "everyday" tier was repriced UP by both hyperscale labs in the same week, which no prior generation did. OpenAI's GPT-5.6 Luna launched at $1/$6 per 1M against the outgoing GPT-5.4 mini's $0.75/$4.50, while Sol ($5/$30) and Terra ($2.50/$15) held exactly. Google priced Gemini 3.5 Flash-Lite at $0.30/$2.50 — above Gemini 3.1 Flash-Lite's $0.25/$1.50 — even as Gemini 3.6 Flash cut output 17% ($9.00 → $7.50). Deflation held at the top and the middle; the cheap tier is where both raised.

See all pricing trivia

For buyers

Re-baseline token-cost assumptions at least twice a year — a model you priced six months ago is usually cheaper now, and a newer model may be cheaper and better. Don't extrapolate the deflation to seats: model the workload as raw metered tokens (falling) separately from per-seat access (rising), and pick the cheaper envelope deliberately.

For vendors

To ride deflation you need per-model rate cards that can change without breaking contracts, plus structural discounts (caching, batch) to cut effective price without touching headline rates. To monetise access, a premium seat tier with priority or uncapped usage captures the value that falling tokens leave on the table.

Outlook — what to watch

Expect another leg down: open-weight models (DeepSeek, Llama-class) keep resetting the floor, and caching/batch discounts are spreading. The deflation only stalls if frontier capability stops commoditising — watch whether any vendor can hold a premium token price on a genuinely differentiated model. Access tiers still have room to climb.

Bottom line

Compute is deflating relentlessly while access inflates. Every frontier vendor in the corpus has cut per-token prices at least once; the value is migrating from the token to the seat.

FAQ

Are AI API prices going up or down?

Down. Every frontier vendor in the corpus has cut per-token API prices at least once — usually at each model release — and layered caching and batch discounts on top. Consumer subscriptions are the exception; those are rising.

Why are ChatGPT and Claude subscriptions getting more expensive if tokens are cheaper?

The token (raw compute) and the seat (human access) sit on different curves. Vendors pass compute deflation to API developers but capture value from power users through premium ~$200/mo access tiers.

How often should I re-check my model costs?

At least twice a year. A model you priced six months ago is usually cheaper now, and a newer one is often cheaper and better — Claude 3.5 Sonnet beat Opus at a fraction of the price.

All trends