New 7 companies · First observed July 2026 · Updated August 2026 Explore in the graph

Zero-markup resale transmits a model cut in days

Quick answer

Between 2026-08-04 and 2026-08-11, five corpus companies — Glean, Cursor, Augment Code, LiveKit and Vercel — repriced the same two OpenAI models by exactly the same two percentages: GPT-5.6 Luna down 80%, GPT-5.6 Terra down 20%. None of them announced a price change, because none of them made one: each already billed third-party model tokens at cost, so an upstream rate-card move passed straight through. The synchrony is the finding, not the discount — one decision surfaced as five separate 'price change' events, on four dates, in three different denominations.

5 vendors cut the same two OpenAI SKUs by exactly 80% and 20% inside 11 days

What's happening — and why

What's happening: OpenAI launched the GPT-5.6 line on 2026-07-23 — Sol $5/$30, Terra $2.50/$15, Luna $1/$6 per 1M tokens. Glean's Model Hub table then cut Luna 80% ($1.00 to $0.20 in, $6.00 to $1.20 out) and Terra 20% ($2.50 to $2.00, $15.00 to $12.00) on 08-04. Cursor matched on 08-06, and Augment Code, LiveKit and Vercel all matched on 08-11. Netlify ran the identical mechanic on a different upstream three days later, halving Gemini 3.6 Flash ($1.50/$7.50 to $0.75/$3.75 per 1M) at a fixed 180-credits-per-dollar conversion, with no plan change at all (Free $0 / Personal $9 / Pro $20 unchanged).

Why: each of these vendors had already published an at-cost passthrough policy before the window opened. Cursor's docs state that its Other Models pool bills third-party API pricing at cost; Vercel attributes its move to the zero-markup AI Gateway; Netlify converts provider rates at a fixed 180 credits per $1. Having ceded pricing authority for third-party tokens, they don't decide these numbers — they inherit them. Their model prices are consequences, not decisions.

What makes it hard to see: the denomination changes at each hop. Glean, Cursor and Augment quote per 1M tokens; LiveKit quotes per minute ($0.0040 to $0.0008/min) and still lands on exactly 80%; Netlify quotes in credits. Vercel's v0 Mini fell from $1/$5 to $0.20/$1.20 — landing on Luna's exact new rate — yet the repricing is invisible on Vercel's own pricing page, because v0's plan cards never show token rates at all. And the direction of the underlying move matters: Luna launched at $1/$6, above the outgoing GPT-5.4 mini at $0.75/$4.50, so the 80% cut twelve days later was a correction to a launch price. Every reseller inherited both moves.

How it works

ONE UPSTREAM MOVE · FIVE PUBLISHED PRICE CHANGES · NOBODY DECIDED OpenAI GPT-5.6 upstream rate card · launched 07-23 LUNA $1/$6 → $0.20/$1.20 TERRA $2.50/$15 → $2/$12 DATERESELLERMOVE PUBLISHEDDENOMINATED IN 08-0408-0608-1108-1108-11 GleanCursorAugment CodeLiveKitVercel v0 −80% / −20% −80% / −20% −80% / −20% −80% / −20% −80% / −20% per 1M tokens per 1M tokens per 1M tokens per MINUTE, still 80% not on the plan cards 08-11 SambaNova +73% / +95% own rate, does not transmit ● TRANSMITS AT COST (5) × SETS ITS OWN RATE (3) SAME MECHANIC, DIFFERENT UPSTREAM · NETLIFY 08-14 · GEMINI 3.6 FLASH −50% AT 180 CREDITS/$1
One upstream rate card fans out into five resellers that publish the identical move on four dates in three denominations (blue); the margin-taking branch does not transmit (orange).

Evidence over time

7 supporting · 0 counter — hover or tap a point for detail, click to jump to the row.

supports ↑ challenges ↓ 2026
supporting evidence counterexample

Evidence

Company Date What happened
OpenAI Jul 2026 Launched the GPT-5.6 line — Sol $5/$30, Terra $2.50/$15, Luna $1/$6 per 1M tokens — replacing the GPT-5.5/5.4 headline lineup. Luna at $1/$6 sat ABOVE the outgoing GPT-5.4 mini at $0.75/$4.50, i.e. the everyday tier launched more expensive.
Glean Aug 2026 Model Hub Usage table cut GPT-5.6 Terra 20% ($2.50 to $2.00 in, $15.00 to $12.00 out) and GPT-5.6 Luna 80% ($1.00 to $0.20 in, $6.00 to $1.20 out) — the first published price decreases on that card since Glean began disclosing dollar rates in July 2026. Luna was simultaneously reclassified Premium to Standard, moving it inside the included 100/user/week allowance instead of always billing FlexCredits.
Cursor Aug 2026 Other Models pool: Luna $1/$1.25/$0.10/$6 to $0.20/$0.25/$0.02/$1.20; Terra $2.50/$3.125/$0.25/$15 to $2/$2.50/$0.20/$12. Sol unchanged. Cursor's docs state the pool bills third-party API pricing at cost, so the cut flows straight through. Re-confirmed 2026-08-11.
Augment Code Aug 2026 Luna $1.00/$6.00 to $0.20/$1.20 (80%) and Terra $2.50/$15.00 to $2.00/$12.00 (20%) — the largest single-model rate move since Augment began publishing per-model pricing. The flat $100/month Business plan (up to 50 seats) was unchanged.
LiveKit Aug 2026 Per-minute denomination, same magnitudes: Luna $0.0040 to $0.0008/min (80%), Terra $0.0101 to $0.0081/min (20%), Sol held at $0.0203/min. The pricing page's default $0.0672/min voice-agent estimate did not move because it runs on Gemma 4 31B, which was not repriced.
Vercel Aug 2026 v0 Mini fell $1/$5 to $0.20/$1.20 per 1M — landing on Luna's exact new rate — and v0 Pro $3/$15 to $2/$10. Vercel's own note attributes it to the zero-markup AI Gateway passthrough. Because v0's plan cards never show token rates, the repricing is invisible on the pricing page.
Netlify Aug 2026 Same mechanic, different upstream: Gemini 3.6 Flash halved ($1.50/$7.50 to $0.75/$3.75 per 1M, cache write $0.15 to $0.07) and Gemini 3.7 Flash listed at the identical rate. Netlify converts provider rates at a fixed 180 credits per $1, so the cut passes through to credit consumption without any plan change (Free $0 / Personal $9 / Pro $20 unchanged).

Counterexamples

  • Perplexity · — — Launched the Gateway API on 2026-08-11 explicitly to STOP passing through: it hosts open-weight models at Perplexity's own set rates (kimi-k3 $3.00/$15.00, glm-5.2 $1.40/$4.40), inverting its Agent API's zero-markup resale. Three days later (2026-08-14) it also raised Agent API fetch_url 100% back to $0.0005, undoing its own 2026-07-29 cut.
  • SambaNova · — — Hosts its own silicon and sets its own rates: on 2026-08-11 it REVERSED its 2026-07-23 gemma-4-31B-it cut, taking the model from $0.22/$0.59 back up to $0.38/$1.15 (+73% input, +95% output) — a move no passthrough vendor could make.
  • Sarvam AI · — — Moved upstream rates the opposite direction on its own docs surface: Sarvam-105B repriced roughly 7x higher (Rs 4 / Rs 2.5 / Rs 16 to Rs 29.28 / Rs 10.98 / Rs 73.2 per 1M) on 2026-08-14, while its own marketing pricing page still advertises the old rate.

Trivia

  • Three companies published the identical two percentage moves on the same calendar day — 2026-08-11 — without coordinating: Augment Code, LiveKit and Vercel all took GPT-5.6 Luna down 80% and Terra down 20%, seven days after Glean did the same, with Cursor's 08-06 move re-confirmed the same day. LiveKit denominates in minutes rather than tokens ($0.0040 to $0.0008/min) and still landed on exactly 80%.

  • Vercel's v0 Mini fell to $0.20/$1.20 per 1M — landing on GPT-5.6 Luna's exact new rate — and the repricing is invisible on Vercel's own pricing page, because v0's plan cards never show token rates at all.

  • The everyday tier launched more expensive than the model it replaced: GPT-5.6 Luna arrived on 2026-07-23 at $1/$6 per 1M against the outgoing GPT-5.4 mini at $0.75/$4.50. The 80% cut twelve days later was a correction to a launch price, and every reseller inherited both moves.

See all pricing trivia

For buyers

Ask which models in a platform's catalog are marked up and which are passed through at cost, and get the answer in writing — it decides whether the vendor can hold your rate when upstream moves, and whether a favourable cut reaches you at all. For a passthrough vendor, monitor the upstream provider's rate card rather than the reseller's: the reseller has no notice obligation for a change it did not make, and the move may not even be visible where you look (Vercel's v0 plan cards show no token rates, so an 80% cut landed with nothing on the pricing page). Then treat the reverse case as the real risk — an at-cost pipe transmits increases with exactly the fidelity it transmitted this cut, and Luna itself launched at $1/$6 above the GPT-5.4 mini it replaced at $0.75/$4.50. Finally, read the packaging alongside the rate: Glean simultaneously reclassified Luna from Premium to Standard, moving it inside the included 100/user/week allowance instead of always billing FlexCredits, so the effective change was larger than the percentage.

For vendors

Running zero-markup resale means shipping a gateway that rates third-party tokens at cost and a published conversion buyers can audit — Netlify's fixed 180 credits per $1, Cursor's documented at-cost Other Models pool, Vercel's zero-markup AI Gateway. The payoff is that a cheaper upstream reaches your customer automatically with no repricing work; the cost is that you have no model margin and no ability to hold a rate, in either direction. If you take margin instead, you get independence and you should say so on the pricing page: Perplexity launched a Gateway API on 2026-08-11 specifically to stop passing through, hosting open-weight models at its own set rates (kimi-k3 $3.00/$15.00, glm-5.2 $1.40/$4.40) while its Agent API keeps reselling at cost, and SambaNova — on its own silicon — reversed a cut the same day. Either policy is defensible; what isn't is leaving the buyer to guess. And if you pass through, ship a changelog anyway, because your rate card now moves without you.

Outlook — what to watch

Logged new on seven corpus companies, five of them moving inside eleven days, so the mechanic is documented rather than inferred — the vendors state the passthrough policy themselves. Expect it to spread as gateway layers standardise: Netlify's 08-14 Gemini pass-through is already the pattern running on a non-OpenAI upstream. The move to watch is the first transmitted increase; Perplexity's Agent API fetch_url went 100% back up to $0.0005 on 2026-08-14, three days after it launched a margin-taking Gateway and sixteen days after its own cut, which shows how fast direction flips. It sharpens if another upstream move produces a same-week cluster across a different set of resellers. It weakens or is falsified if a zero-markup reseller holds a rate after its upstream moves — buffering the change behind a notice period — or if the same synchrony shows up among vendors that do take model margin.

Bottom line

The corpus now holds two populations that look identical on a rate card: vendors whose model prices are decisions, and vendors whose model prices are consequences. Only the second group moves in lockstep, and between 2026-08-04 and 2026-08-11 five of them published the same two percentage moves without making a pricing decision. A change feed that counts vendor pricing events will count five where one was made.

FAQ

How can I tell whether a platform marks up model tokens or passes them through at cost?

Look for a stated policy first — Cursor's docs say its Other Models pool bills third-party API pricing at cost, Vercel attributes v0's rates to its zero-markup AI Gateway, and Netlify publishes a fixed 180-credits-per-dollar conversion. Absent a policy, the tells are behavioural: the listed rate matches the provider's public rate to the cent (Vercel's v0 Mini landed on GPT-5.6 Luna's exact new $0.20/$1.20), the change lands the same week as other vendors', and no announcement accompanies it. It matters because a marked-up vendor can hold your rate when upstream moves — and can also keep a cut instead of passing it on.

Why did my AI coding bill drop without any announcement?

Because your vendor probably didn't make a decision. Between 2026-08-04 and 2026-08-11, Glean, Cursor, Augment Code, LiveKit and Vercel each cut GPT-5.6 Luna 80% and GPT-5.6 Terra 20% — a single upstream rate-card move flowing through five at-cost billing pipes. There was nothing to announce, and no notice obligation, because none of the five set those prices.

Does at-cost passthrough transmit price increases too?

Yes, with exactly the same fidelity. GPT-5.6 Luna launched on 2026-07-23 at $1/$6 per 1M — above the outgoing GPT-5.4 mini at $0.75/$4.50 — so every reseller inherited an increase before it inherited the 80% correction twelve days later. Perplexity shows how fast direction can flip: it raised its Agent API fetch_url 100% back to $0.0005 on 2026-08-14, undoing its own 2026-07-29 cut. Zero markup is not a discount guarantee; it is an abdication of rate control in both directions.

If prices pass through, what should I actually monitor?

The upstream provider's rate card, not the reseller's. Then normalise for denomination, because the same move arrives in different units: LiveKit quotes GPT-5.6 Luna per minute ($0.0040 to $0.0008/min) and still lands on exactly 80%, Netlify converts to credits at 180 per dollar, and Vercel's v0 plan cards never display token rates at all. Also check what is not repriced — LiveKit's headline $0.0672/min voice-agent estimate didn't move because it runs on Gemma 4 31B, which wasn't part of the cut.

All trends