Zero-markup resale transmits a model cut in days
Between 2026-08-04 and 2026-08-11, five corpus companies — Glean, Cursor, Augment Code, LiveKit and Vercel — repriced the same two OpenAI models by exactly the same two percentages: GPT-5.6 Luna down 80%, GPT-5.6 Terra down 20%. None of them announced a price change, because none of them made one: each already billed third-party model tokens at cost, so an upstream rate-card move passed straight through. The synchrony is the finding, not the discount — one decision surfaced as five separate 'price change' events, on four dates, in three different denominations.
What's happening — and why
What's happening: OpenAI launched the GPT-5.6 line on 2026-07-23 — Sol $5/$30, Terra $2.50/$15, Luna $1/$6 per 1M tokens. Glean's Model Hub table then cut Luna 80% ($1.00 to $0.20 in, $6.00 to $1.20 out) and Terra 20% ($2.50 to $2.00, $15.00 to $12.00) on 08-04. Cursor matched on 08-06, and Augment Code, LiveKit and Vercel all matched on 08-11. Netlify ran the identical mechanic on a different upstream three days later, halving Gemini 3.6 Flash ($1.50/$7.50 to $0.75/$3.75 per 1M) at a fixed 180-credits-per-dollar conversion, with no plan change at all (Free $0 / Personal $9 / Pro $20 unchanged).
Why: each of these vendors had already published an at-cost passthrough policy before the window opened. Cursor's docs state that its Other Models pool bills third-party API pricing at cost; Vercel attributes its move to the zero-markup AI Gateway; Netlify converts provider rates at a fixed 180 credits per $1. Having ceded pricing authority for third-party tokens, they don't decide these numbers — they inherit them. Their model prices are consequences, not decisions.
What makes it hard to see: the denomination changes at each hop. Glean, Cursor and Augment quote per 1M tokens; LiveKit quotes per minute ($0.0040 to $0.0008/min) and still lands on exactly 80%; Netlify quotes in credits. Vercel's v0 Mini fell from $1/$5 to $0.20/$1.20 — landing on Luna's exact new rate — yet the repricing is invisible on Vercel's own pricing page, because v0's plan cards never show token rates at all. And the direction of the underlying move matters: Luna launched at $1/$6, above the outgoing GPT-5.4 mini at $0.75/$4.50, so the 80% cut twelve days later was a correction to a launch price. Every reseller inherited both moves.
How it works
Evidence over time
7 supporting · 0 counter — hover or tap a point for detail, click to jump to the row.
Evidence
| Company | Date | What happened |
|---|---|---|
| OpenAI | Jul 2026 | Launched the GPT-5.6 line — Sol $5/$30, Terra $2.50/$15, Luna $1/$6 per 1M tokens — replacing the GPT-5.5/5.4 headline lineup. Luna at $1/$6 sat ABOVE the outgoing GPT-5.4 mini at $0.75/$4.50, i.e. the everyday tier launched more expensive. |
| Glean | Aug 2026 | Model Hub Usage table cut GPT-5.6 Terra 20% ($2.50 to $2.00 in, $15.00 to $12.00 out) and GPT-5.6 Luna 80% ($1.00 to $0.20 in, $6.00 to $1.20 out) — the first published price decreases on that card since Glean began disclosing dollar rates in July 2026. Luna was simultaneously reclassified Premium to Standard, moving it inside the included 100/user/week allowance instead of always billing FlexCredits. |
| Cursor | Aug 2026 | Other Models pool: Luna $1/$1.25/$0.10/$6 to $0.20/$0.25/$0.02/$1.20; Terra $2.50/$3.125/$0.25/$15 to $2/$2.50/$0.20/$12. Sol unchanged. Cursor's docs state the pool bills third-party API pricing at cost, so the cut flows straight through. Re-confirmed 2026-08-11. |
| Augment Code | Aug 2026 | Luna $1.00/$6.00 to $0.20/$1.20 (80%) and Terra $2.50/$15.00 to $2.00/$12.00 (20%) — the largest single-model rate move since Augment began publishing per-model pricing. The flat $100/month Business plan (up to 50 seats) was unchanged. |
| LiveKit | Aug 2026 | Per-minute denomination, same magnitudes: Luna $0.0040 to $0.0008/min (80%), Terra $0.0101 to $0.0081/min (20%), Sol held at $0.0203/min. The pricing page's default $0.0672/min voice-agent estimate did not move because it runs on Gemma 4 31B, which was not repriced. |
| Vercel | Aug 2026 | v0 Mini fell $1/$5 to $0.20/$1.20 per 1M — landing on Luna's exact new rate — and v0 Pro $3/$15 to $2/$10. Vercel's own note attributes it to the zero-markup AI Gateway passthrough. Because v0's plan cards never show token rates, the repricing is invisible on the pricing page. |
| Netlify | Aug 2026 | Same mechanic, different upstream: Gemini 3.6 Flash halved ($1.50/$7.50 to $0.75/$3.75 per 1M, cache write $0.15 to $0.07) and Gemini 3.7 Flash listed at the identical rate. Netlify converts provider rates at a fixed 180 credits per $1, so the cut passes through to credit consumption without any plan change (Free $0 / Personal $9 / Pro $20 unchanged). |
Counterexamples
- Perplexity · — — Launched the Gateway API on 2026-08-11 explicitly to STOP passing through: it hosts open-weight models at Perplexity's own set rates (kimi-k3 $3.00/$15.00, glm-5.2 $1.40/$4.40), inverting its Agent API's zero-markup resale. Three days later (2026-08-14) it also raised Agent API fetch_url 100% back to $0.0005, undoing its own 2026-07-29 cut.
- SambaNova · — — Hosts its own silicon and sets its own rates: on 2026-08-11 it REVERSED its 2026-07-23 gemma-4-31B-it cut, taking the model from $0.22/$0.59 back up to $0.38/$1.15 (+73% input, +95% output) — a move no passthrough vendor could make.
- Sarvam AI · — — Moved upstream rates the opposite direction on its own docs surface: Sarvam-105B repriced roughly 7x higher (Rs 4 / Rs 2.5 / Rs 16 to Rs 29.28 / Rs 10.98 / Rs 73.2 per 1M) on 2026-08-14, while its own marketing pricing page still advertises the old rate.
Trivia
-
Three companies published the identical two percentage moves on the same calendar day — 2026-08-11 — without coordinating: Augment Code, LiveKit and Vercel all took GPT-5.6 Luna down 80% and Terra down 20%, seven days after Glean did the same, with Cursor's 08-06 move re-confirmed the same day. LiveKit denominates in minutes rather than tokens ($0.0040 to $0.0008/min) and still landed on exactly 80%.
-
Vercel's v0 Mini fell to $0.20/$1.20 per 1M — landing on GPT-5.6 Luna's exact new rate — and the repricing is invisible on Vercel's own pricing page, because v0's plan cards never show token rates at all.
-
The everyday tier launched more expensive than the model it replaced: GPT-5.6 Luna arrived on 2026-07-23 at $1/$6 per 1M against the outgoing GPT-5.4 mini at $0.75/$4.50. The 80% cut twelve days later was a correction to a launch price, and every reseller inherited both moves.
For buyers
Ask which models in a platform's catalog are marked up and which are passed through at cost, and get the answer in writing — it decides whether the vendor can hold your rate when upstream moves, and whether a favourable cut reaches you at all. For a passthrough vendor, monitor the upstream provider's rate card rather than the reseller's: the reseller has no notice obligation for a change it did not make, and the move may not even be visible where you look (Vercel's v0 plan cards show no token rates, so an 80% cut landed with nothing on the pricing page). Then treat the reverse case as the real risk — an at-cost pipe transmits increases with exactly the fidelity it transmitted this cut, and Luna itself launched at $1/$6 above the GPT-5.4 mini it replaced at $0.75/$4.50. Finally, read the packaging alongside the rate: Glean simultaneously reclassified Luna from Premium to Standard, moving it inside the included 100/user/week allowance instead of always billing FlexCredits, so the effective change was larger than the percentage.
For vendors
Running zero-markup resale means shipping a gateway that rates third-party tokens at cost and a published conversion buyers can audit — Netlify's fixed 180 credits per $1, Cursor's documented at-cost Other Models pool, Vercel's zero-markup AI Gateway. The payoff is that a cheaper upstream reaches your customer automatically with no repricing work; the cost is that you have no model margin and no ability to hold a rate, in either direction. If you take margin instead, you get independence and you should say so on the pricing page: Perplexity launched a Gateway API on 2026-08-11 specifically to stop passing through, hosting open-weight models at its own set rates (kimi-k3 $3.00/$15.00, glm-5.2 $1.40/$4.40) while its Agent API keeps reselling at cost, and SambaNova — on its own silicon — reversed a cut the same day. Either policy is defensible; what isn't is leaving the buyer to guess. And if you pass through, ship a changelog anyway, because your rate card now moves without you.
Outlook — what to watch
Logged new on seven corpus companies, five of them moving inside eleven days, so the mechanic is documented rather than inferred — the vendors state the passthrough policy themselves. Expect it to spread as gateway layers standardise: Netlify's 08-14 Gemini pass-through is already the pattern running on a non-OpenAI upstream. The move to watch is the first transmitted increase; Perplexity's Agent API fetch_url went 100% back up to $0.0005 on 2026-08-14, three days after it launched a margin-taking Gateway and sixteen days after its own cut, which shows how fast direction flips. It sharpens if another upstream move produces a same-week cluster across a different set of resellers. It weakens or is falsified if a zero-markup reseller holds a rate after its upstream moves — buffering the change behind a notice period — or if the same synchrony shows up among vendors that do take model margin.
Bottom line
The corpus now holds two populations that look identical on a rate card: vendors whose model prices are decisions, and vendors whose model prices are consequences. Only the second group moves in lockstep, and between 2026-08-04 and 2026-08-11 five of them published the same two percentage moves without making a pricing decision. A change feed that counts vendor pricing events will count five where one was made.
FAQ
How can I tell whether a platform marks up model tokens or passes them through at cost?
Look for a stated policy first — Cursor's docs say its Other Models pool bills third-party API pricing at cost, Vercel attributes v0's rates to its zero-markup AI Gateway, and Netlify publishes a fixed 180-credits-per-dollar conversion. Absent a policy, the tells are behavioural: the listed rate matches the provider's public rate to the cent (Vercel's v0 Mini landed on GPT-5.6 Luna's exact new $0.20/$1.20), the change lands the same week as other vendors', and no announcement accompanies it. It matters because a marked-up vendor can hold your rate when upstream moves — and can also keep a cut instead of passing it on.
Why did my AI coding bill drop without any announcement?
Because your vendor probably didn't make a decision. Between 2026-08-04 and 2026-08-11, Glean, Cursor, Augment Code, LiveKit and Vercel each cut GPT-5.6 Luna 80% and GPT-5.6 Terra 20% — a single upstream rate-card move flowing through five at-cost billing pipes. There was nothing to announce, and no notice obligation, because none of the five set those prices.
Does at-cost passthrough transmit price increases too?
Yes, with exactly the same fidelity. GPT-5.6 Luna launched on 2026-07-23 at $1/$6 per 1M — above the outgoing GPT-5.4 mini at $0.75/$4.50 — so every reseller inherited an increase before it inherited the 80% correction twelve days later. Perplexity shows how fast direction can flip: it raised its Agent API fetch_url 100% back to $0.0005 on 2026-08-14, undoing its own 2026-07-29 cut. Zero markup is not a discount guarantee; it is an abdication of rate control in both directions.
If prices pass through, what should I actually monitor?
The upstream provider's rate card, not the reseller's. Then normalise for denomination, because the same move arrives in different units: LiveKit quotes GPT-5.6 Luna per minute ($0.0040 to $0.0008/min) and still lands on exactly 80%, Netlify converts to credits at 180 per dollar, and Vercel's v0 plan cards never display token rates at all. Also check what is not repriced — LiveKit's headline $0.0672/min voice-agent estimate didn't move because it runs on Gemma 4 31B, which wasn't part of the cut.