Unit proliferation at the infra layer — with the corpus vocabulary now saturating and consolidation a live countermove
Infrastructure and platform vendors keep adding metered dimensions as they ship features — Vercel now meters eight — while consumer apps consolidate to a single credit or seat. The metering surface is fragmenting at the infra layer and consolidating at the app layer.
What's happening — and why
What's happening: infrastructure vendors keep adding new things they bill for — compute, bandwidth, storage, requests, function invocations — so one platform can meter eight or more units at once. Consumer apps move the opposite way, collapsing everything into a single credit or seat.
Why: at the infra layer each capability has a distinct cost driver, and buyers want the price to track the resource they actually control, so granular metering wins. At the app layer users just want a simple, predictable bill — so the same complexity gets hidden behind one number.
How it works
Evidence over time
22 supporting · 11 counter — hover or tap a point for detail, click to jump to the row.
Evidence
| Company | Date | What happened |
|---|---|---|
| Vercel | Jun 2025 | Active CPU billing added; now meters eight distinct units — seats, bandwidth-GB, edge requests, function invocations, CPU, memory-GB-hours, tokens, builds. |
| Vercel | Apr 2024 | Granular metering + Pro add-ons split billing into more dimensions vs the 2022 bundled-allocation model. |
| Modal | Sep 2025 | Per-second GPU/CPU/memory plus storage-GB and function invocations — five metered units. |
| ElevenLabs | May 2025 | PAYG shift spread metering across characters, credits, hours, minutes and seats. |
| Deepgram | Dec 2024 | Split into per-product STT / TTS / Audio-Intelligence tables billed per-minute, per-character and per-token. |
| RunPod | Feb 2026 | Pods / Serverless / Instant Clusters relabelled into separate metered product lines with reservations. |
| LangChain (LangSmith) | Jun 2026 | LLMops platform with 6 distinct units: seats, traces, workflow-executions, units, cpu-hours, storage-gb — an observability and evaluation platform accreting units as it ships new capabilities (traces for monitoring, storage for datasets, cpu-hours for evaluation runs). |
| LiveKit | Jun 2026 | Real-time agent infra with 5 metered units: agent-session minutes, WebRTC media minutes, inference credits, telephony minutes, and bandwidth-gb — each representing a distinct cost driver in an AI voice-agent stack. |
| Weaviate | Jun 2026 | Vector DB with 4 units: vectors-indexed, tokens (for AI-powered operations), api-calls, and storage-gb — multi-dimensional infra metering extending from GPU cloud into managed data infrastructure. |
| Anthropic | Jun 2026 | Added two new meters alongside per-token billing in one release: Claude Managed Agents bills $0.08 per session-hour of runtime (a session-runtime dimension, replacing Code Execution container-hours for agent sessions), and a research-preview Fast mode adds a premium-speed price band ($10/$50 on Opus 4.8, $30/$150 on Opus 4.6/4.7). A frontier model API accreting non-token units as it ships agent features. |
| Vercel | Jul 2026 | Surfaced four new BETA metered SKUs on one pricing page on top of its existing eight platform dimensions: Vercel Agent ($0.25/1M tokens), Vercel Services ($0.50/1M requests), Vercel Queues ($0.60/1M API operations), and Container Registry Image Storage ($0.10/GB-month). New units: api-operations and image-storage-gb-month — a platform accreting metered dimensions as it ships primitives. |
| AssemblyAI | Jul 2026 | Broke Guardrails into four separately-priced sub-meters (Profanity $0.01/hr, PII Audio Redaction $0.05/hr, PII Text Redaction $0.08/hr, Content Moderation $0.15/hr) and added Medical Mode ($0.15/hr) and Keyterms Prompting ($0.05/hr) add-ons across six product tabs — per-feature per-hour meters multiplying inside a speech-AI platform. |
| RunPod | Jun 2026 | Its new per-request Public Endpoints layer meters by request, token, OR character depending on the model — $0.05/1000 chars (Whisper V3), $0.02/megapixel (FLUX.1), $1.20/request (SORA 2 Pro video) — layered on top of per-hour Pods, per-second Serverless, and multi-node Clusters. A single vendor now spans per-hour, per-second, per-request, per-token, per-character, and per-megapixel meters. |
| Braintrust | Jul 2026 | Proliferation by converting a feature gate into a meter — the cheapest way to add a unit. The comparison row 'Default retention: 14 days (Starter) / 30 days (Pro) / Custom (Enterprise)' was renamed 'Included data retention' and Pro now reads '30-day retention, then $0.50 per GB per month', with a matching '+ $0.50/GB/mo' badge on the plan card. Tier fees unchanged. An LLM-observability vendor going 3 meters to 4 one day after its closest competitor went 7 to 2. Source: changes/braintrust-2026-07-22-packaging.md. |
| Bland AI | Jul 2026 | Unbundling as proliferation: Bland dropped telephony from its all-in per-minute rate, which turns one headline meter into two billable components. The inverse of Lovable's same-week merge of three balances into one — two vendors resolving the same question in opposite directions inside a single capture batch. Source: changes/bland-ai-2026-07-21-packaging.md. |
| Modal | Jul 2026 | A per-second compute vendor adding a completely different metering BASIS rather than another dimension of the same one: Shared API launched with token-based pricing alongside Modal's existing per-second GPU/CPU/memory billing, on top of the 2026-07-14 additions (a separately-metered Sandbox+Notebooks tier at roughly 3x standard function rates, plus published region 1.5-1.75x and non-preemptible 3x modifiers). Per-second, per-token, per-GB and two multipliers on one rate card. |
| Glean | Jul 2026 | The proliferation pattern reaching enterprise search: Glean expanded its Enterprise Flex rate card to 13 metered capabilities — a seat-and-platform vendor publishing a 13-line consumption menu beside the license. Source: changes/glean-2026-07-22-packaging.md. |
| Composio | Jul 2026 | Six-dimension usage pricing announced effective 2026-08-15, alongside retiring the $229 tier and adding a $599 Business tier — an agent-tooling vendor stating its meter count as a headline feature of the new packaging. Source: changes/composio-2026-07-23-packaging.md. |
| Vercel | Aug 2026 | Two new meters inside a single add-on. AI Gateway's Trace Drains forwards an OpenTelemetry trace of every request to a customer's own observability tooling and bills on BOTH the number of trace events delivered ($0.05 per 1,000 AI Gateway Traces) and the volume of trace data transferred ($0.50 per 1 GB of egress), on Pro and Enterprise. It joins the existing Custom Reporting, provider-allowlist and Zero Data Retention surcharges, and charges post against the team's general Drains usage rather than depleting AI Gateway Credits — so it is a new meter on a new balance. The same 2026-08-11 batch published Custom Environments overage at $50/month per additional pack of five (max 16 Pro, 22 Enterprise). |
| Fireworks AI | Aug 2026 | A new pricing AXIS that creates no new billing unit — the shape that lets the corpus vocabulary freeze while rate cards get harder. Fireworks began pricing region-restricted dedicated deployments at a flat 1.5x the standard on-demand rate, gated behind a Contact Sales request, for GPUs pinned to US-only or Europe-only infrastructure. It is the first geographic-routing premium published for its dedicated GPU product and parallels the existing 10% US-only Serverless token premium at fifteen times the markup. In the same capture GB300 288GB joined the card at $18.00/hr, the highest published on-demand rate. |
| Glean | Aug 2026 | Meter proliferation by MODALITY inside existing SKUs. GPT Realtime 1.5 and GPT Realtime 2 had each been a single ambiguous row on Glean's Model Hub Usage table showing $5.00 per million input tokens, $0.50 cache read and no output rate at all. Each is now split into three modality-specific lines — Audio $32.00 in / $64.00 out, Text $4.00 in / $16.00-$24.00 out, Image $5.00 in — with GPT Realtime 2.1 added on the same structure, plus new rows for GPT-4o Mini TTS ($0.60 text in / $12.00 audio out) and Deepgram Nova-3 Multilingual ($0.0117/min output, 2 channels). Three published rows where there was one, and real-time audio token rates disclosed for the first time. No existing rate moved. |
| Perplexity AI | Aug 2026 | A fifth metered developer surface. The Gateway API — Perplexity-hosted open-weight models priced per token with no per-request fee — launched alongside, not inside, the existing Sonar API (tokens plus a per-request search fee), Search API ($5.00/1K requests), Agent API (third-party models at cost plus metered tool calls) and Embeddings API. The catalog went from 3 models to 5 in three days. It follows the per-session sandbox meter added to the Agent API on 2026-07-21 ($0.03 per container session, 20-minute billing window), Perplexity's first meter counting neither tokens nor requests. |
Counterexamples
- ZenRows · Aug 2026 — The second major INFRA consolidation, and the cleanest currency merge in the corpus — three billing BASES collapsed into one. ZenRows replaced cost-per-1,000-requests on the Universal Scraper API, per-GB billing on Residential Proxies and a $0.09-per-session-hour fee on the Scraping Browser with a single shared credits balance spanning four renamed primitives: Fetch, Extract (new, beta), Batch (new, beta) and Browser Sessions. Credit weights are fixed and published — 1 credit for a standard request, 5 for JavaScript rendering, 10 for Premium Proxies, 25 for both — and Residential Proxies is now labelled '(Legacy)' as a standalone add-on. Like LangSmith's seven-to-two collapse, the merge arrived inside a full ladder rebuild: Free became a permanent $0 plan with 5,000 credits (previously a 14-day trial), and the paid ladder went to Build $19-$39 / Launch $69-$129 / Growth $199-$399 / Scale $549-$999, three rungs each, with Enterprise above 12.5M credits/month.
- Diffbot · Aug 2026 — A vendor declining to proliferate, which is rarer than consolidating. Diffbot added a fifth product — a Web Search API — and priced it at 1 credit per query, the SAME rate as a full page extraction, folded into the existing 'all APIs included' entitlement on every tier including Free. No new unit, no new tier, and Free/Startup/Plus/Enterprise unchanged at $0/$299/$899/custom, with Crawl still the only feature gated to $899+. The relaunch also surfaced a previously undocumented mechanic: the Free plan does not auto-bill overage — exceeding 10,000 credits/month or the rate limit returns a 429 until reset — while paid plans bill pro rata.
- Character.ai · Feb 2026 — Single unit only (active users); monetises free users with ads rather than new meters.
- Suno · May 2026 — One credit unit across Free / Pro / Premier.
- Harvey · May 2026 — Seats only — no usage metering.
- Fathom · Jun 2026 — AI meeting notetaker with a single unit (seats) — consumer/prosumer tools resist unit proliferation.
- Superhuman · Jun 2026 — AI email client with seats-only billing — no usage dimensions added despite AI features.
- LangChain (LangSmith) · Jul 2026 — The sharpest counterexample in the trend's history, because it is an INFRA vendor consolidating — the layer the hypothesis says accretes. LangChain collapsed seven distinct LangSmith meters (per-1k traces, $0.005 per deployment run, uptime minutes at $0.0036/min production and $0.0007/min development, $0.05 per Fleet run, $1.50/LCU Engine, per-second dollar-denominated sandbox rates) into TWO normalized units: 1 LCU = $1.50 for work and compute, 1 LSU = $1.00 for traces and storage. Deployments now meter on resources consumed (0.045 LCU/vCPU-hr runtime compute, 0.177 LSU/vCPU-hr database compute); Fleet is LCU-metered with 5 LCU (Developer) or 25 LCU (Plus) included; seats held at $0/$39. This directly contradicts the trend's own earlier evidence entry for LangSmith (2026-06-09, '6 distinct units … accreting units as it ships new capabilities'). Read the price alongside it, though: base traces went to 0.005 LSU each, double the old $2.50/1k, and extended retention fell from 400 days to 180. Sources: changes/langchain-2026-07-21-packaging.md, changes/langsmith-2026-07-21-price-change.md.
- Lovable · Jul 2026 — Consolidation at the app layer, executed as a currency merge: three separate consumption balances — build credits, a dollar-denominated Lovable Cloud balance (hosting, database, storage, network, compute, realtime) and a dollar-denominated AI-gateway balance — collapsed into ONE credit balance, with remaining dollars converted at the plan's credit rate and monthly grants reissued as 20 Cloud + 4 AI credits on top of 5 daily build credits. Lovable's billing_units are now credits and nothing else. Source: changes/lovable-2026-07-21-packaging.md.
- Together AI · Jul 2026 — Consolidation inside an inference vendor's own catalog: Together unified its speech-to-text pricing in the same release that raised GPU cluster reserved rates — meters merging in one product line while prices rose in another, which is why counting meters alone tells you nothing about direction of price.
- Braintrust · Jul 2026 — The same vendor that ADDED a meter on 2026-07-22 consolidated eight days later, and reduced disclosure while doing it: the 'Topics credit' became a shared 'Model credits' pool, and the pricing page swapped its explicit per-tier token-rate line for a generic 'then token rates' plus a 'View detailed pricing' link. Fewer visible meters, less published rate detail — consolidation and transparency moving in opposite directions. Source: changes/braintrust-2026-07-30-packaging.md.
Trivia
-
Vercel meters eight distinct dimensions on its Pro plan (seats, bandwidth-GB, edge requests, function invocations, CPU, memory-GB-hours, tokens, builds) — the highest unit count for a single-vendor plan in the corpus — and added active CPU billing as its eighth unit in June 2025. The proliferation tracks directly with Vercel's product expansion: each new capability (serverless functions, edge compute, AI inference) introduced its own cost driver, and the pricing page grew to match.
-
The corpus has grown from roughly 20 distinct billing units (at 43 companies) to 34 (at 158 companies), adding units like actions (Rox), workflow-executions (Upstash), documents (Nomic), and browser-hours (Browserbase) — each representing a new product capability that has no natural mapping to an existing unit. The unit count grows with the product surface, not with the number of companies, which is why infra vendors lead the proliferation.
-
Consumer apps and the unit-proliferation trend diverge sharply: Character.ai (single unit: active users), Suno (single unit: credits), and Harvey (single unit: seats) prove that the app layer actively resists proliferation. The divergence is a deliberate UX choice — every additional billing dimension adds cognitive overhead for a buyer who just wants to know their monthly cost, so app-layer products pay a real adoption cost to add metered dimensions that infra buyers accept as normal.
-
Anthropic added two non-token meters in a single June 2026 release — a $0.08/session-hour runtime charge for Claude Managed Agents and a latency-priced Fast mode SKU — making a frontier model API, not a data/compute infra vendor, the place unit proliferation now shows up. The session-hour meter is metered to the millisecond and accrues only while a session is 'running', so the same Claude API that once billed purely per token now charges by tokens, cached tokens, batch tokens, per-tool actions, container/session time, and output speed at once.
-
RunPod (2026-06-30) is the corpus's most unit-diverse single vendor: its pricing page bills the same account by per-hour Pods, per-second Serverless workers, multi-node Cluster commits, AND per-request Public Endpoints that themselves meter by request ($1.20/request SORA 2 Pro), token, character ($0.05/1000 chars Whisper), OR megapixel ($0.02/megapixel FLUX) depending on the model — six billing bases surfaced on one page, the clearest single-vendor illustration of infra-layer metering fragmentation.
-
Vercel added four BETA metered SKUs in one 2026-07-06 release (Agent, Services, Queues, Container Registry) on top of its eight existing platform dimensions — a net +4 meters in a single capture, while the app-layer counterexamples (Character.ai, Suno, Harvey) still hold to a single unit. The infra-accretes / app-consolidates divergence widened rather than converged this cycle.
-
Two direct LLM-observability competitors moved in opposite directions ONE DAY apart. LangChain collapsed seven LangSmith meters into two normalized units on 2026-07-21 (1 LCU = $1.50, 1 LSU = $1.00, replacing per-1k traces, per-deployment-run, uptime minutes, per-Fleet-run, Engine LCUs and per-second dollar sandbox rates). On 2026-07-22 Braintrust went from three meters to four, putting $0.50/GB/mo on data retention that had been a plain feature gate. Meter count is not a one-way ratchet even inside one segment.
-
Consolidation is not simplification. LangSmith's collapse from seven meters to two doubled the price of the thing most customers actually buy — base traces went to 0.005 LSU each, "double the old $2.50/1k" — and cut extended retention from 400 days to 180. The rate card got shorter and the comparison got harder, because two abstract units (LCU, LSU) now sit between the buyer and the dollar.
-
The corpus's metering vocabulary did not just slow down — it stopped. 65 distinct billing units at 380 companies, the same 65 as at 353: twenty-seven companies were added and contributed ZERO new units, a marginal rate of 0.00 against 0.13 at the prior step and roughly 0.47 in the early passes (20 units at 43 companies). The tail also CONTRACTED, from 15 single-company units to 13, because two former one-offs found adopters — `prompts-tracked` is now on 4 companies and `sites` on 8 — and no new one-off appeared.
-
Vendors kept making rate cards harder without inventing a single new unit, by adding AXES instead. Four landed in seventeen days: Fireworks' flat 1.5x for US- or Europe-pinned dedicated GPUs (2026-08-11), DeepSeek's peak/off-peak split (2026-08-14), Sarvam's `editor_flow: true` request flag that exactly doubles a Dubbing rate already multiplied by target-language count (2026-08-15), and Glean's split of GPT Realtime into three modality lines per model — Audio $32.00/$64.00, Text $4.00/$16.00-24.00, Image $5.00 in (2026-08-11). Counting billing units now understates complexity.
-
ZenRows ran the LangSmith play at the scraping layer and went further: on 2026-08-04 it replaced THREE dollar-denominated billing bases — cost-per-1,000-requests on the Scraper API, per-GB on Residential Proxies, and $0.09 per session-hour on the Scraping Browser — with one shared credits currency across four primitives at published fixed weights (1 credit standard, 5 with JavaScript rendering, 10 with Premium Proxies, 25 for both), demoting Residential Proxies to "(Legacy)". Like LangSmith's, the merge shipped inside a full ladder rebuild — a shorter meter list and a harder comparison.
For buyers
More units means more places the bill can surprise you. For infra vendors, ask which dimensions are new this year and which dominate a typical invoice; for app vendors, the single unit is simpler but often hides the same complexity behind a credit.
For vendors
A new metered dimension only pays off if you can attribute cost to it cleanly and explain it on the invoice. Each unit raises billing-system and support load — proliferate at the infra layer where customers model cost, consolidate at the app layer where they don't.
Outlook — what to watch
Infra metering will keep fanning out as new cost drivers appear (active CPU, GPU-seconds, agent steps). The countervailing force is tooling: as cost-attribution and forecasting improve, expect a swing back toward a few 'headline' units with the rest folded into bundles. Watch for vendors that lead with one number and meter the rest underneath.
Bottom line
The corpus exposes 34 distinct billing units. They fragment at the infrastructure layer (Vercel meters eight) and consolidate to a single credit or seat at the app layer.
FAQ
What is a 'billing unit' in AI pricing?
The thing you're charged for — tokens, seats, GPU-hours, API calls, credits, edge requests, and so on. The corpus uses 34 distinct ones; a single infra vendor can meter many at once.
Why do infrastructure vendors meter so many units?
Each new capability has a distinct cost driver (compute, bandwidth, storage, invocations), and infra buyers want to model cost precisely. Vercel, Modal and RunPod expose multiple units for exactly that reason.
Is more granular metering better?
For infra buyers, yes — it maps price to the cost they control. For consumer apps it's noise, which is why those vendors collapse everything into one credit or seat.