Per-action pricing for agent tools
Search- and agent-native vendors now bill discrete tool actions — a web search, a code run, a research task — as their own unit, on top of (or instead of) tokens. Six corpus companies do it, including a general-model API (Anthropic) and browser-agent infra (Browserbase); the billable unit is shifting from the token to the action.
What's happening — and why
What's happening: when an AI agent uses a tool — searches the web, opens a page, runs code, kicks off deep research — some vendors now charge for that action directly, per call, separate from the tokens the model generates.
Why: agent workloads make tool calls a real, attributable cost (an API hit, compute, a third-party fee) that doesn't map cleanly to tokens. Pricing the action lets vendors recover it and signal it to buyers. It's early — mostly search/agent-native APIs — but it points at a future where the unit of value is the task an agent completes, not the text it emits.
How it works
Evidence over time
18 supporting · 8 counter — hover or tap a point for detail, click to jump to the row.
Evidence
| Company | Date | What happened |
|---|---|---|
| Groq | Jan 2026 | Built-in tools priced per use — web search $5–$8/1K, website visits $1/1K, code execution $0.18/hour, billed on top of token rates. |
| Perplexity AI | Jan 2026 | API restructured around a Search API plus an Agentic Research tier — research priced per task, above raw per-call search. |
| You.com | Mar 2026 | Research API launched with effort tiers (lite $6.50 → exhaustive $300) — agentic research billed per call by effort, separate from token rates. |
| Exa | Apr 2026 | Per-endpoint cards with a dedicated Agent endpoint; base Search at $7/1k, Deep Search $15/1k — each action billed per call. |
| Anthropic | May 2026 | Web search tool billed $10 per 1,000 searches and code execution at $0.05/hour per container — discrete per-action fees on top of tokens, on a general-model API. |
| Browserbase | Mar 2026 | Fetch API launched at $1/$4/$7 per 1k calls, adding per-call Search/Fetch metering on top of browser-hours — action pricing on browser-agent infra. |
| Composio | Jul 2025 | Core meter renamed to 'tool calls' (after 'executions' 2024 and 'API calls' Jan 2025); premium tools (search APIs, E2B sandboxes, ML inference) bill at exactly 3x the standard per-call rate — $0.897 vs $0.299 per 1K overage calls on the $29 plan. |
| Zapier | Jun 2026 | Meters Agents by 'activities' and Chatbots by bot count, on top of platform tasks — agentic actions as a separate billable dimension from workflow steps. |
| Vapi | Jun 2026 | Voice agent platform bills $0.05/min hosting plus per-tool passthrough (STT, LLM, TTS, telephony at cost) — a multi-action billing model where each voice agent call spans multiple tool invocations priced separately. |
| Anthropic | Jun 2026 | Claude Managed Agents bill $0.08 per session-hour of runtime (millisecond-metered, running-only) on top of tokens — a session-time meter for agent workloads, replacing the Code Execution container-hour model. Adds a wall-clock dimension to the per-action billing already in place ($10/1k web searches, code execution per container). |
| Perplexity AI | Jul 2026 | A new shape: the per-session meter with an explicit billing window. Perplexity added `sandbox` as a fifth Agent API tool — 'an isolated container for executing code during an Agent API request' — at $0.03 per container session, and the docs are precise that 'a session covers up to 20 minutes of active use for billing purposes — this is the billing window, not a runtime cap.' It is Perplexity's first meter that counts neither tokens nor requests. SDK search queries issued from inside the sandbox bill separately. Third-party model resale also widened from 4 providers to 7 (adding Z.AI, Moonshot AI, NVIDIA) at direct provider rates with no markup. |
| Perplexity AI | Jul 2026 | The corpus's first deflation of the action unit: web_search halved from $0.005 to $0.0025 per invocation ($2.50 per 1,000) and fetch_url from $0.0005 to $0.00025, with the same cut flowing through to SDK search queries run inside the sandbox tool (now documented as 'same as web_search'). Corroborated by the page's own worked examples. people_search and finance_search ($0.005 each) and the $0.03 sandbox session fee were unchanged, as were Sonar token/request pricing, the $5.00-per-1,000 Search API, and Embeddings — so the cut is targeted at the two highest-volume agent tools. |
| Composio | Jul 2026 | The first large action-unit HIKE, and a split into six meters. Effective 2026-08-15 the single tool-call meter becomes six metered dimensions (tool calls, trigger events, LLM tokens, premium-tools dollar credit, sandbox GB-hr, filesystem GB), tool-call overage rises from $0.299 per 1K to $4 per 1K ($3 via Sessions), and the $29 Pro plan's included tool calls fall from 200,000 to 50,000. Plans re-tiered from Totally Free / Ridiculously Cheap $29 / Serious Business $229 to Free / Pro $29 / Business $599 / Enterprise, with existing customers grandfathered through 2026-12-31. |
| MiniMax | Jul 2026 | New adopter of per-call tool billing on an inference API: MiniMax added an MCP billing surface (API-vlm at $0.01 per call, cut from $0.06) and a server-tool line (web_search at $0.01 per call), alongside Music 3.0 at $0.15 per up-to-5-minute track. Flagship token rates held at $0.30/$1.20 per 1M for M3 and M2.7 ($0.60/$2.40 for M2.7-highspeed, cache reads $0.06/M), with M1/M2/M2.5 moved to Legacy — so the tool meter arrived without a token change. |
| Zhipu AI | Jul 2026 | New adopter: the GLM API card added per-use web search plus per-image, per-video and per-agent SKUs alongside the token rates (GLM-5.2/5.1 $1.4/$4.4, GLM-5 $1/$3.2, GLM-5-Turbo $1.2/$4.0, GLM-4.5-X $2.2/$8.9 per 1M). Discrete non-token SKUs appearing on a Chinese frontier lab's rate card in the same release that raised the GLM Coding Plan to $18/$72/$160 a month. |
| Glean | Jul 2026 | Surface-dependent action pricing, the most novel mechanic of the window: a Basic Search Query costs 0 FlexCredits inside Glean's own apps but 1 FlexCredit when it arrives through the Client API — the same action, priced by the door it comes through. The 13-line Enterprise Flex rate card also mixes per-query, per-session and per-minute meters: Adaptive Reasoning ~11/~38 credits (standard) and ~26/~83 (premium), Voice Session ~3/~26, Meeting Notes ~9/~18 per MINUTE, Code Writer ~9/~32, Image Generation ~7/~9, Slide Generation ~45/~142, Deep Research ~33/~144, Agent Run ~7/~114, Advanced Agent Runs ~150/~450. No dollar figure is published anywhere. |
| Perplexity AI | Aug 2026 | The corpus's first action-unit ROUND TRIP, 16 days wide. fetch_url went back to $0.0005 per invocation — exactly its pre-2026-07-29 rate — undoing the halving to $0.00025 logged as this trend's first-ever action-unit cut. web_search stayed at the cut price ($0.0025), people_search and finance_search stayed at $0.005, and the $0.03 sandbox per-session fee did not move. The reversal is confirmed independently of the rate table: the docs' worked 'Agent API Research Preset' example (2,000 input + 1,000 output tokens, 1 web_search, 1 fetch_url) totalled $0.00675 on 2026-08-11 and totals $0.007 now — a delta that only reconciles at $0.0005. Volatility in this unit is not just fast, it is non-monotonic: the same SKU halved and restored inside one synthesis cycle. |
| Groq | Aug 2026 | The category's founding vendor stopped publishing the category's rate card. groq.com/pricing now 308-redirects to the Groq homepage, which carries no pricing content and no 'Pricing' nav link; the homepage instead leads with an 'ANNOUNCING OUR $650 MILLION FUNDRAISE' banner and an LPX platform repositioning. No rate moved — every figure still in the GroqCloud developer docs is identical to the 2026-07-21 capture (Llama 3.1 8B $0.05/$0.08, Llama 3.3 70B $0.59/$0.79, GPT OSS 120B $0.15/$0.60, GPT OSS 20B $0.075/$0.30, Whisper $0.111/$0.04 per hour, Orpheus $22/$40 per 1M characters). But the three built-in agentic tools this trend was founded on — web search, website visits, code execution — now have no live public source anywhere, because their tool docs defer to the dead pricing page rather than printing a rate. Groq supplied this trend's first evidence entry (2026-01-22, web search $5-$8/1K, website visits $1/1K, code execution $0.18/hour) and pulled a priced tool seven days after launching it (2026-07-21). It has now retracted the publication as well as the SKU. |
Counterexamples
- DeepSeek · Aug 2026 — The holdout cohort hardened rather than narrowed. DeepSeek repriced its API for the first time in the tracked range and did it entirely inside the token meter, inventing a new axis rather than adding a tool SKU: a 2026-08-11 footnote warned of an unspecified 'significant' increase, and on 2026-08-14 it published peak/off-peak billing effective 16:00 UTC on 2026-08-16 — peak windows 01:00-04:00 and 06:00-10:00 UTC, all other hours at exactly half the peak rate. V4-Flash cache-miss input moves from a flat $0.14 per 1M to $0.22 off-peak / $0.44 peak; V4-Pro from $0.435 to $0.66 / $1.32. Roughly 3-12x on peak line items and 1.5-2.5x off-peak, with the steepest relative jump on cache-hit input. A vendor that will build time-of-use pricing before it will price a tool call is a stronger counterexample than one that has simply not moved. Mistral and Cohere remain unchanged holdouts.
- OpenAI · May 2026 — Built-in tools fold into token usage rather than a discrete per-call line item.
- Google · Apr 2026 — Gemini grounding/tool use is billed within token and request usage, not as a separate per-action SKU.
- Groq · Jul 2026 — Added a Browser Automation built-in tool at $0.08/hour, joining its existing per-use tool card — web search ($5-$8 per 1,000 requests), website visits ($1 per 1,000), and code execution ($0.18/hour). Discrete agent tools metered by action or by hour on an LPU inference API.
- Google · Jul 2026 — AlphaEvolve on Vertex AI prices the agent as the base Gemini model rate PLUS an explicit 2x agent surcharge (3x all-in — e.g. Gemini 3.1 Pro $6/$36 vs $2/$12 base). The agent premium as a MULTIPLE of the underlying token rate rather than a flat per-action fee — a token-scaled agent charge, distinct from Google's own grounding/tool use which stays folded into token/request usage.
- Groq · Jul 2026 — The first WITHDRAWAL of a priced tool SKU in the corpus, and a direct retraction of evidence added at the last review: the Browser Automation built-in tool at $0.08/hour — launched 2026-07-14 — was pulled from the Built-In Tools (Compound) table on 2026-07-21, seven days later, the shortest priced-SKU life the corpus has recorded. The table reverted to four tools (basic search, advanced search, visit website, code execution). The rest of the card held (web search $5-$8 per 1,000, website visits $1 per 1,000, code execution $0.18/hour), and retained model prices were stable — so the retraction was specific to the new tool, not a general repricing.
- OpenAI · Jul 2026 — Partial flip of this trend's oldest counterexample: alongside the GPT-5.6 line, OpenAI published Containers tool pricing — 1 GB for $0.03 and 64 GB for $1.92 per container, moving to per-20-minute-session billing on 2026-03-31. That is a discrete, non-token SKU with a session billing window, the same shape Perplexity's sandbox uses. OpenAI still folds most built-in tool use into token usage, so it remains a counterexample in the main, but no longer a pure one.
- Mistral AI · Jul 2026 — Ships agentic models with no separate tool SKU at all — the raise it made this window landed on token rates (Small 4 $0.1/$0.3 → $0.15/$0.6 per 1M; OCR doubled to $4/1K pages), not on a per-action tool line. Together with Cohere and DeepSeek, Mistral is the remaining holdout cohort: agentic capability priced entirely inside the token meter.
Trivia
-
Composio's billable unit changed identity three times in roughly twelve months — 'executions' (2024), 'API calls' plus authenticated users (January 2025), then 'tool calls' (July 2025) — and the final form prices premium tool calls at exactly 3.0x the standard rate ($0.897 vs $0.299 per 1K), the corpus's first premium-vs-standard tiering *inside* the action unit itself.
-
Groq was the first corpus vendor to price discrete agent tool calls as their own line item (January 2026), breaking the convention that web search was just a model capability bundled into token usage — its $5-$8 per 1,000 searches made the action unit legible as a cost driver separate from the conversation tokens around it.
-
You.com's effort-tier Research API (lite $6.50 → exhaustive $300 per call, launched March 2026) produces the widest single-parameter cost range in the corpus: a 46× spread from a single configuration knob. No other billing dimension in the corpus moves cost by that ratio within one API endpoint, making "effort" the most financially consequential agent parameter the corpus has documented.
-
Anthropic's May 2026 decision to price web search at $10 per 1,000 calls and code execution by the container-hour on its general-model API — rather than folding tool use into tokens as OpenAI and Google do — was the decisive event that graduated per-action pricing from a search-native curiosity to an established billing unit, because it meant a frontier general-purpose API now had discrete per-action SKUs.
-
Groq's Browser Automation tool holds the record for the shortest priced-SKU life in the corpus: launched at $0.08/hour on 2026-07-14 and pulled from the Built-In Tools (Compound) table on 2026-07-21, seven days later. It was logged as this trend's newest evidence at the 2026-07-15 review and had to be re-scored as a withdrawal at the next one — the tool rate card can retract as fast as it expands, even while the rest of Groq's card (web search $5–$8/1k, website visits $1/1k, code execution $0.18/hour) holds steady.
-
Perplexity added a meter and then halved two others inside nine days. On 2026-07-21 it shipped `sandbox` at $0.03 per container session — its first meter counting neither tokens nor requests, with the docs stressing that the 20 minutes is "the billing window, not a runtime cap". On 2026-07-29 it cut `web_search` from $0.005 to $0.0025 per invocation and `fetch_url` from $0.0005 to $0.00025, the corpus's first documented deflation of the action unit itself. The sandbox session fee did not move.
-
Composio is taking the action unit up 13x while cutting what is included: from 2026-08-15 its tool-call overage goes from $0.299 per 1K to $4 per 1K ($3 via Sessions), and the $29 Pro plan's included tool calls fall from 200,000 to 50,000 — a 75% entitlement cut at an unchanged price, on top of a 13x overage rate. The single tool-call meter it adopted at its July-2025 relaunch also splits into six dimensions (tool calls, trigger events, LLM tokens, premium credit, sandbox GB-hr, filesystem GB), with existing customers grandfathered only through 2026-12-31.
-
Glean (2026-07-22) prices the identical action differently depending on which door it comes through: a Basic Search Query costs 0 FlexCredits inside Glean's own apps and 1 FlexCredit via the Client API. Its 13-line rate card also mixes three meters in one product — per-query (Adaptive Reasoning ~11/~38 standard, ~26/~83 premium), per-session (Voice ~3/~26), and per-MINUTE (Meeting Notes ~9/~18) — making surface and duration billing axes alongside the action.
-
The trend's oldest counterexample partly flipped: OpenAI, cited here since 2026-05 for folding built-in tools into token usage, published Containers tool pricing on 2026-07-23 — 1 GB for $0.03 and 64 GB for $1.92 per container, moving to per-20-minute-session billing on 2026-03-31. That leaves Mistral, Cohere and DeepSeek as the vendors still shipping agentic models with no separate tool SKU at all.
For buyers
Where tools are billed per action, forecast on the agent's plan — how many searches, page visits and code runs per task — not just token volume. A verbose agent can rack up tool fees that dwarf its token cost; cap tool calls per run where the vendor allows it.
For vendors
To price actions you need per-tool metering and a clear line item for each (search, visit, execution, research), plus a free monthly allotment to ease adoption. Make the per-call price legible so buyers can model an agent run end to end.
Outlook — what to watch
Anthropic has already moved tool use out of tokens into explicit per-action fees (web search $10/1k, code execution per hour); watch whether OpenAI and Google follow — if they do, per-action becomes the default. The natural endpoint is pricing the completed task, which edges into outcome-based pricing.
Bottom line
Per-action tool pricing has graduated from emerging to established: six vendors — search-native plus a general-model API (Anthropic) and browser infra (Browserbase) — now meter tool calls directly. The action is overtaking the token as the unit that matters.
FAQ
What is per-action (agent-tool) pricing?
Charging for each tool an AI agent invokes — a web search, a page visit, a code execution, a research task — as its own billed unit, separate from the tokens the model generates.
Which vendors price agent tools per action?
In the corpus: Groq (built-in web search / visits / code execution), Perplexity (Agentic Research tier), You.com (Research API effort tiers), Exa (per-endpoint + Agent), plus Anthropic (web search $10/1k, code execution per hour) and Browserbase (Fetch/Search per 1k). OpenAI and Google still fold tool use into tokens.
How do I budget for an agent that uses tools?
Estimate tool calls per task (searches, visits, runs), not just tokens — tool fees can exceed token cost for search-heavy agents. Use any per-run caps or free allotments the vendor offers.