Perplexity launches Gateway API for its own hosted open-weight models
Perplexity added a fifth developer API, the Gateway API, offering unified access to open-weight models it hosts itself (DeepSeek, Kimi, GLM) at Perplexity-set per-token prices from $0.13/1M input tokens, distinct from the Agent API's at-cost third-party resale.
Perplexity's developer platform had four API surfaces: Sonar API (token + per-request search fee), Search API ($5.00/1K requests), Agent API (third-party models resold at direct provider rates with no markup, plus metered tool calls), and Embeddings API. No Perplexity-hosted, Perplexity-priced model inference product existed.
A fifth surface, the Gateway API, launched at docs.perplexity.ai/docs/gateway/models: three Perplexity-hosted open-weight models — perplexity/deepseek-v4-flash-0731 ($0.13 input / $0.26 output / $0.028 cache-read per 1M tokens), perplexity/kimi-k3 ($3.00 / $15.00 / $0.30), and perplexity/glm-5.2 ($1.40 / $4.40 / $0.14) — billed per token with no per-request fee, reachable via OpenAI-compatible Chat Completions and Anthropic-compatible Messages endpoints under one API key.
Perplexity’s developer documentation quietly added a new top-level product on 2026-08-11: the Gateway API, described as giving “unified access to open-weight models hosted by Perplexity through a single endpoint and a single API key.” It sits alongside — not inside — the existing Agent API, Sonar API, Search API, and Embeddings API in the docs sidebar.
The structural distinction matters. Perplexity’s Agent API resells third-party frontier models (OpenAI, Anthropic, Google, xAI, Z.AI, Moonshot AI, NVIDIA) “at direct provider rates with no markup” — Perplexity takes zero model margin there and instead charges for the retrieval/tool layer wrapped around the tokens. The Gateway API inverts that: Perplexity hosts the inference itself for three open-weight models (from DeepSeek, Moonshot AI, and Z.AI) and sets its own per-token rate card, complete with a discounted cache-read rate and reasoning tokens billed at the output rate. It is the first Perplexity developer product priced with real hosting margin rather than at-cost pass-through, positioning Perplexity alongside inference-hosting platforms like Together AI, Fireworks, and DeepInfra for open-weight model serving.
No existing rate moved: Sonar token and request-fee pricing, the $5.00-per-1,000 Search API, all Agent API tool prices, and Embeddings API rates were all re-verified unchanged on the same pass.
Perplexity's developer docs added a fifth API surface, the Gateway API, offering "unified access to open-weight models hosted by Perplexity through a single endpoint and a single API key" via OpenAI-compatible Chat Completions and Anthropic-compatible Messages formats. Unlike the Agent API's zero-markup third-party resale, Gateway models are hosted and priced by Perplexity itself: perplexity/deepseek-v4-flash-0731 ($0.13 input / $0.26 output / $0.028 cache-read per 1M tokens), perplexity/kimi-k3 ($3.00 / $15.00 / $0.30), and perplexity/glm-5.2 ($1.40 / $4.40 / $0.14), billed per token with no per-request fees. No existing Sonar, Search API, Agent API tool, or Embeddings rate moved.