Ask
Launch

Modal launches Shared API — token-based pricing alongside per-second GPU billing

Modal pricing

Modal added a new OpenAI-compatible Shared API billed by token rather than GPU-second, launching first with Moonshot's Kimi K3 model; per-token rates are not yet published on Modal's pricing page.

Before

Modal billed only by GPU/CPU/memory-second (plus flat Starter/Team/Enterprise plan fees) across all compute products.

After

A parallel Shared API product now bills by token for select hosted models (starting with Kimi K3), alongside the existing per-second GPU rate card; Starter's $30/month free credit applies to Shared API usage too.

Modal’s pricing page began surfacing a banner (“Kimi K3 is live. Try the new Shared API with token-based pricing”) pointing to a new product line: an OpenAI-compatible Shared API endpoint metered by token rather than by GPU-second. This is Modal’s first departure from pure per-second compute billing since founding — it runs alongside, not instead of, the existing GPU/CPU/memory rate card, giving customers a choice between dedicated per-second GPU capacity and shared per-token inference for supported models.

The first model available on the Shared API is Moonshot’s Kimi K3. Modal has not yet published a per-token rate card on its public pricing page or billing docs, so exact input/output token prices are unknown as of this capture. A follow-up discovery pass is needed once Modal publishes the rate card.

From Modal's pricing timeline
Shared API Launched — Token-Based Pricing Alongside Per-Second GPU

Modal launched an OpenAI-compatible Shared API metered by token rather than by GPU-second, going live first with Moonshot's Kimi K3 model. The Shared API runs in parallel with — not instead of — the existing per-second GPU/CPU/memory rate card, and Starter's $30/month free credit grant applies to Shared API usage too. This is Modal's first departure from pure per-second compute billing since founding; per-token input/output rates were not yet published on modal.com/pricing as of this capture.

About Modal
modal.com ↗

Modal runs a per-second pure-usage compute model: GPU rates from T4 at $0.000164/sec to B300 at $0.001972/sec, with B200 at $0.001736/sec, H200 at $0.001261/sec, H100 at $0.001097/sec, RTX PRO 6000 at $0.000842/sec, A100 80GB at $0.000694/sec, A100 40GB at $0.000583/sec, L40S at $0.000542/sec, A10 at $0.000306/sec, and L4 at $0.000222/sec.

Free tier
Yes
Commits
Available
Transparency
public

Modal pricing history

  1. Aug 2026
    Shared API Access Narrowed to Team and Enterprise
  2. Aug 2026
    Team Plan Container Concurrency Raised 5x to 5,000
  3. Jul 2026
    Shared API Launched — Token-Based Pricing Alongside Per-Second GPU
  4. Jul 2026
    B300 GPU + Sandbox/Notebooks Pricing + Rate Modifiers
  5. Feb 2026
    AWS + GCP Marketplace Billing for Enterprise
Full Modal timeline

More Modal activity

All pricing activity