Ask
Launch

Together AI launches Provisioned Throughput (PTU) and cuts Dedicated Inference rates

Together AI pricing

Together AI launched a Provisioned Throughput (PTU) SKU at $0.05/PTU-min and cut Dedicated Inference rates — on-demand H100 to $5.49/hr, B200 to $8.99/hr.

Before

No reserved-throughput product; Dedicated Inference priced per instance — 1x H100 80GB $6.49/hr, 1x HGX B200 180GB $11.95/hr, H200 Contact us.

After

New Provisioned Throughput SKU reserves capacity in throughput units at $0.05/PTU-minute (MiniMax M3, GLM-5.2); Dedicated Inference restructured to per-GPU-per-hour, on-demand vs reserved — on-demand H100 $5.49/hr, B200 $8.99/hr, with H200/B300/GB200/GB300 quoted Contact us and all reserved capacity Contact sales.

Together AI added a fifth pricing surface, Provisioned Throughput (PTU), which reserves dedicated inference capacity in throughput units billed per PTU-minute ($0.05/PTU-min on MiniMax M3 and GLM-5.2). Each PTU delivers a fixed, model-specific tokens-per-minute rate, and an on-page calculator sizes the PTUs required for a traffic profile and estimates monthly cost and savings versus a commercial model’s list price (assuming 24/7 provisioning).

In the same update, Dedicated Inference was restructured from a per-instance table into a per-GPU-per-hour grid split into on-demand (pay-as-you-go) and reserved (Contact sales) columns. On-demand HGX H100 dropped from $6.49 to $5.49/hr (-15%) and HGX B200 from $11.95 to $8.99/hr (-25%), while newly listed HGX H200, HGX B300, GB200 NVL72, and GB300 NVL72 lines are quoted “Contact us”. Serverless per-token and GPU Cluster rates were unchanged. The pricing page also carries a new Series C funding banner.

From Together AI's pricing timeline
Provisioned Throughput (PTU) launch + Dedicated Inference restructure & price cuts

Together launched Provisioned Throughput — a new SKU that reserves dedicated capacity in throughput units (PTUs) billed per PTU-minute ($0.05/PTU-min on MiniMax M3 and GLM-5.2), with an on-page calculator that sizes PTUs and estimates monthly cost vs. commercial-model list prices. In the same update, Dedicated Inference was restructured to a per-GPU-per-hour table split into on-demand (pay-as-you-go) vs reserved (Contact sales) columns: on-demand HGX H100 fell to $5.49/hr (from $6.49) and HGX B200 to $8.99/hr (from $11.95), and NVIDIA HGX H200, HGX B300, GB200 NVL72, and GB300 NVL72 lines were added (quoted "Contact us"). GPU Cluster and serverless rates were unchanged. The pricing page also carries a new Series C funding banner.

About Together AI
together.ai ↗

Together AI runs a multi-SKU pure-usage cloud: per-token serverless inference for popular open-weight models (Llama 3.3 70B at $1.04/$1.04, DeepSeek V4 Pro at $1.74/$3.48 with $0.20 cached input, Qwen3.5 9B at $0.17/$0.25, GLM-5.1 at $1.40/$4.40), image generation billed per image or per megapixel (FLUX.2 [dev] $0.0154 per image, FLUX.1 [schnell] $0.0027 per megapixel, SD XL $0.0019 per megapixel), and per-hour dedicated and cluster GPUs. The same serverless API meters speech per 1M characters (Cartesia Sonic-2 and Sonic-3 at $65.00, Orpheus TTS at $15, Kokoro-82M at $4.00 after a 60% cut on 21 July 2026 — a roughly 16x spread that makes model choice, not volume, the driver of a speech bill), transcription per audio minute — all four serverless speech-to-text models (Whisper Large v3, Parakeet TDT 0.6B v3, and both Nemotron 3 and 3.5 ASR Streaming variants) unified onto one flat $0.0015/min rate on 29 July 2026 — and video per video (Google Veo 3.0 at $1.60, Sora 2 at $0.80).

Free tier
Yes
Commits
Available
Transparency
public

Together AI pricing history

  1. Aug 2026
    Compute rate card and full model catalog flat again; Specialized fine-tuning tab re-actuated after a two-week capture gap and shows catalog rotation, not repricing
  2. Aug 2026
    Compute rate card flat again; DeepSeek V4 Pro 0813 + 6 HappyHorse video models added; GLM-5.1 and Kimi K2.6 rotated off the catalog
  3. Aug 2026
    Compute rate card flat a fifth straight capture; Cogito v2.1 671B, Rnj-1 Instruct, and Inkling Small's dual-meter billing confirmed
  4. Aug 2026
    Compute rate card flat a fourth straight week; Qwen3.8-2.4T-A95B and Inkling Small added
  5. Aug 2026
    Provisioned Throughput capacity increase + Kimi K3 added; PrismML confirmed Free
Full Together AI timeline

More Together AI activity

All pricing activity