Ask
Price change

Together AI reprices serverless and cuts GPU cluster rates

Together AI pricing

Together cut GPU cluster rates (on-demand H100 $5.49→$4.79, reserved 7–30d H100 $4.99→$4.19), repriced serverless models, and began publishing cached-input rates on serverless inference.

Before

On-demand H100 $5.49/hr, reserved 7–30d H100 $4.99/hr, B200 reserved $9.65/hr; DeepSeek V4 Pro $2.10/$4.40; Llama 3.3 70B $0.88/$0.88; no published cached-input discount.

After

On-demand H100 $4.79/hr, reserved 7–30d H100 $4.19/hr (as low as $3.29/hr on 91–180d), B200 reserved $7.99/hr; DeepSeek V4 Pro $1.74/$3.48 with $0.20 cached input; Llama 3.3 70B $1.04/$1.04; cached-input rates now published.

Together AI repriced its serverless rate card and cut GPU cluster rates across the board. On-demand cluster H100 dropped to $4.79/hr (from $5.49) and B200 to $8.19/hr (from $9.95); reserved 7–30 day H100 fell to $4.19/hr (from $4.99) and B200 to $7.99/hr (from $9.65), with H100 reaching $3.29/hr on a 91–180 day reservation. H200 now appears on the cluster rate card (on-demand $5.99/hr, reserved $4.99–$3.99/hr).

On serverless, DeepSeek V4 Pro dropped to $1.74/$3.48 (from $2.10/$4.40), Qwen3.5 9B rose to $0.17/$0.25 (from $0.10/$0.15), and Llama 3.3 70B rose to $1.04/$1.04 (from $0.88/$0.88). Most notably, Together now publishes cached-input rates on serverless models (e.g. DeepSeek V4 Pro $0.20, GLM-5.1/5.2 $0.26, Kimi K2.6 $0.20) — closing the previously-flagged competitive gap against Fireworks, OpenAI, and Anthropic, which all shipped cached-input discounts earlier.

Dedicated endpoint rates (H100 $6.49/hr, HGX B200 180GB $11.95/hr), Code Sandbox ($0.0446/vCPU-hour, $0.0149/GiB-hour), Code Interpreter ($0.03/session), storage ($0.16/GiB-month), and the standard fine-tuning rate card were unchanged.

From Together AI's pricing timeline
Serverless re-pricing + GPU cluster rate cuts + cached input

Together repriced its serverless rate card and cut GPU cluster rates. DeepSeek V4 Pro dropped to $1.74/$3.48 (from $2.10/$4.40) and now shows a $0.20 cached-input rate; Qwen3.5 9B rose to $0.17/$0.25 (from $0.10/$0.15); Llama 3.3 70B rose to $1.04/$1.04 (from $0.88/$0.88). On-demand cluster H100 fell to $4.79/hr (from $5.49) and B200 to $8.19/hr (from $9.95); reserved 7–30 day H100 fell to $4.19/hr (from $4.99) and B200 to $7.99/hr (from $9.65), with H100 as low as $3.29/hr on a 91–180 day reservation. Cached-input pricing is now published on serverless models.

About Together AI
together.ai ↗

Together AI runs a multi-SKU pure-usage cloud: per-token serverless inference for popular open-weight models (Llama 3.3 70B at $1.04/$1.04, DeepSeek V4 Pro at $1.74/$3.48 with $0.20 cached input, Qwen3.5 9B at $0.17/$0.25, GLM-5.1 at $1.40/$4.40), image generation billed per image or per megapixel (FLUX.2 [dev] $0.0154 per image, FLUX.1 [schnell] $0.0027 per megapixel, SD XL $0.0019 per megapixel), and per-hour dedicated and cluster GPUs. The same serverless API meters speech per 1M characters (Cartesia Sonic-2 and Sonic-3 at $65.00, Orpheus TTS at $15, Kokoro-82M at $4.00 after a 60% cut on 21 July 2026 — a roughly 16x spread that makes model choice, not volume, the driver of a speech bill), transcription per audio minute — all four serverless speech-to-text models (Whisper Large v3, Parakeet TDT 0.6B v3, and both Nemotron 3 and 3.5 ASR Streaming variants) unified onto one flat $0.0015/min rate on 29 July 2026 — and video per video (Google Veo 3.0 at $1.60, Sora 2 at $0.80).

Free tier
Yes
Commits
Available
Transparency
public

Together AI pricing history

  1. Aug 2026
    Compute rate card and full model catalog flat again; Specialized fine-tuning tab re-actuated after a two-week capture gap and shows catalog rotation, not repricing
  2. Aug 2026
    Compute rate card flat again; DeepSeek V4 Pro 0813 + 6 HappyHorse video models added; GLM-5.1 and Kimi K2.6 rotated off the catalog
  3. Aug 2026
    Compute rate card flat a fifth straight capture; Cogito v2.1 671B, Rnj-1 Instruct, and Inkling Small's dual-meter billing confirmed
  4. Aug 2026
    Compute rate card flat a fourth straight week; Qwen3.8-2.4T-A95B and Inkling Small added
  5. Aug 2026
    Provisioned Throughput capacity increase + Kimi K3 added; PrismML confirmed Free
Full Together AI timeline

More Together AI activity

All pricing activity