Ask
Price change

Together AI cuts GPU cluster rates again — reserved H100 floor now $3.09/hr

Together AI pricing

Together cut GPU Cluster rates: on-demand HGX H100 to $3.99/hr (from $4.79) and reserved 7–30 day H100 to $3.59/hr (from $4.19), with the 91–180 day reserved floor at $3.09/hr.

Before

On-demand H100 $4.79/hr; reserved 7–30d H100 $4.19/hr, 31–90d $3.45, 91–180d $3.29

After

On-demand H100 $3.99/hr; reserved 7–30d H100 $3.59/hr, 31–90d $3.29, 91–180d $3.09

Together AI cut its managed GPU Cluster rates for the second time in a week. On-demand NVIDIA HGX H100 dropped to $3.99/hr (from $4.79), and the reserved tiers stepped down across the board: 7–30 day reservations to $3.59/hr (from $4.19), 31–90 day to $3.29 (from $3.45), and 91–180 day to $3.09/hr (from $3.29) — making $3.09/hr the new published H100 floor on a 91–180 day commit.

On-demand HGX H200 ($5.99/hr) and HGX B200 ($8.19/hr), and reserved H200/B200 rates, were unchanged. A new 1× H200 140GB dedicated-endpoint line appeared (priced “Contact us”), and the standard fine-tuning tier now states a $4.00 per-job minimum charge. The cut deepens Together’s cost-leadership position on managed Hopper-class GPUs versus peers like Fireworks AI and Baseten.

From Together AI's pricing timeline
Reserved + on-demand GPU cluster rate cut (H100 down 12–16%)

Together cut its GPU Cluster rates again. On-demand HGX H100 fell to $3.99/hr (from $4.79); reserved 7–30 day H100 fell to $3.59/hr (from $4.19), 31–90 day to $3.29 (from $3.45), and 91–180 day to $3.09/hr (from $3.29) — making the reserved H100 floor $3.09/hr. On-demand H200/B200 and reserved H200/B200 rates were unchanged. A 1× H200 140GB dedicated-endpoint line was added (priced "Contact us"). The standard fine-tuning tier now states a $4.00 per-job minimum charge.

About Together AI
together.ai ↗

Together AI runs a multi-SKU pure-usage cloud: per-token serverless inference for popular open-weight models (Llama 3.3 70B at $1.04/$1.04, DeepSeek V4 Pro at $1.74/$3.48 with $0.20 cached input, Qwen3.5 9B at $0.17/$0.25, GLM-5.1 at $1.40/$4.40), image generation billed per image or per megapixel (FLUX.2 [dev] $0.0154 per image, FLUX.1 [schnell] $0.0027 per megapixel, SD XL $0.0019 per megapixel), and per-hour dedicated and cluster GPUs. The same serverless API meters speech per 1M characters (Cartesia Sonic-2 and Sonic-3 at $65.00, Orpheus TTS at $15, Kokoro-82M at $4.00 after a 60% cut on 21 July 2026 — a roughly 16x spread that makes model choice, not volume, the driver of a speech bill), transcription per audio minute — all four serverless speech-to-text models (Whisper Large v3, Parakeet TDT 0.6B v3, and both Nemotron 3 and 3.5 ASR Streaming variants) unified onto one flat $0.0015/min rate on 29 July 2026 — and video per video (Google Veo 3.0 at $1.60, Sora 2 at $0.80).

Free tier
Yes
Commits
Available
Transparency
public

Together AI pricing history

  1. Aug 2026
    Compute rate card and full model catalog flat again; Specialized fine-tuning tab re-actuated after a two-week capture gap and shows catalog rotation, not repricing
  2. Aug 2026
    Compute rate card flat again; DeepSeek V4 Pro 0813 + 6 HappyHorse video models added; GLM-5.1 and Kimi K2.6 rotated off the catalog
  3. Aug 2026
    Compute rate card flat a fifth straight capture; Cogito v2.1 671B, Rnj-1 Instruct, and Inkling Small's dual-meter billing confirmed
  4. Aug 2026
    Compute rate card flat a fourth straight week; Qwen3.8-2.4T-A95B and Inkling Small added
  5. Aug 2026
    Provisioned Throughput capacity increase + Kimi K3 added; PrismML confirmed Free
Full Together AI timeline

More Together AI activity

All pricing activity