Ask
Price change

Together AI raises GPU cluster reserved rates, unifies speech-to-text pricing

Together AI pricing

Together AI raised GPU cluster reserved H100 rates and unified speech-to-text pricing at $0.0015/audio-minute, cutting Nemotron 3.5 ASR pricing 67%.

Before

GPU Cluster reserved H100: $3.59/hr (7-30d), $3.29 (31-90d), $3.09 (91-180d); STT metered per-character for Parakeet ($0.0035/1M chars) and Nemotron 3 ASR ($0.0015/1M chars), per-minute for Whisper ($0.0015/min), Whisper Streaming ($0.0035/min), and Nemotron 3.5 ASR ($0.0045/min)

After

GPU Cluster reserved H100: $3.69/hr (7-30d), $3.45 (31-90d), $3.19 (91-180d); all four serverless speech-to-text models (Whisper Large v3, Parakeet TDT 0.6B v3, Nemotron 3 ASR Streaming, Nemotron 3.5 ASR Streaming) now bill at a unified $0.0015 per audio minute

Together’s managed GPU Cluster reserved H100 rate rose for the first time since the aggressive back-to-back cuts of June 2026 — up to $3.69/hr on a 7-30 day reservation (from $3.59), $3.45/hr on 31-90 days (from $3.29), and a $3.19/hr floor on 91-180 days (from $3.09). On-demand H100, and every published H200/B200 on-demand and reserved rate, held flat, so the increase is isolated to the reserved-H100 tenors.

In the same capture cycle, Together’s speech-to-text catalog was re-metered onto a single unit. Previously the four serverless STT models split across two different meters — per-1M-characters for Parakeet TDT 0.6B and Nemotron 3 ASR, per-audio-minute for Whisper and Nemotron 3.5 ASR — at four different rates. All four now bill at a flat $0.0015 per audio minute, which is a like-for-like 67% cut on Nemotron 3.5 ASR (from $0.0045/min) and a unit change (character to minute) for Parakeet and Nemotron 3 ASR. The serverless chat catalog also grew (Kimi K3, Gemma 4 31B, Qwen3.7-Plus, LFM2.5-8B-A1B) without moving any previously-published per-model rate.

From Together AI's pricing timeline
GPU Cluster reserved H100 rates rise (first increase since June cuts) + speech-to-text unified to $0.0015/audio-minute

Together's GPU Cluster reserved H100 rate rose across all three commitment tenors for the first time since the back-to-back June 2026 cuts: 7–30 day $3.59→$3.69/hr, 31–90 day $3.29→$3.45/hr, and the 91–180 day floor $3.09→$3.19/hr. On-demand H100 and every H200/B200 on-demand and reserved rate held flat, so the increase is isolated to reserved-H100 tenors. In the same capture cycle, all four serverless speech-to-text models (Whisper Large v3, Parakeet TDT 0.6B v3, and both Nemotron 3 and 3.5 ASR Streaming variants) were unified onto a single $0.0015-per-audio-minute meter — a 67% cut on Nemotron 3.5 ASR (from $0.0045/min) and a billing-unit change (from per-1M-characters to per-audio-minute) for Parakeet TDT 0.6B and Nemotron 3 ASR. Four new serverless chat models (Kimi K3, Gemma 4 31B, Qwen3.7-Plus, LFM2.5-8B-A1B) joined the catalog without moving any previously-published rate.

About Together AI
together.ai ↗

Together AI runs a multi-SKU pure-usage cloud: per-token serverless inference for popular open-weight models (Llama 3.3 70B at $1.04/$1.04, DeepSeek V4 Pro at $1.74/$3.48 with $0.20 cached input, Qwen3.5 9B at $0.17/$0.25, GLM-5.1 at $1.40/$4.40), image generation billed per image or per megapixel (FLUX.2 [dev] $0.0154 per image, FLUX.1 [schnell] $0.0027 per megapixel, SD XL $0.0019 per megapixel), and per-hour dedicated and cluster GPUs. The same serverless API meters speech per 1M characters (Cartesia Sonic-2 and Sonic-3 at $65.00, Orpheus TTS at $15, Kokoro-82M at $4.00 after a 60% cut on 21 July 2026 — a roughly 16x spread that makes model choice, not volume, the driver of a speech bill), transcription per audio minute — all four serverless speech-to-text models (Whisper Large v3, Parakeet TDT 0.6B v3, and both Nemotron 3 and 3.5 ASR Streaming variants) unified onto one flat $0.0015/min rate on 29 July 2026 — and video per video (Google Veo 3.0 at $1.60, Sora 2 at $0.80).

Free tier
Yes
Commits
Available
Transparency
public

Together AI pricing history

  1. Aug 2026
    Compute rate card and full model catalog flat again; Specialized fine-tuning tab re-actuated after a two-week capture gap and shows catalog rotation, not repricing
  2. Aug 2026
    Compute rate card flat again; DeepSeek V4 Pro 0813 + 6 HappyHorse video models added; GLM-5.1 and Kimi K2.6 rotated off the catalog
  3. Aug 2026
    Compute rate card flat a fifth straight capture; Cogito v2.1 671B, Rnj-1 Instruct, and Inkling Small's dual-meter billing confirmed
  4. Aug 2026
    Compute rate card flat a fourth straight week; Qwen3.8-2.4T-A95B and Inkling Small added
  5. Aug 2026
    Provisioned Throughput capacity increase + Kimi K3 added; PrismML confirmed Free
Full Together AI timeline

More Together AI activity

All pricing activity