Ask
All companies
technology

Together AI pricing

together.ai facts checked analysis reviewed
Estimate your Together AI cost — model your usage, see overages, and find the cheapest plan. Open calculator →
Quick summary
In this page
AI Summary
  • Together AI runs a multi-SKU pure-usage cloud: per-token serverless inference for popular open-weight models (Llama 3.3 70B at $1.04/$1.04, DeepSeek V4 Pro at $1.74/$3.48 with $0.20 cached input, Qwen3.5 9B at $0.17/$0.25, GLM-5.1 at $1.40/$4.40), image generation billed per image or per megapixel (FLUX.2 [dev] $0.0154 per image, FLUX.1 [schnell] $0.0027 per megapixel, SD XL $0.0019 per megapixel), and per-hour dedicated and cluster GPUs. The same serverless API meters speech per 1M characters (Cartesia Sonic-2 and Sonic-3 at $65.00, Orpheus TTS at $15, Kokoro-82M at $4.00 after a 60% cut on 21 July 2026 — a roughly 16x spread that makes model choice, not volume, the driver of a speech bill), transcription per audio minute — all four serverless speech-to-text models (Whisper Large v3, Parakeet TDT 0.6B v3, and both Nemotron 3 and 3.5 ASR Streaming variants) unified onto one flat $0.0015/min rate on 29 July 2026 — and video per video (Google Veo 3.0 at $1.60, Sora 2 at $0.80).
  • Dedicated inference endpoints (restructured to on-demand vs reserved, priced per GPU per hour) at $5.49/hr on-demand HGX H100 and $8.99/hr HGX B200 (down from $6.49 and $11.95), with H200/B300/GB200/GB300 lines quoted "Contact us" and all reserved capacity "Contact sales"; on-demand GPU clusters at $3.99/hr H100, $5.99/hr H200, and $8.19/hr B200; reserved cluster rates (7–30 day commits) at $3.69/hr H100 (ticked up slightly as of 2026-07-29) and $7.99/hr B200, dropping to $3.19/hr H100 on 91–180 day reservations — among the lowest published in the market.
  • Fine-tuning priced per 1M training tokens (LoRA / full-parameter): up to 16B at $0.48 / $1.20 SFT; 17–69B at $1.50 / $3.75; 70–100B at $2.90 / $7.25. A specialized per-model tier (DeepSeek-R1, GLM-5, Qwen3.5, gpt-oss) runs SFT LoRA $3–$40 and DPO LoRA $7.50–$100 per 1M tokens with per-model minimum charges of $6–$60.
  • Batch API offers up to 50% off, but only on a named list of serverless models — six other serverless models (DeepSeek-R1, DeepSeek-V3.1, DeepSeek V4 Pro, MiniMax M2.7, Kimi K2.5, Kimi K2.6) cannot run batch at all, and the discount never applies to dedicated inference. Code Sandbox at $0.0446/vCPU-hour and $0.0149/GiB-hour for agentic code execution; Code Interpreter at $0.03 per 60-minute session; storage at $0.16/GiB-month. A July 2026 Provisioned Throughput (PTU) SKU reserves dedicated capacity in throughput units billed per PTU-minute ($0.05/PTU-min on MiniMax M3 and GLM-5.2), sized via an on-page calculator that estimates cost vs. commercial-model list prices assuming 24/7 provisioning.
  • Founders include Stanford CRFM director Percy Liang and Stanford ML researcher Chris Re — making Together the rare commercial cloud with top-tier academic-lab architecture credibility on top of standard founder-CEO leadership.
  • Together raised a $305M Series B in February 2025 led by General Catalyst at $3.3B post-money; NVIDIA, Salesforce Ventures, and others participated. A Series C reported in late 2025 pushed the valuation past $5B, and on 1 July 2026 Together raised a further $800M at an $8.3B valuation — with the CEO declining to rule out an IPO the following year — a war chest that underwrites the repeated 2026 GPU rate cuts as continued price leadership rather than margin recovery.
Pricing summary
Together AI 2026 — Multi-SKU AI Acceleration Cloud
Serverless tokens + dedicated endpoints + on-demand/reserved clusters + Code Sandbox; up to 50% Batch discount on select models
Free trial
Sign up free
Evaluating Together for proof-of-value
Annual commit
Enterprise
Custom
Sustained workloads, regulated industries
Dedicated endpoints
$5.49 /hr (on-demand H100)
Single-tenant per-GPU-hour inference
GPU clusters
From $3.69 /hr (reserved H100)
Training and large-batch inference
New
Provisioned Throughput
$0.05 /PTU-minute
Reserved capacity in throughput units (PTUs)
No monthly fee. Dedicated endpoints were restructured and cut on 2026-07-14 (on-demand H100 now $5.49/hr, B200 $8.99/hr) and a Provisioned Throughput (PTU) SKU launched at $0.05/PTU-min; the compute rate card was flat through 2026-07-21, then GPU Cluster reserved H100 rates ticked up slightly on 2026-07-29 to $3.69/hr (7–30 day). Reserved cluster rates require 7–30 day commits (as low as $3.19/hr H100 on a 91–180 day reserve). The Batch API discount is up to 50% and applies only to a named six-model list; six other serverless models (including DeepSeek V4 Pro and Kimi K2.6) cannot run batch at all, and no batch discount applies to dedicated inference. Code Sandbox bills per vCPU-hour and GiB-hour separately.

About

Together AI is a San Francisco-based generative AI cloud company founded in June 2022 by Vipul Ved Prakash (ex-Topsy CEO and Cloudmark founder), Ce Zhang (then ETH Zurich systems professor, now at the University of Chicago), Chris Re (Stanford ML and Snorkel co-founder), and Percy Liang (Stanford CRFM director). The product is an AI Acceleration Cloud — a managed inference, training, and code-execution platform optimized for open-source models and customer-fine-tuned variants — combining per-token serverless inference, per-hour dedicated endpoints, per-hour GPU clusters (on-demand and reserved), Code Sandbox / Code Interpreter for agentic workflows, and a fine-tuning service. The runtime is built on Together’s proprietary Together Inference Engine with FlashAttention-3 kernels and speculative decoding pipelines.

By 2026 Together serves Salesforce, Zoom, Pika Labs, Hippocratic AI, Cartesia, Arc Institute, and roughly 1,500 other paying customers spanning enterprise AI infrastructure (RAG systems, multi-tenant fine-tunes, large-batch inference), academic research labs running open-source training, and AI-native startups serving production workloads. The company raised a $305M Series B in February 2025 led by General Catalyst at a $3.3B post-money valuation with NVIDIA, Salesforce Ventures, Coatue, and Kleiner Perkins participation; a Series C reported in late 2025 brought valuation past $5B.

Together competes with Fireworks AI, Baseten, Replicate, Anyscale, and Groq for the managed-inference market, plus hyperscaler offerings (AWS Bedrock, Vertex AI, Azure ML). Its differentiation is the combination of academic-lab founder credibility (Stanford CRFM + Stanford ML), one of the broadest open-source model catalogs in the industry, aggressive reserved cluster pricing (H100 at $3.69/hr is among the lowest published rates), and Code Sandbox as a non-token SKU that captures agentic code-execution workloads without forcing customers onto third-party sandbox providers.


Pricing summary : How Together’s multi-SKU AI Acceleration Cloud is priced

Together runs five parallel pricing surfaces on a unified credits balance. Serverless inference charges per million input/output tokens by model, with per-model rates published inline on the pricing page (rare among competitors who route to docs) — and a parallel set of non-token meters for the same serverless API: per image or per megapixel for image models, per video for video models, per 1M characters for text-to-speech, and per audio minute for transcription. Provisioned Throughput (PTU) — added July 2026, expanded to a third model on 2026-08-12 — reserves dedicated capacity in throughput units billed per PTU-minute ($0.05/PTU-min on MiniMax M3, GLM-5.2, and Kimi K3), sized via an on-page calculator that estimates monthly cost against commercial-model list prices. Dedicated endpoints are single-tenant per-GPU-per-hour rentals at $5.49/hr on-demand H100 and $8.99/hr on-demand B200 (reserved capacity is “Contact sales”), optimized for sustained-QPS workloads. GPU clusters are multi-node per-hour rentals for training and large-batch inference at $3.99/hr on-demand H100 and $3.69/hr reserved H100 (7–30 day commit), with attached storage at $0.16/GiB-month. Code Sandbox and Code Interpreter bill per vCPU-hour, GiB-hour, and per-session for agentic code execution.

A Batch API offers up to 50% off for asynchronous workloads, toggled directly on the rate card — but the discount is narrower than the toggle implies: Together’s docs name six discounted models (Llama 3.3 70B Instruct Turbo, Llama 3 70B chat, Qwen2.5 7B Instruct Turbo, Mixtral 8x7B, GLM-4.5-Air FP8, Whisper large-v3), everything else that can run batch runs at standard rates, six more serverless models (DeepSeek-R1, DeepSeek V3.1, DeepSeek V4 Pro, MiniMax M2.7, Kimi K2.5, Kimi K2.6) cannot run batch at all, and the discount never applies to dedicated inference. Fine-tuning is priced per 1M training tokens by model size and method, with a specialized tier ($3–$40 SFT LoRA, up to $100 DPO LoRA per 1M, plus $6–$60 per-model minimum charges) for frontier architectures like DeepSeek-R1 and GLM-5. Enterprise commitments unlock volume discounts on top of reserved cluster rates and enable VPC deployment, custom SLAs, and dedicated solutions engineering. This multi-SKU pure-usage architecture — token / PTU-minute / image / megapixel / video / character / audio-minute / GPU-hour / vCPU-hour / GiB-month — is one of the most expansive usage-based rate cards in AI infrastructure.

Latest move (as of 2026-08-04): The pricing page was restructured from a tabbed layout into one continuous scroll with anchor navigation covering all nine model categories (Chat, Vision, Image, Audio, Video, Transcribe, Embeddings, Rerank, Moderation) plus GPU Clusters, Sandbox, Storage, Dedicated Inference, and Fine-Tuning — the “IMAGE” / “SPECIALIZED PRICING” tab buttons now scroll rather than filter, so every section is visible in a single capture. Every compute meter (GPU Clusters, Dedicated Inference, Provisioned Throughput, Code Sandbox, Storage, both fine-tuning tiers) held exactly at its 2026-07-29 rate. The one confirmed price correction: the Transcribe tab’s speech-to-text rates are not uniformly $0.0015/audio-minute as recorded here on 2026-07-29 — NVIDIA Nemotron 3.5 ASR bills at $0.0045/min and a “Whisper Large v3 (Streaming)” line at $0.0035/min (see the Speech/audio/video table below for the full breakdown). Moderation also gained its first serverless model, Llama Guard 4 12B at $0.20/1M tokens. The serverless chat catalog kept growing (GLM-5.2, Inkling Small, DeepSeek V4 Flash 0731, NVIDIA Nemotron 3 Ultra, Kimi K2.7 Code, Qwen3.7-Max, Cogito v2.1 671B, and others) without moving any previously-published per-model rate. Confirmed flat (2026-08-11): every compute meter (GPU Clusters, Dedicated Inference, Provisioned Throughput, Code Sandbox, Storage, both fine-tuning tiers) held at its 2026-07-29 rate for a second consecutive week; the only rate-card movement was three new serverless model additions — Muse Glimmer joined the Chat and Vision catalogs at $0.35 input / $1.50 output per 1M tokens with a $0.04 cached-input rate, and ByteDance Seedance 2.5 ($0.115/video) and FLUX 3 ($0.17/video) joined the Video catalog — none of which moved a previously-published rate. Provisioned Throughput expanded + repriced by capacity (2026-08-12): Kimi K3 joined Provisioned Throughput as a third supported model, and the published per-PTU capacity rose for MiniMax M3 — input 138,840→166,667 TPM/PTU and cached 694,200→833,333 TPM/PTU (+20% each), output 23,140→41,667 TPM/PTU (+80%) — while GLM-5.2’s output capacity rose 9,620→11,364 TPM/PTU (+18%); the sticker price held flat at $0.05/PTU-minute for all three models, so buyers now get materially more throughput per dollar reserved. Every other compute meter (GPU Clusters, Dedicated Inference, Code Sandbox, Storage, both fine-tuning tiers) held its 2026-07-29 rate for a third consecutive week. Together’s model-catalog docs also newly confirm PrismML Ternary Bonsai 27B as fully “Free” (input and output both Free), resolving the ambiguous blank-output cell this page previously flagged. The chat, image, and video catalogs kept growing (e.g. Inkling Small, DeepSeek V4 Flash 0731, Cogito v2.1 671B on chat; Ideogram 4.0, Gemini 3.1 Flash Image “Nano Banana 2”, Wan 2.6 Image on the image side; Kling 2.1, Vidu 2.0/Q1, Wan 2.2, and two new Google Veo 3.0 Fast variants on video) without moving any other previously-published rate. The specialized fine-tuning tier also gained three new per-model rows: Llama 4 Maverick ($8.00 SFT LoRA / $20.00 DPO LoRA / $16.00 min), Qwen3-Coder-480B-A35B-Instruct (since delisted — see the 2026-08-26 update below), and Qwen3.5-122B-A10B ($6.00 / $15.00 / $10.00 min). Confirmed flat a fourth week (2026-08-13): a full per-tab re-capture (Chat, Image, Audio, Video, Transcribe, Specialized Fine-Tuning) confirms GPU Clusters, Dedicated Inference, Provisioned Throughput, Code Sandbox, Storage, and both fine-tuning tiers all held their 2026-07-29 rates for a fourth straight week — including every one of the three specialized-tier rows added on 2026-08-12. The catalog grew by exactly two SKUs: Qwen3.8-2.4T-A95B joined Chat at $2.50 input / $6.25 output per 1M tokens ($0.50 cached), and Inkling Small joined the speech catalog at $0.50 per 1M characters. The re-capture also surfaced one data-quality flag worth watching: the Video tab lists “Qwen3.6-Plus” — the same name as an existing Chat model — at $0.50/video, which reads as a likely rate-card artifact rather than a real per-video SKU. Confirmed flat again (2026-08-14): a fifth full per-tab re-capture found every price on every tab — GPU Clusters, Dedicated Inference, Provisioned Throughput, Code Sandbox, Storage, and both fine-tuning tiers — identical to the prior capture; the “Qwen3.6-Plus” $0.50/video artifact flagged the day before is also still unchanged, still unresolved. Catalog disclosure was the only movement: two chat models named in earlier updates without a price now show confirmed rates (Cogito v2.1 671B $1.25/$1.25, Rnj-1 Instruct $0.15/$0.15), and two more surfaced with prices for the first time (Qwen2.5 7B Instruct Turbo $0.30/$0.30 — one of the six Batch-discounted models — and Llama 3 8B Instruct Lite $0.14/$0.14). The most useful find: Inkling Small, previously recorded only on the speech catalog ($0.50 per 1M characters), is also confirmed on the Chat table at $0.50 input / $1.20 output per 1M tokens ($0.10 cached) — a second model, after Inkling itself, that bills on two separate meters depending on which endpoint it’s called through. Confirmed flat again (2026-08-25): the entire compute rate card — GPU Clusters, Dedicated Inference, Provisioned Throughput, Code Sandbox, Storage, and the standard fine-tuning tier — held its 2026-07-29 rate for an eleven-day span. Catalog churn continued: DeepSeek V4 Pro 0813 joined as a distinct dated snapshot ($1.32/$3.96, $0.13 cached) alongside the unchanged original DeepSeek V4 Pro, six new Alibaba HappyHorse video models arrived ($0.24 and $0.14/video), and GLM-5.1 and Kimi K2.6 were rotated off the chat catalog. NVIDIA Nemotron 3 ASR Streaming 0.6B moved from the Audio (per-character) tab to the Transcribe (per-audio-minute) tab, still at $0.0015 but on a different meter — the same billing-unit-move pattern flagged on 2026-07-29, now recurring on a different model. Confirmed flat again (2026-08-26): a full per-tab re-capture (Chat, Image, Audio, Video, Transcribe) found the entire compute rate card and every serverless model row on the live pricing page byte-identical to 2026-08-25 — no model added, removed, or repriced anywhere on the live page. A corrected tab selector also re-actuated the Specialized fine-tuning tab for the first time since 2026-08-12 (it had silently no-opped onto the default Standard tab in every capture from 2026-08-13 through 2026-08-25): the re-capture shows catalog rotation rather than repricing — Qwen3-235B-A22B, GLM-4.6/GLM-4.7, and Qwen3-Coder-480B-A35B-Instruct are no longer listed, DeepSeek-V4 Flash 0731 joined at $6.00/$15.00/$12.00, and three surviving rows picked up expanded aliases (“DeepSeek-R1 / V3” → “DeepSeek-V3.1”, “Kimi K2” → “Kimi K2.6 / Kimi K2.7-Code”, “GLM-5 / GLM-5.1” → “GLM-5.2 / GLM-5.1 / GLM-5”) at their exact prior price (see the Fine-tuning specialized tier table below). The only other movement was on the second-source docs catalog, which now discloses roughly a dozen video and image models not yet visible on the live pricing page (Sora 2 Pro, three Veo 3.1 variants, three Wan 2.7 models, two Vidu Q3 variants, PixVerse v5.6/v6, Seedream 5.0 Lite, Gemini 3.1 Flash-Lite Image, Pruna AI P-Image-Ideogram) — none confirmed billable until they appear on the live rate card — plus a new docs-vs-page conflict on Vidu 2.0 ($0.80/video in docs vs $0.28/video on the live page).

What makes this different: Reserved cluster pricing at $3.69/hr H100 (7–30 day commit) — and as low as $3.19/hr on a 91–180 day reservation — sits well below typical on-demand H100 rates from peers like Fireworks AI and Baseten. Together accepts a higher utilization risk (customer commits 7–30 days regardless of usage) in exchange for delivering lower per-hour cost — a structural choice that captures large-batch training and inference customers who can guarantee sustained utilization.


Pricing by product

Serverless inference (per-token, chat models)

ModelInput ($/1M)Output ($/1M)Cached input ($/1M)
Kimi K3$3.00$15.00$0.30
Qwen3.8-2.4T-A95B$2.50$6.25$0.50
DeepSeek V4 Pro$1.74$3.48$0.20
GLM-5.2$1.40$4.40$0.26
DeepSeek V4 Pro 0813$1.32$3.96$0.13
Qwen3.7-Max$1.25$3.75$0.13
Cogito v2.1 671B$1.25$1.25
Llama 3.3 70B$1.04$1.04
Inkling$1.00$4.05$0.17
Kimi K2.7 Code$0.95$4.00$0.19
NVIDIA Nemotron 3 Ultra$0.60$3.60$0.20
Qwen3.5-397B-A17B$0.60$3.60$0.35
Inkling Small$0.50$1.20$0.10
Qwen3.6-Plus$0.50$3.00
Gemma 4 31B$0.39$0.97
Muse Glimmer 30B$0.35$1.50$0.04
Qwen3.7-Plus$0.32$1.28
MiniMax M3$0.30$1.20$0.06
MiniMax M2.7$0.30$1.20$0.06
Qwen2.5 7B Instruct Turbo$0.30$0.30
Gemma-4-31B-it-Pearl$0.28$0.86
Qwen3 235B A22B Instruct 2507 FP8 Throughput$0.20$0.60
Qwen3.5 9B$0.17$0.25
Rnj-1 Instruct$0.15$0.15
gpt-oss-120B$0.15$0.60
DeepSeek V4 Flash 0731$0.14$0.28$0.03
Llama 3 8B Instruct Lite$0.14$0.14
Gemma 3n E4B Instruct$0.06$0.12
gpt-oss-20B$0.05$0.20
LFM2.5-8B-A1B$0.03$0.12
Ternary Bonsai 27B (PrismML)$0.00$0.00

Inkling and PrismML Ternary Bonsai 27B were added to the chat rate card in the week to 2026-07-21; PrismML Ternary Bonsai 27B was listed at an input rate of 0.00 with a blank output cell at the time. Update (2026-08-12): Together’s model-catalog docs now publish PrismML Ternary Bonsai 27B as Free for both input and output, resolving that earlier ambiguity — it is a genuinely free model, not an unpriced placeholder. Update (2026-08-13): Qwen3.8-2.4T-A95B joined the catalog at $2.50 input / $6.25 output per 1M tokens with a $0.50 cached-input rate — a new addition, not a repricing of any existing model — while every previously-published chat rate held. By 2026-07-29 the catalog had grown further with Kimi K3 (now the single highest-priced model on the card at $3.00 in / $15.00 out), Gemma 4 31B, Qwen3.7-Plus, Qwen3.6-Plus, LFM2.5-8B-A1B, Cogito v2.1 671B, and Rnj-1 Instruct — without any previously-published per-model rate moving. Muse Glimmer joined the Chat and Vision catalogs on 2026-08-11 at $0.35 input / $1.50 output ($0.04 cached) — again without moving any previously-published rate. The catalog kept growing again on 2026-08-12 (Inkling Small $0.50/$1.20 with $0.10 cached, DeepSeek V4 Flash 0731 $0.14/$0.28 with $0.03 cached, Qwen3.5-397B-A17B $0.60/$3.60 with $0.35 cached, Gemma-4-31B-it-Pearl $0.28/$0.86, Gemma 3n E4B Instruct $0.06/$0.12, gpt-oss-20B $0.05/$0.20, MiniMax M2.7 $0.30/$1.20 with $0.06 cached, among others) without moving any previously-published per-model rate. A “Batch API price” toggle above the table swaps the card to the asynchronous Batch rates, but only a named six-model list is actually discounted (up to 50% off) — several headline models here, including DeepSeek V4 Pro and Kimi K2.6, cannot run batch jobs at all. Cached-input rates are published on some, but not all, chat models. Update (2026-08-14): Cogito v2.1 671B ($1.25/$1.25) and Rnj-1 Instruct ($0.15/$0.15), both name-only catalog additions since 2026-07-29, show confirmed prices for the first time; Qwen2.5 7B Instruct Turbo ($0.30/$0.30, one of the six Batch-discounted models) and Llama 3 8B Instruct Lite ($0.14/$0.14) also surfaced with prices. Inkling Small is confirmed billing on two separate meters: $0.50 input / $1.20 output per 1M tokens ($0.10 cached) on this Chat table, and $0.50 per 1M characters on the Speech table below — the same dual-meter pattern the original Inkling model already carries. Every rate already published above held exactly flat. Update (2026-08-25): a full re-capture (all nine tabs plus a second-source re-check of Together’s serverless model-catalog docs) found the entire compute rate card — GPU Clusters, Dedicated Inference, Provisioned Throughput, Code Sandbox, Storage, and the standard fine-tuning tier — still identical to 2026-07-29. The catalog moved instead: DeepSeek V4 Pro 0813 joined as a new dated snapshot at $1.32 input / $3.96 output per 1M tokens ($0.13 cached) — a new SKU alongside the unchanged original DeepSeek V4 Pro ($1.74/$3.48), confirmed on both the pricing page and the docs catalog. DeepSeek V4 Flash 0731 ($0.14/$0.28, $0.03 cached) and Qwen3 235B A22B Instruct 2507 FP8 Throughput ($0.20/$0.60) — both previously only named in this page’s prose — are now formalized into the table above with confirmed rates. Two previously-listed models are no longer on the live rate card: GLM-5.1 and Kimi K2.6 are both absent from the 2026-08-25 capture (GLM-5.2 and Kimi K3/K2.7 Code appear to supersede them) — this reads as catalog rotation, not a repricing of any surviving model. Two display-name changes: “Muse Glimmer” → “Muse Glimmer 30B” and “PrismML Ternary Bonsai 27B” → “Ternary Bonsai 27B,” and the latter’s output cell — ambiguous since 2026-07-21 — is now confirmed as an explicit $0.00 (not blank) directly on the pricing page, matching the docs catalog’s “Free” listing. No previously-published rate on any surviving model moved.

Image generation (per image / per megapixel)

ModelPrice per imagePrice per MPDefault steps
Nano Banana Pro (Gemini 3 Pro Image)$0.134
GPT Image 2$0.053
Gemini 3.1 Flash Image (Nano Banana 2)$0.05
Gemini Flash Image 2.5 (Nano Banana)$0.039
Qwen Image 2.0 Pro$0.08
Qwen Image 2.0$0.04
Ideogram 4.0$0.06
GPT Image 1.5$0.034
Wan 2.6 Image$0.03
FLUX.2 [pro]$0.03
FLUX.2 [flex]$0.03
FLUX.2 [dev]$0.0154
FLUX.2 [max]$0.07050
FLUX.1 Kontext [max]$0.0828
FLUX.1 Kontext [pro]$0.0428
FLUX1.1 [pro]$0.04
Ideogram 3.0$0.06
Google Imagen 4.0 Ultra$0.06
ByteDance Seedream 4.0$0.03
Google Imagen 4.0 Preview$0.04
Juggernaut Pro Flux$0.0049
Qwen Image$0.0058
ByteDance Seedream 3.0$0.018
Google Imagen 4.0 Fast$0.02
FLUX.1 [schnell]$0.00274
HiDream-I1-Full$0.009
Juggernaut Lightning Flux$0.0017
SD XL$0.0019

Together bills image models on two different meters — per generated image for hosted third-party models, per megapixel (with a default step count) for the FLUX / SD family. Quoted prices assume the default steps shown; exceeding them adds cost. The rate card also footnotes that “displayed prices refer to the lowest resolution/duration settings” — hosted models such as Nano Banana Pro carry a higher rate above 2K, so the per-image figures here are floors, not flat rates. The catalog grew again on 2026-08-12 (Ideogram 4.0 $0.06/image, Gemini 3.1 Flash Image “Nano Banana 2” $0.05/image, Wan 2.6 Image $0.03/image, Qwen Image 2.0 $0.04/image, FLUX.2 [flex] $0.03/image, Juggernaut Pro Flux $0.0049/MP, GPT Image 1.5 $0.034/image, and several Google Imagen 4.0 / ByteDance Seedream variants, among others) without moving any previously-published rate. Update (2026-08-14): exact prices for the Google Imagen 4.0 / ByteDance Seedream variants named above are now confirmed — Imagen 4.0 Fast $0.02/MP, Preview $0.04/MP, Ultra $0.06/MP, Seedream 3.0 $0.018/MP, Seedream 4.0 $0.03/MP — alongside further catalog growth (Ideogram 3.0 $0.06/MP, HiDream-I1-Full $0.009/MP, Juggernaut Lightning Flux $0.0017/MP, Gemini Flash Image 2.5 “Nano Banana” $0.039/image), none of which moved a previously-published rate. Update (2026-08-25): the rows above are now confirmed directly in the pricing-page table rather than only in this prose (a full Image-tab re-capture). Two new additions: Nano Banana Pro (Gemini 3.1 Flash Image’s higher-tier sibling, “Gemini 3 Pro Image”) at $0.134/image, and Qwen Image 2.0 Pro at $0.08/image — both new SKUs, not repricings. One data-quality flag: Together’s serverless model-catalog docs list Qwen Image 2.0 at $0.035/MP and Qwen Image 2.0 Pro at $0.075/MP (both lower, and on a different meter, than this table’s $0.04/image and $0.08/image pricing-page figures) — treat the pricing-page numbers above as authoritative per this page’s standing practice, and the docs figures as unconfirmed until Together reconciles them. A second flag: the Image tab lists a “Qwen3.7-Plus” row with no price in any column (input, MP, or per-image) — the same kind of chat-model-name-on-the-wrong-tab artifact already flagged for “Qwen3.6-Plus” on the Video tab; not a real image SKU.

Speech, audio, and video (per 1M characters, per audio minute, per video)

ModelMeterRate
Cartesia Sonic-3 / Sonic-2 (TTS)per 1M characters$65.00
Orpheus TTSper 1M characters$15
Kokoro-82M TTSper 1M characters$4.00
Inkling (speech)per 1M characters$1.00
Inkling Small (speech)per 1M characters$0.50
NVIDIA Parakeet TDT 0.6B V3 Realtime (streaming ASR)per 1M characters$0.0035
Whisper Large v3 (transcribe, batch)per audio minute$0.0015
NVIDIA Parakeet TDT 0.6B v3 (transcribe, batch)per audio minute$0.0015
NVIDIA Nemotron 3 ASR Streaming 0.6B (transcribe)per audio minute$0.0015
Whisper Large v3 (Streaming) (transcribe)per audio minute$0.0035
NVIDIA Nemotron 3.5 ASR (transcribe)per audio minute$0.0045
Google Veo 3.0 + Audioper video$3.20
Google Veo 3.0 Fast + Audioper video$1.20
Google Veo 3.0per video$1.60
Kling 2.1 Masterper video$0.92
ByteDance Seedance 1.0 Proper video$0.57
Google Veo 2.0per video$2.50
MiniMax Hailuo 02per video$0.49
Kling 2.1 Proper video$0.32
Wan 2.2 T2Vper video$0.66
Sora 2per video$0.80
Google Veo 3.0 Fastper video$0.80
PixVerse v5per video$0.30
MiniMax 01 Directorper video$0.28
Vidu 2.0per video$0.28
Wan 2.2 I2Vper video$0.31
HappyHorse 1.0 T2V / I2V / R2Vper video$0.24
Vidu Q1per video$0.22
Kling 1.6 Standardper video$0.19
Kling 2.1 Standardper video$0.18
FLUX 3per video$0.17
ByteDance Seedance 2.0per video$0.16
HappyHorse 1.1 T2V / I2V / R2Vper video$0.14
ByteDance Seedance 1.0 Liteper video$0.14
ByteDance Seedance 2.5per video$0.115
Qwen3.6-Plus (unconfirmed artifact — see note)per video$0.50

ByteDance Seedance 2.5 and FLUX 3 joined the Video catalog on 2026-08-11 at $0.115 and $0.17 per video respectively — new catalog entries, not a repricing of ByteDance Seedance 2.0 ($0.16), which held flat. Kokoro-82M TTS was cut from $10.00 to $4.00 per 1M characters in the week to 2026-07-21 — a 60% reduction on the cheapest self-hosted-class TTS line, while Cartesia Sonic-2/-3 held at $65.00. Correction (2026-08-04): Together actually runs two separate speech-to-text billing surfaces that reuse near-identical model names at different rates — the pricing page’s Audio tab prices real-time streaming ASR per 1M characters (NVIDIA Parakeet TDT 0.6B V3 Realtime at $0.0035, NVIDIA Nemotron 3 ASR Streaming 0.6B at $0.0015), while its separate Transcribe tab prices batch transcription per audio minute, where only two of four listed rows are $0.0015 (Whisper Large v3, NVIDIA Parakeet TDT 0.6B v3) — NVIDIA Nemotron 3.5 ASR is $0.0045/min and a “Whisper Large v3 (Streaming)” line is $0.0035/min. This is not the flat $0.0015-per-audio-minute figure recorded here on 2026-07-29 for “all four” STT models — that reading traced to Together’s docs.together.ai model catalog (still showing all four Speech-to-Text models at a flat $0.0015/min as of this capture, and pointing readers to a separate “Transcription and Translations” doc page not captured here) rather than the live pricing page’s own Transcribe table, which is the actual billed rate. Embeddings (Multilingual e5 large instruct, $0.02/1M tokens) run on the same per-token meter; Rerank has no serverless models on the rate card, and Moderation now lists one model, Llama Guard 4 12B, at $0.20/1M tokens (Together’s docs confirm $0.20 on both the input and output columns for this model — the same effective rate shown on the pricing page’s single “price” column). The video catalog grew again on 2026-08-12 (Kling 2.1 Master/Pro/Standard and Kling 1.6 Standard at $0.18–$0.92, Vidu 2.0 and Vidu Q1 at $0.22–$0.28, Wan 2.2 I2V/T2V at $0.31/$0.66, two new Google Veo 3.0 Fast variants at $0.80/$1.20, PixVerse v5 at $0.30, two ByteDance Seedance 1.0 variants, MiniMax Hailuo 02 and MiniMax 01 Director, and Google Veo 2.0 at $2.50) without moving any previously-published rate; Whisper Large v3, Parakeet TDT 0.6B v3, and both Nemotron ASR variants confirmed unchanged at their 2026-07-29 rates on both the pricing page and Together’s docs. Update (2026-08-13): a full per-tab re-capture (Audio, Video, Transcribe) confirms every rate in both tables above held exactly flat, and fills in exact prices for rows the 2026-08-12 update had only named in prose: ByteDance Seedance 1.0 Lite $0.14/video and 1.0 Pro $0.57/video, MiniMax Hailuo 02 $0.49/video and MiniMax 01 Director $0.28/video, and Kling 1.6 Standard specifically at $0.19/video (vs. 2.1 Standard’s $0.18 floor and 2.1 Master’s $0.92 ceiling). Inkling Small also joined the speech catalog at $0.50 per 1M characters, half the price of the existing Inkling line. One data-quality flag: the Video tab lists a row named “Qwen3.6-Plus” at $0.50/video — identical to the existing Chat-tab model of the same name (priced $0.50 input / $3.00 output per 1M tokens) — which looks like a rate-card artifact (a model name duplicated across tabs) rather than a genuine per-video SKU; treat it as unconfirmed pending clarification from Together. Update (2026-08-14): a fifth re-capture confirms every rate on this table held exactly flat, including the unresolved “Qwen3.6-Plus” $0.50/video row. Inkling Small is confirmed billing on two meters, not one: $0.50 per 1M characters here on Speech, and $0.50 input / $1.20 output per 1M tokens on the Chat table — the same dual-meter pattern the original Inkling model already carries (see the Serverless inference table above for the full detail). Update (2026-08-25): a full re-capture of the Audio, Video, and Transcribe tabs found a surface move, not a price move — NVIDIA Nemotron 3 ASR Streaming 0.6B no longer appears on the Audio (streaming, per-1M-character) tab where it was priced at $0.0015; it now appears only on the Transcribe (batch, per-audio-minute) tab, still at $0.0015, but on a different meter. The number is unchanged; the unit it bills on is not, and — as with the 2026-07-29 unit change on this same table — the pricing page carries no on-page flag of the move. Six new video models joined the catalog: Alibaba’s HappyHorse 1.0 (T2V/I2V/R2V variants, $0.24/video each) and HappyHorse 1.1 (T2V/I2V/R2V, $0.14/video each, roughly 42% cheaper than 1.0). Kling 2.1 Pro’s exact rate ($0.32/video, previously only implied by the $0.18–$0.92 Kling range) is now confirmed. Every other rate on this table, including the still-unresolved “Qwen3.6-Plus” $0.50/video artifact, held exactly flat. Together’s serverless model-catalog docs independently confirm NVIDIA Nemotron 3 ASR Streaming 0.6B at $0.0015/audio-min (consistent with the live pricing page) but continue to publish NVIDIA Nemotron 3.5 ASR Streaming 0.6B at $0.0015/audio-min — conflicting with the live pricing page’s $0.0045/min for “NVIDIA Nemotron 3.5 ASR” — the same docs-vs-live-page discrepancy flagged since 2026-08-04; the live pricing page remains the priced-against source of record. Update (2026-08-26): a full re-capture of the Audio, Video, and Transcribe tabs found every rate in both tables above held exactly flat, including the still-unresolved “Qwen3.6-Plus” $0.50/video artifact. A new docs-vs-live-page conflict surfaced: Together’s serverless model-catalog docs price Vidu 2.0 at $0.80/video (listed as 720p/8s) versus $0.28/video on the live pricing page’s Video table — a second Vidu-family/Qwen-Image-style discrepancy between the two surfaces; the live pricing page remains the priced-against source of record. The docs catalog also newly lists roughly a dozen video and image models not yet visible on the live pricing page — Sora 2 Pro ($2.40/video), three Google Veo 3.1 variants ($0.05–$0.08/video), three Wan 2.7 T2V/I2V/R2V models ($0.10/video each), two Vidu Q3 variants ($0.0975–$0.195/video), PixVerse v5.6 and v6 ($0.09–$0.13/video), ByteDance Seedream 5.0 Lite ($0.035/MP), Google Gemini 3.1 Flash-Lite Image “Nano Banana 2 Lite” ($0.069/image), and Pruna AI P-Image-Ideogram ($0.00225/MP) — none are added to the tables above until they are confirmed on the live rate card.

Provisioned Throughput (PTU) — reserved capacity in throughput units

New in July 2026. PTUs reserve dedicated capacity billed per PTU-minute; each PTU delivers a fixed, model-specific tokens-per-minute (TPM) rate that differs for input, cached, and output tokens. An on-page calculator (“Estimate your PTUs & cost”) sizes the PTUs required for a target traffic profile and estimates monthly cost and savings vs. a selected commercial model’s list price (assuming continuous 24/7 provisioning, ~43,800 min/mo).

ModelInput TPM/PTUCached TPM/PTUOutput TPM/PTUPrice ($/PTU-min)
MiniMax M3166,667833,33341,667$0.05
Kimi K316,667166,6673,333$0.05
GLM-5.235,714192,30811,364$0.05

Kimi K3 joined Provisioned Throughput on 2026-08-12 as a third supported model. In the same update, MiniMax M3’s published per-PTU capacity rose (input 138,840→166,667, cached 694,200→833,333, output 23,140→41,667 TPM/PTU) and GLM-5.2’s output capacity rose (9,620→11,364 TPM/PTU) — the $0.05/PTU-minute sticker price held flat for all three models, so the capacity increase is an effective cut in cost per unit of reserved throughput.

Dedicated endpoints (single-tenant, per GPU per hour)

Restructured 2026-07-14 into on-demand (pay-as-you-go) vs reserved (Contact sales) columns, with a per-GPU-per-hour cut on the priced lines (on-demand H100 now $5.49/hr, B200 $8.99/hr):

HardwareOn-demand (PAYG)Reserved
NVIDIA HGX H100$5.49Contact sales
NVIDIA HGX H200Contact usContact sales
NVIDIA HGX B200$8.99Contact sales
NVIDIA HGX B300Contact usContact sales
NVIDIA GB200 NVL72Contact usContact sales
NVIDIA GB300 NVL72Contact usContact sales

GPU clusters (multi-node)

On-demand (per hour):

HardwareHourly rate
NVIDIA HGX H100$3.99
NVIDIA HGX H200$5.99
NVIDIA HGX B200$8.19

Reserved — rate steps down with longer reservation (minimum 6 days):

Hardware7–30 days31–90 days91–180 days181+ days
NVIDIA HGX H100$3.69$3.45$3.19Contact us
NVIDIA HGX H200$4.99$4.15$3.99Contact us
NVIDIA HGX B200$7.99$7.79$6.79Contact us
NVIDIA GB200 NVL72Contact usContact usContact usContact us
NVIDIA GB300 NVL72Contact usContact usContact usContact us
NVIDIA HGX B300Contact usContact usContact usContact us

Rates above are the ones on Together’s main pricing page. Update (2026-08-25): NVIDIA HGX B300 now appears as its own row in the GPU Clusters table (previously only listed under Dedicated Inference) — on-demand shows no rate (”—”) and every reserved tenor reads “Contact us”; every priced row above (H100, H200, B200) held exactly flat. Beware a second, conflicting vendor surface: the GPU Clusters landing page still advertises Together’s pre-June-2026 rate card (H100 at $5.49/hr on-demand and “starting at” reserved rates), and the cluster docs link to that page rather than publishing figures of their own — treat it as stale and price against the rate card above. On 2026-07-29 the reserved H100 rate ticked up slightly across all three tenors to $3.69/hr (7–30 day), $3.45/hr (31–90 day), and a $3.19/hr floor (91–180 day) — the first increase since the June 2026 cuts; on-demand H100, and all H200/B200 reserved and on-demand rates, were unchanged.

Fine-tuning — standard tier (per 1M training tokens)

Base model sizeSFT LoRASFT fullDPO LoRADPO full
Up to 16B$0.48$0.54$1.20$1.35
17B – 69B$1.50$1.65$3.75$4.12
70B – 100B$2.90$3.20$7.25$8.00

Each standard fine-tuning job is subject to a minimum charge of $4.00.

Fine-tuning — specialized per-model tier (per 1M training tokens)

Frontier architectures are priced per model on a separate “Specialized” tab (SFT LoRA / DPO LoRA / per-job minimum charge shown). This tab’s tab-click selector silently no-opped in every re-capture from 2026-08-13 through 2026-08-25 (the capture kept landing on the default Standard tab under the Specialized filename); a corrected selector re-actuated it for the first time in two weeks on 2026-08-26:

ModelSFT LoRADPO LoRAMinimum charge
Llama 4 Scout Instruct$3.00$7.50$6.00
gpt-oss-120B$5.00$12.50$6.00
DeepSeek-V4 Flash 0731 / DeepSeek-V4 Flash$6.00$15.00$12.00
Qwen3.5-122B-A10B$6.00$15.00$10.00
Llama 4 Maverick Instruct$8.00$20.00$16.00
Qwen3.5-397B-A17B$8.00$20.00$22.00
DeepSeek-V3.1$10.00$25.00$20.00
Kimi K2.6 / Kimi K2.7-Code$15.00$37.50$60.00
GLM-5.2 / GLM-5.1 / GLM-5$40.00$100.00$60.00

Specialized rates span roughly $3–$40 SFT LoRA and $7.50–$100 DPO LoRA per 1M tokens, with per-model minimum charges ($6–$60) on every listed model. Update (2026-08-26): the first successful re-capture of this tab since 2026-08-12 shows catalog rotation, not repricing — every surviving model’s rate is byte-identical to what this page recorded on 2026-08-12. Three rows from the 2026-08-12 snapshot are no longer listed — Qwen3-235B-A22B, GLM-4.6/GLM-4.7, and Qwen3-Coder-480B-A35B-Instruct are absent from today’s capture and not confirmed billable elsewhere on the live page (their 2026-08-12 rates are preserved in the Pricing evolution section below); one new row appeared, DeepSeek-V4 Flash 0731 at $6.00/$15.00/$12.00; and three rows carry expanded aliasing that reads as a rename rather than a reprice — “DeepSeek-R1 / V3” is now labeled “DeepSeek-V3.1”, “Kimi K2” is now “Kimi K2.6 / Kimi K2.7-Code”, and “GLM-5 / GLM-5.1” is now “GLM-5.2 / GLM-5.1 / GLM-5” — each at the exact prior price point. Because this tab went unverified for two weeks, this rotation could have happened on any date between 2026-08-13 and 2026-08-25; it is dated to 2026-08-26 here as the first-confirmed date, not necessarily the effective date.

Code Sandbox + Code Interpreter

ResourceRate
Code Sandbox (vCPU)$0.0446/vCPU-hour
Code Sandbox (memory)$0.0149/GiB-hour
Code Interpreter (session)$0.03/session (60-minute session, reusable within the window)
Storage (sandbox or model)$0.16/GiB-month

Sales motions across products: PLG / self-serve for serverless, on-demand clusters, Provisioned Throughput sizing, and Code Sandbox; sales-led for reserved dedicated/cluster capacity, Enterprise annual contracts, and VPC deployments.


Hidden costs : What Together AI customers actually pay beyond the rate card

Archetype A: AI-native startup running Llama 3.3 70B serverless with bursty traffic

A growth-stage AI assistant startup serving ~75K requests/day, average 1.5K input + 400 output tokens, with traffic concentrated in business hours:

Line itemMonthly cost
Input tokens (3.4M/day × 30 = 101M, Llama 70B at $1.04/1M)$105
Output tokens (900K/day × 30 = 27M, Llama 70B at $1.04/1M)$28
Batch API for nightly summarization workflows (10M tokens, -50%)$5
Code Interpreter for occasional agent execution (300 sessions × $0.03)$9
Estimated total~$147/month

For bursty traffic without sustained QPS, serverless dominates and the bill is dominated by per-token cost. Moving to a dedicated H100 endpoint ($5.49/hr on-demand, cut from $6.49 on 2026-07-14) would cost ~$4,000/month — only economical if sustained QPS rises above ~4 req/sec.

Archetype B: Mid-market team running a Llama 70B fine-tune on reserved H100 cluster

A team that fine-tuned Llama 3.3 70B (full-parameter SFT, 25M training tokens) and runs sustained inference on a reserved H100 cluster:

Line itemMonthly cost
Initial fine-tuning (one-time, 25M tokens × $3.75)$94
Reserved H100 (24h × 30 × $3.69/hr)$2,657
Storage for model artifacts + sandbox (50GiB × $0.16)$8
Code Sandbox for agent execution (200 vCPU-hours × $0.0446)$9
Estimated total~$2,700/month (after one-time $94 fine-tune)

Reserved cluster pricing dominates the bill at sustained utilization — and the $3.69/hr H100 reserved rate (up slightly from $3.59 on 2026-07-29; dropping to $3.19/hr on a 91–180 day reservation) makes Together one of the cheapest published managed-inference platforms for training and large-batch workloads. The trade-off is the reservation commit: even idle hours cost the customer.

Want to estimate your own Together AI bill? Use the Together AI pricing calculator to model serverless tokens, dedicated GPU hours, reserved cluster commits, and Code Sandbox costs.


Pricing evolution : Together’s pricing history from decentralized GPU pooling to AI Acceleration Cloud

Cadence

QuarterPrice changesProduct / SKU additionsNotes
2022 Q201Together founded; decentralized GPU pooling product
2023 Q401Inference Cloud GA + Series A ($102.5M)
2024 Q101Dedicated endpoints + fine-tuning launched
2024 Q311GPU Clusters launched at $5.49/hr on-demand H100
2025 Q100Series B ($305M) at $3.3B valuation
2025 Q201Batch API + Code Sandbox + Code Interpreter
2025 Q401FLUX.2 + FLUX-schnell + Stable Diffusion 3 image SKUs
2026 Q110Specialized fine-tuning tier (DeepSeek-R1, GLM-5) at $10–$100+
2026 Q2212026-06-24 broad serverless re-pricing + GPU cluster rate cuts (on-demand H100 $5.49→$4.79, reserved 7–30d H100 $4.99→$4.19); cached-input rates first published; HGX H200 cluster line added; 2026-06-30 second GPU cluster cut in a week (on-demand H100 $4.79→$3.99, reserved 7–30d H100 $4.19→$3.59, 91–180d floor $3.29→$3.09); 1× H200 140GB dedicated line added; standard fine-tuning $4.00 per-job minimum stated
2026 Q3432026-07-01 $800M raise at an $8.3B valuation (CEO would not rule out an IPO the following year); no rate move, but the balance-sheet backdrop for the quarter’s cuts. 2026-07-14 Provisioned Throughput (PTU) launched at $0.05/PTU-min (MiniMax M3, GLM-5.2) with an on-page sizing calculator; Dedicated Inference restructured from per-instance to a per-GPU-per-hour on-demand-vs-reserved grid and cut (on-demand H100 $6.49→$5.49, B200 $11.95→$8.99); HGX H200/B300 + GB200/GB300 NVL72 dedicated lines added (Contact us), all reserved dedicated capacity moved to Contact sales; Series C funding banner appeared. 2026-07-21 compute rate card held flat while the model catalog moved: Kokoro-82M TTS cut $10.00→$4.00 per 1M characters (−60%); Inkling added at $1.00/$4.05 with $0.17 cached input plus a $1.00 per-1M-character speech line; PrismML Ternary Bonsai 27B added at 0.00 input with no output rate published; 2026-07-29 GPU Cluster reserved H100 rates ticked up for the first time since June’s cuts (7–30d $3.59→$3.69, 31–90d $3.29→$3.45, 91–180d $3.09→$3.19); all four serverless speech-to-text models unified onto a flat $0.0015/audio-minute meter (a 67% cut on Nemotron 3.5 ASR, a character-to-minute unit change for Parakeet TDT 0.6B and Nemotron 3 ASR); four new chat models added (Kimi K3, Gemma 4 31B, Qwen3.7-Plus, LFM2.5-8B-A1B) without moving existing rates. 2026-08-13 compute rate card confirmed flat for a fourth straight week via a full per-tab re-capture; only movement was two new catalog SKUs (Qwen3.8-2.4T-A95B chat model, Inkling Small speech model). 2026-08-14 compute rate card confirmed flat a fifth time; catalog disclosure only — Cogito v2.1 671B and Rnj-1 Instruct (both name-only since 2026-07-29) got confirmed prices, and Inkling Small was confirmed as a dual-meter model billing on both Chat (per-token) and Speech (per-character). 2026-08-25 compute rate card confirmed flat again; DeepSeek V4 Pro 0813 added ($1.32/$3.96, $0.13 cached), six Alibaba HappyHorse video models added ($0.24 and $0.14/video), GLM-5.1 and Kimi K2.6 rotated off the chat catalog, NVIDIA HGX B300 added to the GPU Clusters table (Contact us), and NVIDIA Nemotron 3 ASR Streaming 0.6B moved from the Audio (per-character) tab to the Transcribe (per-audio-minute) tab at the same $0.0015 number. 2026-08-26 compute rate card and full live-page model catalog confirmed flat again (no additions, removals, or repricings anywhere on the live page); a corrected selector re-actuated the Specialized fine-tuning tab for the first time since 2026-08-12 and found catalog rotation with no repricing (3 rows delisted, 1 new row, 3 rows relabeled at unchanged prices); the second-source docs catalog also disclosed roughly a dozen unconfirmed video/image SKUs and a new Vidu 2.0 docs-vs-page price conflict ($0.80 docs vs $0.28 live page)

Tracked range: 2022 Q2–2026 Q3. Quarters not listed above were verified stable (0 price changes, 0 SKU additions).

Notable changes

  • 2023-11-29 — Inference Cloud GA with per-token serverless API; established Together as a Cloudflare-for-LLM-inference contender.
  • 2024-03-12 — Dedicated endpoints + fine-tuning launched; expanded from single-SKU per-token to multi-SKU platform.
  • 2024-09-20 — GPU Clusters launched at $5.49/hr on-demand H100 and $4.99/hr reserved (7–30 day commit); some of the lowest published H100 rates in managed inference.
  • 2025-06-18 — Batch API at 50% discount launched; Code Sandbox + Code Interpreter added non-token SKUs to the rate card.
  • 2025-10-08 — FLUX.2 + FLUX-schnell + Stable Diffusion 3 image generation SKUs launched at per-image rates.
  • 2026-01-15 — Specialized fine-tuning tier launched for DeepSeek-R1, GLM-5, and other large-context frontier architectures; reflected higher infrastructure cost of training on newer architectures.
  • 2026-06-24 — Broad serverless re-pricing and GPU cluster rate cuts: on-demand cluster H100 $5.49→$4.79/hr and B200 $9.95→$8.19/hr; reserved 7–30 day H100 $4.99→$4.19/hr and B200 $9.65→$7.99/hr (H100 as low as $3.29/hr on a 91–180 day reserve); an HGX H200 line was added at $5.99/hr on-demand. On serverless, DeepSeek V4 Pro fell to $1.74/$3.48 while Qwen3.5 9B and Llama 3.3 70B rose; cached-input rates (e.g. GLM-5.1/5.2 $0.26, DeepSeek V4 Pro / Kimi K2.6 $0.20) were published for the first time, and the specialized fine-tuning tier began carrying $6–$60 per-model minimum charges.
  • 2026-06-30 — A second GPU cluster rate cut within the same week: on-demand HGX H100 $4.79→$3.99/hr and reserved H100 stepping down across every window (7–30 day $4.19→$3.59, 31–90 day $3.45→$3.29, 91–180 day $3.29→$3.09), making $3.09/hr the new published reserved-H100 floor. On-demand and reserved H200/B200 rates were unchanged. A new 1× H200 140GB dedicated-endpoint line appeared (priced “Contact us”), and the standard fine-tuning tier began stating a $4.00 per-job minimum charge.
  • 2026-07-01 — Together raised $800M at an $8.3B valuation (Reuters / TechCrunch / Business Wire), with the CEO declining to rule out an IPO the following year (Axios). No rate move accompanied the raise, but it is the balance-sheet backdrop for the cuts on this timeline: a war chest of this size reads as funding for continued price leadership rather than a pivot to margin recovery, and it recasts the twice-in-a-week June H100 cuts and the July dedicated-inference reductions as the start of a sustained posture rather than a one-off promotion. The IPO signal is the one new uncertainty — a near-term public listing would put quarterly-margin scrutiny against the very rate aggression this page tracks. The Series C funding banner that surfaced on the pricing page on 2026-07-14 is the on-page marker of this round.
  • 2026-07-14 — Provisioned Throughput (PTU) launched — a fifth pricing surface that reserves dedicated capacity in throughput units at $0.05/PTU-minute (MiniMax M3, GLM-5.2), sized via an on-page calculator that estimates monthly cost and savings vs. commercial-model list prices. In the same update Dedicated Inference was restructured from a per-instance table to a per-GPU-per-hour on-demand-vs-reserved grid and cut: on-demand HGX H100 $6.49→$5.49/hr (−15%) and HGX B200 $11.95→$8.99/hr (−25%), with HGX H200/B300 and GB200/GB300 NVL72 lines added (quoted “Contact us”) and all reserved dedicated capacity moved to “Contact sales”. Serverless per-token and GPU cluster rates were unchanged. A Series C funding banner also appeared.
  • 2026-07-21 — The compute rate card stopped moving and the model catalog took over. Every compute meter (dedicated, clusters, PTU, Code Sandbox, storage, both fine-tuning tiers) held its 2026-07-14 rate, while Kokoro-82M TTS was cut 60% from $10.00 to $4.00 per 1M characters, Inkling was added at $1.00/$4.05 per 1M tokens with a $0.17 cached-input rate and a parallel $1.00 per-1M-character speech line, and PrismML Ternary Bonsai 27B was listed at an input rate of 0.00 with no output rate published. After four consecutive weeks of GPU and dedicated-capacity repricing, this is the first week where the rate movement is entirely on the catalog side.
  • 2026-07-29 — GPU Cluster reserved H100 rates rose across all three commitment tenors for the first time since the June 2026 cuts: 7–30 day $3.59→$3.69/hr, 31–90 day $3.29→$3.45/hr, and the 91–180 day floor $3.09→$3.19/hr; on-demand H100 and every H200/B200 on-demand and reserved rate held flat, so the increase is isolated to reserved-H100 tenors. In the same cycle, all four serverless speech-to-text models (Whisper Large v3, Parakeet TDT 0.6B v3, and both Nemotron 3 and 3.5 ASR Streaming variants) were unified onto a single $0.0015-per-audio-minute meter — a 67% cut on Nemotron 3.5 ASR (from $0.0045/min) and a billing-unit change (from per-1M-characters to per-audio-minute) for Parakeet TDT 0.6B and Nemotron 3 ASR — and four new chat models (Kimi K3, Gemma 4 31B, Qwen3.7-Plus, LFM2.5-8B-A1B) joined the catalog without moving any previously-published rate.
  • 2026-08-13 — A full per-tab re-capture (Chat, Image, Audio, Video, Transcribe, Specialized Fine-Tuning) confirmed the entire compute rate card — GPU Clusters, Dedicated Inference, Provisioned Throughput, Code Sandbox, Storage, and both fine-tuning tiers, including the three specialized-tier rows added 2026-08-12 — held flat for a fourth consecutive week. The only catalog movement was two new SKUs: Qwen3.8-2.4T-A95B on Chat ($2.50 input / $6.25 output per 1M tokens, $0.50 cached) and Inkling Small on the speech catalog ($0.50 per 1M characters, half the existing Inkling line). The re-capture also flagged a likely rate-card data artifact — a “Qwen3.6-Plus” row appears on the Video tab at $0.50/video, duplicating the name of an existing Chat-tab model — worth a follow-up check rather than treating as a confirmed per-video price.
  • 2026-08-14 — A fifth full per-tab re-capture found every price on every tab identical to the prior capture — GPU Clusters, Dedicated Inference, Provisioned Throughput, Code Sandbox, Storage, and both fine-tuning tiers all held their 2026-07-29 rates, and the unresolved “Qwen3.6-Plus” $0.50/video artifact is still present, unchanged. All movement was catalog disclosure: Cogito v2.1 671B ($1.25/$1.25) and Rnj-1 Instruct ($0.15/$0.15), both named without a price since 2026-07-29, got confirmed rates; Qwen2.5 7B Instruct Turbo ($0.30/$0.30) and Llama 3 8B Instruct Lite ($0.14/$0.14) surfaced with prices for the first time; and Inkling Small was confirmed as a second dual-meter model (after Inkling itself) — it bills $0.50 input / $1.20 output per 1M tokens on Chat and $0.50 per 1M characters on Speech, depending on which endpoint a caller invokes.
  • 2026-08-26 — A full per-tab re-capture (Chat, Image, Audio, Video, Transcribe) found the live pricing page byte-identical to 2026-08-25 — the entire compute rate card and every serverless model row held flat, with nothing added, removed, or repriced. A corrected tab selector re-actuated the Specialized fine-tuning tab for the first time since 2026-08-12 (it had silently no-opped onto the default Standard tab in every capture from 2026-08-13 through 2026-08-25): the tab shows catalog rotation, not repricing — 3 rows delisted (Qwen3-235B-A22B, GLM-4.6/GLM-4.7, Qwen3-Coder-480B-A35B-Instruct), 1 new row (DeepSeek-V4 Flash 0731 at $6.00/$15.00/$12.00), and 3 rows relabeled at their exact prior price. The only other movement traced to the second-source docs catalog, which now lists roughly a dozen video/image models (Sora 2 Pro, Veo 3.1 variants, Wan 2.7 variants, Vidu Q3 variants, PixVerse v5.6/v6, Seedream 5.0 Lite, Gemini 3.1 Flash-Lite Image, Pruna AI P-Image-Ideogram) not yet visible on the live page, plus a new Vidu 2.0 docs-vs-page price conflict ($0.80 docs vs $0.28 live page).

The June 2026 repricing in detail

The two June 2026 moves — 2026-06-24 and 2026-06-30, a week apart — are the first broad rate-card changes since GPU Clusters launched in 2024, and together they sharpen Together’s cost-leadership position rather than reversing it. Across the two cuts the reserved 7–30 day H100 rate fell from $4.99 to $3.59/hr and the 91–180 day floor from “$3.29 (post-launch)” down to $3.09/hr, while on-demand H100 dropped from $5.49 to $3.99/hr — pushing Together’s managed-Hopper rate further below typical peer on-demand pricing, so the structural advantage this page has tracked since launch only widened. The serverless re-pricing on the 24th mixed cuts and raises: DeepSeek V4 Pro got materially cheaper ($2.10/$4.40 → $1.74/$3.48) while the small-model floor (Qwen3.5 9B, Llama 3.3 70B) rose modestly, a re-rating toward the heavier-traffic reasoning models.

Three of the moves’ effects close gaps this page previously flagged as weaknesses. First, cached-input pricing is now published on much of the catalog (GLM-5.1/5.2 $0.26, DeepSeek V4 Pro and Kimi K2.6 $0.20, MiniMax M3 $0.06) — so any prior claim that Together shipped “no cached-input discount” is no longer true; the remaining gap is coverage (Llama 3.3 70B and gpt-oss still carry no cached rate), not absence. Second, the specialized fine-tuning tier now states $6–$60 per-model minimum charges, and the 06-30 cut added a $4.00 per-job minimum on the standard fine-tuning tier — reversing the earlier “no per-job minimum” claim across both tiers, so even routine fine-tunes now carry a small floor finance teams must model. Third, the two back-to-back GPU cuts signal that Together is racing the managed-Hopper rate down rather than holding it: $3.09/hr reserved H100 is roughly a 38% cut from the $4.99/hr 2024 launch reserve in under two years.

The July 2026 PTU launch in detail

The 2026-07-14 update is the first structural rate-card change since 2024 rather than another rate cut — it adds a fifth billing dimension (the PTU-minute) and re-shapes how dedicated capacity is sold. Provisioned Throughput answers the one gap Together’s pure-usage model left open for high-volume production buyers: per-token serverless bills bounce with traffic, and per-GPU-hour dedicated forces the customer to reason in hardware rather than throughput. A PTU abstracts both away — the customer reserves a fixed tokens-per-minute envelope ($0.05/PTU-minute on MiniMax M3 and GLM-5.2) and the on-page calculator translates a traffic profile into PTUs and a monthly cost, benchmarked against a commercial model’s list price. That framing — reserved throughput sold against the list price of a closed frontier model — is a direct play for buyers weighing an open-weight deployment versus an OpenAI/Anthropic API bill, and it is the clearest example yet of Together packaging its cost advantage as a budgeting story rather than a raw rate.

The simultaneous Dedicated Inference restructure sharpens rather than reverses the cost-leadership thesis this page has tracked: on-demand HGX H100 fell 15% ($6.49→$5.49/hr) and HGX B200 25% ($11.95→$8.99/hr), and the per-instance table became an on-demand-vs-reserved grid that pushes every serious commitment (“Contact sales”) and every next-gen part (H200, B300, GB200/GB300 NVL72 — “Contact us”) into a sales conversation. So the transparency this page praises on serverless thins on the newest dedicated hardware: the headline H100/B200 rates stay public, but Blackwell-Ultra and Grace-Blackwell capacity is now quote-only — a deliberate trade of published-rate breadth for sales-qualified enterprise capacity as the newest GPUs come online.

A week later, on 2026-07-21, the compute side went quiet and the catalog moved instead — and the shape of that move is more informative than its size. Cutting Kokoro-82M TTS 60% to $4.00 per 1M characters while leaving Cartesia Sonic-2/Sonic-3 at $65.00 stretches the spread on a single meter from roughly 6.5× to 16×, which is the clearest evidence yet that Together runs two pricing regimes on the same per-character unit: open-weight models priced down toward the marginal cost of serving them, and licensed commercial voices priced as a pass-through of someone else’s list price. For buyers that inverts the usual cost lever — on the speech meter, model selection now dominates volume, and a team that standardizes on Sonic for quality is choosing a 16× cost basis, not a 16% premium. The same regime logic explains PrismML Ternary Bonsai 27B arriving at 0.00 input: a 27B ternary-quantized model is cheap enough to serve that Together can list it at zero to seed adoption, the way it has historically discounted small open-weight lines. But the way that line is published — 0.00 in the input column, an empty output column — is the one genuinely awkward cell on a rate card whose whole advantage is that every number is inline and legible. Inkling is the unremarkable half of the move ($1.00 in / $4.05 out, mid-card on input with a 4× output multiplier), with two details worth noting: it ships with a $0.17 cached-input rate on day one, so cached pricing is now the default for new listings rather than a retrofit, and it appears on both the per-token chat table and the per-1M-character speech table — a single model billing on two meters depending on which modality you invoke.

A further week on, 2026-07-29 delivered the first genuine reversal on this timeline: the reserved-H100 tenors that had fallen in every prior 2026 move (7–30 day $3.59→$3.69/hr, 31–90 day $3.29→$3.45/hr, 91–180 day floor $3.09→$3.19/hr) ticked back up. The move is small (2.8–4.9% per tenor) and confined to reserved capacity — on-demand H100 and every H200/B200 line, on-demand or reserved, held flat — which reads less like a change of direction and more like Together finding, and correcting off of, a floor the June cuts pushed past. The same cycle applied the opposite logic to speech-to-text: four models that had split across two meters (per-character for Parakeet TDT 0.6B and Nemotron 3 ASR, per-audio-minute for Whisper and Nemotron 3.5 ASR) at four different rates were collapsed onto a single $0.0015-per-audio-minute price — a 67% cut for Nemotron 3.5 ASR and a straight unit swap, not just a rate change, for the two models that moved off per-character billing. Read together, the two moves show Together adjusting each meter toward what the underlying economics can sustain: GPU capacity nudged up off an unsustainably steep discount, transcription simplified and pushed down toward a single, defensible marginal-cost rate.


What’s unique : Together AI’s distinctive pricing mechanics

1. Per-model serverless rates published inline on the pricing page. Most inference middleware (Fireworks, Baseten) lists discount mechanics on the pricing page but routes to docs for per-model rates. Together’s inline display lets self-serve buyers compare model economics side-by-side without context switching — a pricing transparency UX advantage that materially reduces evaluation friction.

2. Reserved cluster pricing ($3.69/hr H100, dropping to a $3.19/hr floor) for short commits. Most platforms offer either on-demand (high price, no commit) or annual commits (lowest price, year-long lock-in). Together’s 7–30 day reserved tier creates a middle path: customers commit a week to a month at $3.69/hr (ticked up slightly from $3.59 on 2026-07-29, the first increase since June’s cuts) — and step down to a $3.19/hr floor on a 91–180 day reserve — capturing a meaningful discount over the $3.99/hr on-demand rate without annual lock-in. This commitment-flexibility innovation captures large-batch training customers who would balk at annual commits.

3. Code Sandbox as a non-token agentic SKU. Code Sandbox bills per vCPU-hour and GiB-hour — a fundamentally different metric than tokens — and Code Interpreter bills per session. Adding non-token SKUs to a token-dominated rate card lets Together capture agentic code-execution workloads without forcing customers onto third-party sandbox providers (E2B, Modal). The unified billing reduces vendor count for AI-native teams building autonomous agents.

4. Academic-lab founder credibility (Stanford CRFM + Stanford ML). Together’s co-founders include Stanford CRFM director Percy Liang and Stanford ML researcher Chris Re — making the platform’s optimization claims and model curation credible in a way that pure-engineering teams cannot replicate. The CRFM Helm leaderboard, the academic stewardship of open-source models, and the Together rate card share a knowledge base.

5. Multi-mode cluster pricing (on-demand + reserved) at different commit windows. Most clusters force a single mode choice (on-demand-only or reserved-only); Together’s three-tier structure (on-demand, 7–30 day reserved, Enterprise annual commit) lets customers match commit duration to workload predictability. This granular commitment design accommodates training cycles (weeks) and steady production (months) without forcing one model to fit both.

6. Provisioned Throughput (PTU) — reserved capacity sold in throughput units, benchmarked against commercial-model list prices. Launched July 2026, PTUs bill per PTU-minute ($0.05 on MiniMax M3 and GLM-5.2) for a fixed, model-specific tokens-per-minute envelope, and an on-page calculator sizes the reservation and estimates savings against a selected commercial model’s list price. Applying the Azure-OpenAI-style provisioned-throughput unit to open-weight models gives high-volume buyers a predictable-cost alternative to per-token variability — and frames the pitch as open-weight-vs-closed-frontier economics rather than raw GPU-hours.


Strengths & weaknesses

StrengthsWeaknesses
Per-model serverless rates published inline — best transparency in the categoryPer-model rates require reading a long inline table; no comparison filter
Reserved cluster H100 at $3.69/hr (down to a $3.19/hr floor), still among the lowest published in managed inference despite a small 2026-07-29 uptickReserved commits require a 7–30 day duration — idle hours still billed, and the reserved-H100 rate rose for the first time this year on 2026-07-29
Academic-lab founder credibility (CRFM + Stanford ML)Code Sandbox vCPU-hour rates need separate cost-modeling alongside token spend
Code Sandbox + Code Interpreter capture agentic code-execution without third-party toolsCached-input rates published on many but not all serverless models (Llama 3.3 70B, gpt-oss have none)
FLUX.2 + FLUX-schnell + SD XL image SKUs unified in rate cardA100 not prominently listed — A100 capacity available but not on the headline rate card
Multi-mode cluster pricing (on-demand, 7–30 day reserved, annual) accommodates many workload typesSpecialized fine-tuning tier (up to $40 SFT LoRA / $100 DPO LoRA per 1M) prices frontier-model tuning well above the standard tier
Provisioned Throughput (PTU) gives steady high-volume workloads predictable per-PTU-minute cost, with a calculator that quantifies savings vs commercial-model list pricesPTU launched on only two models (MiniMax M3, GLM-5.2), and its savings estimate assumes 24/7 provisioning — idle throughput is still billed; next-gen dedicated GPUs (H200, B300, GB200/GB300) are now quote-only
Open-weight speech models are priced toward marginal serving cost — Kokoro-82M TTS cut 60% to $4.00/1M characters on 2026-07-21 — and new chat listings now ship with a cached-input rate on day one (Inkling $0.17)The same-meter spread that cut creates (Kokoro $4.00 vs Cartesia Sonic-2/-3 $65.00 per 1M characters, ~16×) makes model choice, not volume, the dominant cost lever on speech; and PrismML Ternary Bonsai 27B is published at 0.00 input with a blank output column — the one ambiguous cell on an otherwise fully-priced card
All four serverless speech-to-text models unified onto one $0.0015/audio-minute rate on 2026-07-29 — transcription model choice no longer affects cost at all, and Nemotron 3.5 ASR got a 67% cut in the processThe same move changed the billing unit (character→minute) for Parakeet TDT 0.6B and Nemotron 3 ASR without a flag on the rate card — any customer who budgeted off the old per-character meter should re-check the bill

Billing UX : Together AI’s account controls and payment experience

  • Self-serve signup — Sign up at api.together.ai with email; trial credits applied automatically. Credit card required for production usage.
  • Unified credits balance — Serverless tokens, dedicated GPU hours, GPU cluster hours, fine-tuning training tokens, and Code Sandbox usage all bill against the same workspace credits balance.
  • Per-request usage metadata — API responses include input tokens, output tokens, and per-request cost so client applications can compute and surface real-time cost.
  • Per-model rate visibility — Pricing page displays per-model rates inline behind a nine-way modality tab strip (CHAT / VISION / IMAGE / AUDIO / VIDEO / TRANSCRIBE / EMBEDDINGS / RERANK / MODERATION); dashboard shows live consumption per model and per SKU.
  • “Batch API price” toggle — A switch above the serverless rate table flips the card to its asynchronous Batch equivalent, so buyers can price a batch workload without a separate page. The toggle is broader than the discount: per Together’s batch docs only six named models actually run at 50% off, other batch-eligible models “run at standard rates,” and six serverless models (DeepSeek-R1, DeepSeek V3.1, DeepSeek V4 Pro, MiniMax M2.7, Kimi K2.5, Kimi K2.6) are not available for batch processing at all.
  • Sortable rate card — Model, input, and output columns each carry sort controls (A–Z and by price), so the catalog can be ranked cheapest-first per meter.
  • Spend alerts — Configurable email and webhook alerts at $X spend per period.
  • Payment methods — Credit card and ACH on self-serve; wire transfer, invoice billing, and AWS/GCP Marketplace on Enterprise.
  • PTU sizing calculator — The pricing page’s “Estimate your PTUs & cost” tool takes model, comparison model, peak requests/sec, cache-hit rate, and input/output tokens per request, then returns PTUs required, estimated monthly cost, and estimated monthly savings vs. the selected commercial model’s list price (24/7 provisioning assumed).
  • Cluster reservation booking — 7–30 day GPU cluster reservations bookable directly via the dashboard with confirmed start dates; cancellation policies vary by SKU.
  • Audit logging + RBAC — Workspace-level RBAC on Pro+; SOC 2 audit-log exports on Enterprise via S3 or webhook.
  • Multi-region availability — US and EU regions standard for serverless; reserved clusters available in additional regions on Enterprise commitments.

Strategic wins : Why Together AI’s pricing decisions worked

1. Inline per-model rate publication removed evaluation friction

By publishing per-model rates on the pricing page rather than routing to docs, Together let self-serve buyers compare model economics in a single context. This transparency converts more self-serve customers and reduces sales-led overhead for low-value-deal segments. Most competitors’ docs-routing UX loses cost-sensitive evaluators who never get to the rate card before churning.

2. 7–30 day reserved cluster tier captured the middle of the commit-duration spectrum

Annual commits lock too much for many training and large-batch customers; on-demand is too expensive for sustained workloads. Together’s 7–30 day reserved tier created a middle option that converts customers who would otherwise self-build on raw cloud. The $3.69/hr H100 reserved rate (up slightly from $3.59 on 2026-07-29) — and a $3.19/hr floor on a 91–180 day reserve — is low enough to compete with hyperscaler EDP-discounted rates without forcing year-long commitments.

3. Code Sandbox as a non-token SKU expanded TAM beyond token-only buyers

Adding Code Sandbox ($0.0446/vCPU-hour) and Code Interpreter ($0.03/session) gave Together a SKU that captures agentic code-execution workloads — workloads that would otherwise go to E2B, Modal, or Pyodide. The unified billing balance reduces vendor count for AI-native teams building autonomous agents, locking in wallet share.

4. Academic-lab founder credibility as the platform-runtime trust anchor

Stanford CRFM director Percy Liang and Stanford ML researcher Chris Re as co-founders give Together unusual academic-lab credibility that customers extend to model curation, optimization claims, and platform design. For enterprise procurement leaders evaluating inference middleware, this founder profile distinguishes Together from pure-engineering teams in a way that is hard to replicate.

5. Unifying speech-to-text onto one flat rate removed model choice as a cost variable

On 2026-07-29 Together collapsed four serverless transcription models — previously split across two different meters (per-character and per-audio-minute) at four different price points — onto a single $0.0015-per-audio-minute rate, cutting Nemotron 3.5 ASR 67% in the process. For buyers evaluating STT accuracy, the pricing decision disappears: any of the four models costs the same, so procurement becomes an accuracy/latency bake-off rather than a cost trade-off — the mirror image of the widening TTS spread the 2026-07-21 Kokoro cut created on the adjacent per-character speech meter.


Areas to improve : Gaps in Together’s pricing approach

1. Cached-input discount is published on some, but not all, serverless models

As of June 2026 Together publishes cached-input rates on many serverless models (e.g. DeepSeek V4 Pro $0.20, GLM-5.1 $0.26, MiniMax M3 $0.06) — closing a gap it previously had versus Fireworks, OpenAI, Anthropic, and Baseten. But popular models like Llama 3.3 70B and gpt-oss still show no cached-input rate, so RAG and agent-loop workloads with high prefix re-use only benefit on a subset of the catalog. New listings now arrive with the rate attached — Inkling shipped on 2026-07-21 with $0.17 cached input on day one — which narrows this to a legacy-catalog backfill problem rather than a policy gap: the models missing a cached rate are the ones that predate the June 2026 repricing, and they include two of the most-used lines on the card. Extending cached-input discounting across the full model list would remove the remaining inconsistency.

2. Per-model rate comparison needs better filtering

The inline per-model rate table is comprehensive but long. Customers comparing 5–10 models must scroll and scan rather than filter. Adding a per-model filter / sort / search UI on the pricing page would convert more evaluation traffic into pilots without requiring API exploration.

3. Specialized fine-tuning tier (up to $40 SFT LoRA / $100 DPO LoRA per 1M) creates budget uncertainty

The wide per-model spread — from $3 SFT LoRA on Llama 4 Scout to $40 on GLM-5, and up to $100 DPO LoRA — makes it hard for finance teams to forecast fine-tuning budgets on DeepSeek-R1 or GLM-5 without reading the per-model table. The rates are published per model on a separate “Specialized” tab, so a per-model price calculator would further reduce friction for frontier-fine-tuning workloads that currently go to first-party providers.

4. A100 not on the headline rate card

The on-demand rate card lists H100 and B200 prominently but not A100. For non-frontier workloads that fit comfortably on A100, customers may compare Together’s headline H100 rate to a competitor’s published A100 rate and conclude Together is more expensive. Publishing an A100 rate (even at a “limited availability” disclaimer) would prevent unfavorable comparison.

5. Provisioned Throughput launched on only two models, and its calculator assumes 24/7 provisioning

The July 2026 PTU SKU covers only MiniMax M3 and GLM-5.2 at launch, so buyers who standardized on Llama, DeepSeek, or Qwen cannot yet reserve throughput. The sizing calculator also estimates savings assuming continuous 24/7 provisioning (~43,800 min/mo), which overstates the benefit for bursty or business-hours-only traffic where reserved throughput sits idle overnight. Extending PTU model coverage and adding a duty-cycle input to the calculator would make the savings estimate honest for non-continuous workloads and broaden PTU’s addressable base beyond always-on production endpoints.

6. Next-gen dedicated GPUs moved behind “Contact us” in the 2026-07-14 restructure

The Dedicated Inference restructure kept public rates on HGX H100 ($5.49/hr) and HGX B200 ($8.99/hr) but pushed HGX H200, HGX B300, and the GB200/GB300 NVL72 lines to “Contact us,” with all reserved dedicated capacity now “Contact sales.” That thins the published-rate transparency this page otherwise praises exactly where buyers are evaluating the newest Blackwell-Ultra and Grace-Blackwell hardware. Publishing at least indicative on-demand rates for the next-gen parts would preserve the self-serve legibility that differentiates Together on the rest of the rate card.

7. An unpriced catalog line undercuts the card’s main advantage

Together’s core pricing-page differentiator is that every model carries a visible number. PrismML Ternary Bonsai 27B, added on 2026-07-21, breaks that: it shows 0.00 in the input column and nothing at all in the output column. A buyer cannot tell whether that means genuinely free, promotional-for-now, or simply not-yet-priced — and the difference matters, because “free” is a plan input while “unpriced” is a procurement risk. A single word (“Free (promotional)” or “Pricing TBD”) would resolve it. The related legibility issue is cross-meter models: Inkling now appears on both the per-token chat table and the per-1M-character speech table, so a team costing “Inkling” has to know which modality it is invoking before it can pick a meter. As Together’s catalog spans nine modality tabs, flagging multi-meter models inline would prevent the wrong unit being budgeted.

8. The character-to-minute unit change on two STT models isn’t flagged as a unit change

When Parakeet TDT 0.6B v3 and Nemotron 3 ASR Streaming moved from a per-1M-character meter to a per-audio-minute meter on 2026-07-29, the rate card shows only the new $0.0015 number — there is no on-page signal that the unit itself changed, not just the price. A customer who built a forecast off the old per-character rate has no way to notice the switch without reading the changelog. A footnote flagging “billing unit changed from per-character to per-audio-minute, effective 2026-07-29” alongside the new rate would prevent a budget built on the old meter from silently mispricing a bill built on the new one.


Monetization stack & signals : how Together AI builds & buys its revenue engine

Buys 5 Builds 0 10 open roles

The read — where the monetization investment is going

Together AI runs a bought monetization stack, not an in-house metering build — no engineering-blog disclosure of a home-grown billing/metering service surfaced, and the finance/data org instead names third-party tooling. A RevOps posting describes "our integrated technology stack, including Salesforce" (CRM, stated in-use, double-sourced), and two data-warehouse engineering roles own building and maintaining "dbt transformation projects" on the analytics warehouse (data-platform, stated in-use across two live reqs). An Infrastructure Accounting Manager role names the ERP as NetSuite in-use ("integrations between ERP (NetSuite), procurement, and asset tracking"), corroborating a separate Sr. Revenue Accountant req that lists NetSuite as preferred experience — so rev-rec on NetSuite is a stated, double-sourced signal. The same Sr. Revenue Accountant req lists Metronome (usage-based billing) and Stripe (payments) as preferred experience — strong but inferred signals (preferred-qual framing, not an in-use disclosure) that the usage-metered revenue runs through bought metering + payments rather than a custom meter. Hiring is concentrated in customer-success/solutions for GPU-cluster and inference accounts (10 open roles) plus a finance/billing-data-platform build-out (2 billing-eng, 2 data-platform), reflecting a self-serve-plus-sales-led GPU cloud scaling its revenue-data and post-sale support functions.

Stack — build vs buy
Buys (vendor) · 5
  • Salesforce CRM Job post 1 Job post 2 Jun 2026

    “Oversee and optimize our integrated technology stack, including Salesforce, marketing automation (e.g., HubSpot), and sales engagement tools (e.g., MixMax)”

  • dbt Data platform Job post 1 Job post 2 Jun 2026

    “Build and maintain Airflow orchestrated pipelines and dbt transformation projects (modular, tested, documented)”

  • Metronome Metering inferred Job post Jun 2026

    “Experience with Metronome (usage-based billing) and Stripe (payments infrastructure)”

  • Stripe Payments inferred Job post Jun 2026

    “Experience with Metronome (usage-based billing) and Stripe (payments infrastructure)”

  • NetSuite Revenue recognition Job post 1 Job post 2 Jun 2026

    “Evaluate and enhance systems and integrations between ERP (NetSuite), procurement, and asset tracking tools to support rapid growth”

Open roles in the revenue & lifecycle org — 10
View open roles
  • Sr. Revenue Accountant Billing engineeringRevOps seen Jun 17, 2026
  • Senior Data Engineer Billing engineering seen Jun 17, 2026
  • Sales and Marketing Operations Manager RevOps seen Jun 17, 2026
  • Analytics Engineer — Data Warehouse Data platform seen Jun 17, 2026
  • Data Warehouse Engineer Data platform seen Jun 17, 2026
  • Customer Support Engineer (GPU Cluster) Customer success seen Jun 17, 2026
  • Customer Support Engineer (Inference) Customer success seen Jun 17, 2026
  • Technical Account Manager (TAM), GPU Cluster Customer success seen Jun 17, 2026
  • Solutions Architect (Inference) Customer success seen Jun 17, 2026
  • Forward Deployed Engineer (Inference & Post-Training) Customer success seen Jun 17, 2026
  • +5 more matched roles

Signals reviewed · derived from public job posts

Job postings fill and close over time — once a posting is filled we keep it as a dated citation (the quoted evidence remains); use View open roles for current listings.

Key takeaways

  1. Inline per-model rate publication beats docs-routing for self-serve conversion — but only where the units stay comparable. Together’s pricing page transparency converts evaluators that competitors lose to context switching, so self-serve usage-based platforms should display per-SKU per-model rates directly on the pricing page rather than route to documentation. The 2026-07-21 update shows the limit of that advantage: a single meter can carry two pricing regimes (Kokoro-82M TTS cut to $4.00 per 1M characters while licensed Cartesia voices held at $65.00, a ~16× spread on the identical unit), and one line was added at 0.00 input with no output rate at all. When rows on the same unit aren’t comparable — or aren’t priced — the card stops answering the buyer’s question. The 2026-07-29 speech-to-text unification is the counter-example: collapsing four transcription models onto one $0.0015-per-audio-minute rate removes the comparability problem entirely for that meter — at the cost of a one-time unit change (character→minute) that customers migrating off the old Parakeet/Nemotron 3 ASR rate need to re-check, since the card doesn’t flag that the unit itself moved.

  2. Commitment-based capacity pricing (multi-window reserved clusters + PTU throughput reservation) captures more buyers than on-demand-or-annual binaries. The 7–30 day reserved cluster tier converts customers with training cycles that don’t fit either extreme — Together’s reserved H100 rate ticked up slightly to $3.69/hr on 2026-07-29 (the first increase since June’s cuts, off a $3.59/hr floor), with the 91–180 day floor now $3.19/hr, still among the lowest published in managed inference — and the July 2026 Provisioned Throughput SKU adds a second commitment axis, letting steady high-volume buyers reserve a fixed per-PTU-minute throughput envelope instead of reasoning in GPU-hours.

  3. Non-token SKUs (Code Sandbox, Code Interpreter) expand TAM beyond token-only inference buyers. As agentic workflows scale, code-execution sandbox SKUs are becoming table stakes for inference platforms targeting AI-native teams.

  4. Academic-lab founder credibility is a defensible trust anchor. Stanford CRFM and ML lab credentials extend customer trust from founder vision to model curation, optimization claims, and platform design — a trust multiplier competitors cannot replicate without acquiring similar talent.

  5. Cached input discount is becoming table stakes for serverless inference. Together now publishes cached-input rates on many models (DeepSeek V4 Pro $0.20, GLM-5.1 $0.26) — closing a gap it previously had versus Fireworks, OpenAI, and Anthropic — and new listings arrive with the rate already attached (Inkling shipped 2026-07-21 at $0.17 cached), so the remaining coverage gap is a legacy backfill (Llama 3.3 70B, gpt-oss) rather than a policy choice.


UBP implications

  1. Pricing page transparency converts more self-serve revenue than docs-routing — and a published card is only as trustworthy as its most ambiguous cell. Usage-based platforms should default to inline per-SKU per-model rate display; docs-routing is acceptable for advanced SKUs (specialized fine-tuning, custom enterprise terms) but should not be the default for the top-traffic surfaces. Together’s 2026-07-21 catalog update illustrates the failure mode: a model listed at 0.00 input with a blank output column reads as any of “free,” “promotional,” or “not yet priced,” and a buyer who cannot tell those apart escalates to sales for exactly the SKU the self-serve card was meant to close. Treat those as three distinct published states, and label the meter at the row rather than the tab when a model bills on more than one unit.

  2. Multi-window commitment pricing — plus throughput-unit reservation — captures buyer segments that on-demand-or-annual binaries miss. The 7–30 day reserved tier is the canonical structure for training and large-batch workloads where annual commits over-allocate and on-demand under-allocates; Together’s July 2026 PTU-minute SKU extends the same logic to steady inference, letting buyers reserve a fixed throughput envelope benchmarked against a closed-frontier API’s list price. The design’s honesty depends on the sizing calculator modeling real duty cycles, not only continuous 24/7 provisioning. The first increase in a reserved tenor since Together’s aggressive June 2026 cuts (2026-07-29, +2.8% on the 7–30 day H100 rate) shows that even a structurally cheap capacity provider finds — and corrects off of — a pricing floor; it’s not evidence the cost-leadership thesis is reversing, since on-demand and every H200/B200 line held flat in the same cycle.

  3. Non-token usage SKUs (vCPU-hour, GiB-hour, per-session) are becoming necessary for inference platforms targeting agentic workflows. Token-only rate cards leave code-execution and sandbox workloads on the table for third-party vendors that customers prefer to consolidate.


Sources


Bottom line

Together AI priced its AI Acceleration Cloud around four structural ideas: inline per-model rate publication on the pricing page (best transparency in the category), aggressive reserved cluster pricing at $3.69/hr H100 (down to a $3.19/hr floor after a modest 2026-07-29 uptick) with a 7–30 day commit window (still among the lowest published in managed inference), non-token SKUs (Code Sandbox, Code Interpreter) that capture agentic workloads without third-party vendors, and Stanford CRFM + Stanford ML founder credibility that distinguishes Together from pure-engineering platforms. The multi-mode cluster pricing (on-demand, 7–30 day reserved, annual commit) and five-SKU rate card (token / PTU-minute / image / GPU-hour / vCPU-hour, after the July 2026 Provisioned Throughput launch) make Together one of the most expansive usage-based platforms in AI infrastructure.

For AI engineering teams running training cycles, large-batch inference, and agentic code execution at scale, Together is the most legible commercial platform — and the reserved H100 rate (cut twice within a week in June 2026 to a $3.09/hr floor, then nudged back up to $3.69/hr on 7–30 day terms and a $3.19/hr floor on 2026-07-29, the first increase all year) is itself a structural cost advantage even after that correction. The remaining gaps (partial cached-input coverage on serverless, no A100 on the headline rate card, specialized fine-tuning tier budget uncertainty, an unflagged character-to-minute unit change on two speech-to-text models, and one catalog line added on 2026-07-21 whose output cell is blank on the rate card) are competitive parity and presentation issues rather than structural pricing flaws — the 2026-07-29 GPU Cluster increase was confined to reserved-H100 tenors (2.8–4.9%) while on-demand and every H200/B200 rate held flat, so the cost-leadership thesis above stands substantially intact.

Compare with peers via the blueprint corpus, or model your own spend with the Together AI pricing calculator.

Pricing timeline : Major events on a vertical axis

Each milestone below corresponds to a public pricing change, product launch, or material adjustment. Major events use a filled marker; minor adjustments use a faded one.

Compute rate card and full model catalog flat again; Specialized fine-tuning tab re-actuated after a two-week capture gap and shows catalog rotation, not repricing

A full per-tab re-capture (Chat, Image, Audio, Video, Transcribe) found every price on Together's live pricing page — the entire compute rate card (GPU Clusters, Dedicated Inference, Provisioned Throughput, Code Sandbox, Storage, standard fine-tuning) plus every serverless chat, image, speech, and video row — byte-identical to the 2026-08-25 capture; no model was added, removed, or repriced on the live page itself. A corrected tab selector also re-actuated the Specialized fine-tuning tab for the first time since 2026-08-12 (it had been silently landing on the default Standard tab in every capture from 2026-08-13 through 2026-08-25). The re-capture shows catalog rotation, not repricing: three rows from 2026-08-12 are gone (Qwen3-235B-A22B, GLM-4.6/GLM-4.7, Qwen3-Coder-480B- A35B-Instruct), one new row appeared (DeepSeek-V4 Flash 0731 at $6.00/$15.00/$12.00), and three rows picked up expanded aliases at their exact prior price ("DeepSeek-R1 / V3" → "DeepSeek-V3.1", "Kimi K2" → "Kimi K2.6 / Kimi K2.7-Code", "GLM-5 / GLM-5.1" → "GLM-5.2 / GLM-5.1 / GLM-5") — every surviving SKU's rate held flat. Because the tab went unverified for two weeks, this rotation could have occurred any day in that window; 2026-08-26 is the first-confirmed date, not necessarily the effective date. Separately, the second-source docs catalog (docs.together.ai/docs/inference-models) now lists roughly a dozen video/ image models not yet visible on the live pricing page — Sora 2 Pro ($2.40/video), three Google Veo 3.1 variants ($0.05–$0.08/video), three Wan 2.7 T2V/I2V/R2V models (~$0.10/video), two Vidu Q3 variants ($0.0975–$0.195/video), PixVerse v5.6 and v6 (~$0.09–$0.13/video), ByteDance Seedream 5.0 Lite ($0.035/MP), Google Gemini 3.1 Flash-Lite Image ($0.069/image), and Pruna AI P-Image-Ideogram ($0.00225/MP) — none confirmed billable until they surface on the live rate card. A new docs-vs-page conflict also surfaced: the docs table prices Vidu 2.0 at $0.80/video (720p/8s) versus $0.28/video on the live pricing page — a second Qwen-Image-style discrepancy between the two surfaces, with the live page remaining the priced-against source of record per this page's standing practice.

Compute rate card and full model catalog flat again; Specialized fine-tuning tab re-actuated after a two-week capture gap and shows catalog rotation, not repricing - A full per-tab re-capture (Chat, Image, Audio, Video, Transcribe) found every pr
captured

Compute rate card flat again; DeepSeek V4 Pro 0813 + 6 HappyHorse video models added; GLM-5.1 and Kimi K2.6 rotated off the catalog

A full per-tab re-capture (Chat, Image, Audio, Video, Transcribe, Specialized Fine-Tuning) plus a second-source re-check of Together's serverless model-catalog docs found GPU Clusters, Dedicated Inference, Provisioned Throughput, Code Sandbox, Storage, and the standard fine-tuning tier all still at their 2026-07-29 rates. Catalog movement: DeepSeek V4 Pro 0813 joined at $1.32 input / $3.96 output per 1M tokens ($0.13 cached), alongside the unchanged original DeepSeek V4 Pro ($1.74/$3.48); DeepSeek V4 Flash 0731 and Qwen3 235B A22B Instruct 2507 FP8 Throughput, previously named only in prose, are now confirmed on the rate card; and GLM-5.1 and Kimi K2.6 are no longer listed (catalog rotation, not a repricing of any surviving model). Six new Alibaba HappyHorse video models joined at $0.24/video (1.0 T2V/I2V/R2V) and $0.14/video (1.1 T2V/I2V/R2V), and Kling 2.1 Pro's exact rate ($0.32/video) is now confirmed. NVIDIA HGX B300 appeared as its own row in the GPU Clusters table (Contact us, matching GB200/GB300 NVL72). One surface move: NVIDIA Nemotron 3 ASR Streaming 0.6B no longer appears on the Audio (streaming, per-1M-character) tab; it is now listed only on the Transcribe (batch, per-audio-minute) tab, still at $0.0015 but on a different meter — the number held, the billing unit moved. The Specialized fine-tuning tab did not actuate during this capture cycle (rendered the default Chat/Standard-fine-tuning state), so the specialized per-model fine-tuning tier was not independently re-verified this cycle and is carried forward unchanged from 2026-08-12.

Compute rate card flat again; DeepSeek V4 Pro 0813 + 6 HappyHorse video models added; GLM-5.1 and Kimi K2.6 rotated off the catalog - A full per-tab re-capture (Chat, Image, Audio, Video, Transcribe, Specialized Fi
captured

Compute rate card flat a fifth straight capture; Cogito v2.1 671B, Rnj-1 Instruct, and Inkling Small's dual-meter billing confirmed

A full per-tab re-capture (Chat, Image, Audio, Video, Transcribe, Specialized Fine-Tuning) confirmed GPU Clusters, Dedicated Inference, Provisioned Throughput, Code Sandbox, Storage, and both fine-tuning tiers all held their 2026-07-29 rates again — every price on every tab matched the prior capture exactly. The only movement was catalog disclosure: two chat models named in prior updates without a price (Cogito v2.1 671B, Rnj-1 Instruct) now show confirmed rates ($1.25/$1.25 and $0.15/$0.15 respectively), and two more models — Qwen2.5 7B Instruct Turbo ($0.30/$0.30, one of the six Batch-discounted models) and Llama 3 8B Instruct Lite ($0.14/$0.14) — surfaced with prices for the first time. The more notable find: Inkling Small, previously recorded only as a speech-catalog addition ($0.50 per 1M characters), was also confirmed on the Chat table at $0.50 input / $1.20 output per 1M tokens ($0.10 cached) — it bills on two meters depending on modality, the same dual-meter pattern already documented for the original Inkling model.

Compute rate card flat a fifth straight capture; Cogito v2.1 671B, Rnj-1 Instruct, and Inkling Small's dual-meter billing confirmed - A full per-tab re-capture (Chat, Image, Audio, Video, Transcribe, Specialized Fi
captured

Compute rate card flat a fourth straight week; Qwen3.8-2.4T-A95B and Inkling Small added

Together's full compute rate card — GPU Clusters, Dedicated Inference, Provisioned Throughput (including the per-PTU capacity raised on 2026-08-12), Code Sandbox, Storage, and both fine-tuning tiers — confirmed unchanged for a fourth consecutive week via fresh per-tab captures of Chat, Image, Audio, Video, Transcribe, and Specialized Fine-Tuning. The only rate-card movement was two new serverless catalog additions: Qwen3.8-2.4T-A95B joined Chat at $2.50 input / $6.25 output per 1M tokens ($0.50 cached), and Inkling Small joined the per-1M-character speech catalog at $0.50 (alongside the existing Inkling line at $1.00). Also newly confirmed at exact prices via a full Image-tab capture: Ideogram 4.0 ($0.06/MP), Gemini 3.1 Flash Image "Nano Banana 2" ($0.05/image), Wan 2.6 Image ($0.03/image), and several Google Imagen 4.0 / ByteDance Seedream variants, none of which moved a previously-published rate. One oddity flagged for verification: the Video tab lists "Qwen3.6-Plus" at $0.50/video — the same model name already priced at $0.50 input / $3.00 output per 1M tokens on the Chat tab — which reads as a likely rate-card data artifact (a chat-model name reused on the video table) rather than a real per-video SKU; treat the $0.50/video figure as unconfirmed until Together's docs or a follow-up capture clarifies it.

Compute rate card flat a fourth straight week; Qwen3.8-2.4T-A95B and Inkling Small added - Together's full compute rate card — GPU Clusters, Dedicated Inference, Provision
captured

Provisioned Throughput capacity increase + Kimi K3 added; PrismML confirmed Free

Together added Kimi K3 as a third Provisioned Throughput model and increased published per-PTU capacity for MiniMax M3 (input 138,840→ 166,667 and cached 694,200→833,333 TPM/PTU, +20% each; output 23,140→ 41,667 TPM/PTU, +80%) and GLM-5.2 (output 9,620→11,364 TPM/PTU, +18%), while holding the $0.05/PTU-minute sticker price flat for all three models — an effective cut in cost per unit of reserved throughput. Together's model-catalog docs also newly confirmed PrismML Ternary Bonsai 27B as fully "Free" (input and output), resolving the ambiguous blank-output cell flagged since 2026-07-21. Every other compute meter (GPU Clusters, Dedicated Inference, Code Sandbox, Storage, both fine-tuning tiers) held its 2026-07-29 rate for a third consecutive week; the specialized fine-tuning tier gained three new per-model rows (Llama 4 Maverick, Qwen3-Coder-480B-A35B-Instruct, Qwen3.5-122B-A10B) and the chat/image/video catalogs continued growing without moving any other previously-published rate.

Provisioned Throughput capacity increase + Kimi K3 added; PrismML confirmed Free - Together added Kimi K3 as a third Provisioned Throughput model and increased pub
captured

GPU Cluster reserved H100 rates rise (first increase since June cuts) + speech-to-text unified to $0.0015/audio-minute

Together's GPU Cluster reserved H100 rate rose across all three commitment tenors for the first time since the back-to-back June 2026 cuts: 7–30 day $3.59→$3.69/hr, 31–90 day $3.29→$3.45/hr, and the 91–180 day floor $3.09→$3.19/hr. On-demand H100 and every H200/B200 on-demand and reserved rate held flat, so the increase is isolated to reserved-H100 tenors. In the same capture cycle, all four serverless speech-to-text models (Whisper Large v3, Parakeet TDT 0.6B v3, and both Nemotron 3 and 3.5 ASR Streaming variants) were unified onto a single $0.0015-per-audio-minute meter — a 67% cut on Nemotron 3.5 ASR (from $0.0045/min) and a billing-unit change (from per-1M-characters to per-audio-minute) for Parakeet TDT 0.6B and Nemotron 3 ASR. Four new serverless chat models (Kimi K3, Gemma 4 31B, Qwen3.7-Plus, LFM2.5-8B-A1B) joined the catalog without moving any previously-published rate.

GPU Cluster reserved H100 rates rise (first increase since June cuts) + speech-to-text unified to $0.0015/audio-minute - Together's GPU Cluster reserved H100 rate rose across all three commitment tenor
captured

Kokoro-82M TTS cut 60% + two chat models added; compute rate card flat

Together's compute meters held flat week-over-week — dedicated HGX H100 at $5.49/hr and HGX B200 at $8.99/hr, GPU clusters at $3.99/hr on-demand H100 and $3.59/hr reserved, PTU at $0.05/PTU-min, Code Sandbox, storage, and both fine-tuning tiers all unchanged. The movement was entirely in the model catalog: Kokoro-82M TTS was cut 60% from $10.00 to $4.00 per 1M characters (against Orpheus at $15 and Cartesia Sonic-2/-3 unchanged at $65.00, stretching the TTS spread to ~16×), Inkling was added at $1.00 input / $4.05 output per 1M tokens with a $0.17 cached-input rate and a parallel $1.00 per-1M-character speech line, and PrismML Ternary Bonsai 27B was listed at an input rate of 0.00 with no output rate published.

Kokoro-82M TTS cut 60% + two chat models added; compute rate card flat - Together's compute meters held flat week-over-week — dedicated HGX H100 at $5.49
captured

Provisioned Throughput (PTU) launch + Dedicated Inference restructure & price cuts

Together launched Provisioned Throughput — a new SKU that reserves dedicated capacity in throughput units (PTUs) billed per PTU-minute ($0.05/PTU-min on MiniMax M3 and GLM-5.2), with an on-page calculator that sizes PTUs and estimates monthly cost vs. commercial-model list prices. In the same update, Dedicated Inference was restructured to a per-GPU-per-hour table split into on-demand (pay-as-you-go) vs reserved (Contact sales) columns: on-demand HGX H100 fell to $5.49/hr (from $6.49) and HGX B200 to $8.99/hr (from $11.95), and NVIDIA HGX H200, HGX B300, GB200 NVL72, and GB300 NVL72 lines were added (quoted "Contact us"). GPU Cluster and serverless rates were unchanged. The pricing page also carries a new Series C funding banner.

Provisioned Throughput (PTU) launch + Dedicated Inference restructure & price cuts - Together launched Provisioned Throughput — a new SKU that reserves dedicated cap
captured

$800M raise at an $8.3B valuation (IPO not ruled out)

Together raised $800M at an $8.3B valuation (Reuters, TechCrunch, and Business Wire, 1 July 2026), with the CEO declining to rule out an IPO the following year (Axios, 2 July 2026). No rate move accompanied the raise, but it is the balance-sheet backdrop for the sustained 2026 GPU and dedicated-inference rate cuts this timeline tracks: a war chest of this size funds continued price leadership rather than a pivot to margin recovery. The Series C funding banner that appeared on the pricing page on 2026-07-14 is the on-page marker of this round.

Reserved + on-demand GPU cluster rate cut (H100 down 12–16%)

Together cut its GPU Cluster rates again. On-demand HGX H100 fell to $3.99/hr (from $4.79); reserved 7–30 day H100 fell to $3.59/hr (from $4.19), 31–90 day to $3.29 (from $3.45), and 91–180 day to $3.09/hr (from $3.29) — making the reserved H100 floor $3.09/hr. On-demand H200/B200 and reserved H200/B200 rates were unchanged. A 1× H200 140GB dedicated-endpoint line was added (priced "Contact us"). The standard fine-tuning tier now states a $4.00 per-job minimum charge.

Reserved + on-demand GPU cluster rate cut (H100 down 12–16%) - Together cut its GPU Cluster rates again. On-demand HGX H100 fell to $3.99/hr (f
captured

Serverless re-pricing + GPU cluster rate cuts + cached input

Together repriced its serverless rate card and cut GPU cluster rates. DeepSeek V4 Pro dropped to $1.74/$3.48 (from $2.10/$4.40) and now shows a $0.20 cached-input rate; Qwen3.5 9B rose to $0.17/$0.25 (from $0.10/$0.15); Llama 3.3 70B rose to $1.04/$1.04 (from $0.88/$0.88). On-demand cluster H100 fell to $4.79/hr (from $5.49) and B200 to $8.19/hr (from $9.95); reserved 7–30 day H100 fell to $4.19/hr (from $4.99) and B200 to $7.99/hr (from $9.65), with H100 as low as $3.29/hr on a 91–180 day reservation. Cached-input pricing is now published on serverless models.

Serverless re-pricing + GPU cluster rate cuts + cached input - Together repriced its serverless rate card and cut GPU cluster rates. DeepSeek V
captured

Specialized Model Fine-Tuning Tier

Together added a specialized model fine-tuning tier for DeepSeek-R1, GLM-5, and other large-context models at $10–$100+ per 1M training tokens with $20–$60 per-job minimum charges. Reflected the higher infrastructure cost of training on the latest frontier architectures.

Specialized Model Fine-Tuning Tier screenshot 1
Specialized Model Fine-Tuning Tier screenshot 2

FLUX.2 Image Generation at $0.0154/image

Together added FLUX.2 [dev] image generation at $0.0154/image, FLUX.1 [schnell] at $0.0027/image, and Stable Diffusion 3 at $0.0019/image. Per-image pricing positioned alongside per-token text inference as a unified rate card.

Batch API + Code Sandbox Launched

Together launched Batch API (50% discount on most models) for asynchronous inference workloads, and Code Sandbox ($0.0446/vCPU-hour, $0.0149/GiB-hour) for agentic code execution. Code Interpreter at $0.03/session added a session-billed SKU to the rate card.

Series B ($305M) at $3.3B Valuation

Together raised a $305M Series B led by General Catalyst at a $3.3B post-money valuation. NVIDIA, Salesforce Ventures, Coatue, and others participated. The round funded Code Sandbox, FLUX image generation, and the Together AI Acceleration Cloud rebrand.

GPU Clusters at $5.49/hr H100 On-Demand

Together launched GPU Clusters — on-demand multi-GPU rentals for training and large-batch inference. Pricing at $5.49/hr H100 on-demand and $9.95/hr B200 undercut Fireworks' and Baseten's dedicated rates substantially. Reserved 7–30 day commitments dropped the H100 rate to $4.99/hr.

Dedicated Endpoints + Fine-Tuning Launched

Together added dedicated single-tenant endpoints (per-hour H100, A100) and a fine-tuning service. Established the multi-SKU architecture: serverless per-token + dedicated per-hour + fine-tuning per training-token that remains the canonical structure today.

Series A ($102.5M) + Inference Cloud GA

Together raised a $102.5M Series A led by Kleiner Perkins with NEA, Lux, and others. Inference Cloud went GA with per-token serverless API for Llama 2, Falcon, Code Llama, and Stable Diffusion. Initial pricing was a flat per-million-token rate by model class.

Together Founded

Vipul Ved Prakash (ex-Topsy, Cloudmark) co-founded Together with Ce Zhang (ETH Zurich), Chris Re (Stanford), and Percy Liang (Stanford CRFM director). Initial product was decentralized GPU pooling for open-source model training, evolving rapidly into a managed inference cloud through 2023.

Trivia
  • · Together AI's $3.69/hr H100 reserved rate (7–30 day reservation, dropping to $3.19/hr on a 91–180 day commit) is one of the lowest published rates for any managed Hopper-class GPU — and the $7.99/hr reserved B200 sets a similar floor for Blackwell, both undercutting Fireworks' on-demand rates.
  • · Together was co-founded by Vipul Ved Prakash (ex-Cloudmark, Topsy founder), Ce Zhang (ETH Zurich systems professor), Chris Re (Stanford ML/Snorkel), and Percy Liang (Stanford CRFM director) — making it the rare commercial product where two top academic ML labs are co-architects of the platform.
  • · Together's serverless rate card publishes per-model pricing inline on the pricing page (rare among competitors like Fireworks which route to docs), making per-model side-by-side comparison friction-free.

Questions & answers

How much does Together AI cost per month?
Together has no monthly subscription fee — you pay only for the serverless tokens, dedicated GPU hours, fine-tuning training tokens, and Code Sandbox usage you consume. A small RAG application using Llama 3.3 70B at 30M input + 10M output tokens would cost ~$42/month on serverless; the same workload on a dedicated H100 ($5.49/hr) running 4h/day would cost ~$660/month.
What are Together's serverless per-token rates?
Together publishes per-model rates inline on the pricing page. Sample rates per 1M tokens: Llama 3.3 70B at $1.04 input / $1.04 output; DeepSeek V4 Pro at $1.74 input / $3.48 output ($0.20 cached input); Qwen3.5 9B at $0.17 input / $0.25 output; GLM-5.1 at $1.40 input / $4.40 output; Inkling, added on 2026-07-21, at $1.00 input / $4.05 output ($0.17 cached input). One line reads as free: PrismML Ternary Bonsai 27B, also added on 2026-07-21, shows an input rate of 0.00 on the pricing card with its output cell left blank, while Together's model docs publish the same model at 0.00 input and 0.00 output — so confirm its terms with Together before building on it. Image generation is metered per image or per megapixel: FLUX.2 [dev] at $0.0154 per image, FLUX.1 [schnell] at $0.0027 per megapixel, SD XL at $0.0019 per megapixel — and all quoted media prices assume the lowest resolution/duration settings. Non-text models use their own meters: text-to-speech bills per 1M characters (Cartesia Sonic-2 and Sonic-3 at $65.00, Orpheus TTS at $15, Kokoro-82M at $4.00 after a 60% cut on 2026-07-21), transcription per audio minute — all four serverless speech-to-text models (Whisper Large v3, Parakeet TDT 0.6B v3, and both Nemotron 3 and 3.5 ASR Streaming variants) were unified onto a flat $0.0015/min rate on 2026-07-29 — and video per video (Google Veo 3.0 at $1.60, Sora 2 at $0.80).
What are Together's GPU rates for dedicated endpoints and clusters?
Dedicated inference (per GPU per hour, on-demand): HGX H100 at $5.49/hr, HGX B200 at $8.99/hr (H200/B300/GB200/GB300 quoted "Contact us"; all reserved dedicated capacity is "Contact sales"). On-demand GPU clusters: HGX H100 at $3.99/hr, HGX H200 at $5.99/hr, HGX B200 at $8.19/hr. Reserved clusters (7–30 day commits): H100 at $3.69/hr (up from $3.59 on 2026-07-29), H200 at $4.99/hr, B200 at $7.99/hr, dropping to $3.19/hr H100 on 91–180 day reservations — the reserved rates are among the lowest published in managed inference. Separately, a July 2026 Provisioned Throughput (PTU) SKU reserves capacity in throughput units at $0.05 per PTU-minute (MiniMax M3, GLM-5.2) for buyers who prefer a fixed tokens-per-minute envelope over per-GPU-hour rentals.
Does Together AI have a free tier?
New accounts can start without an upfront commitment, but Together does not publish a specific signup-credit dollar amount on its pricing page or quickstart docs. Calling paid serverless and image models requires a positive credit balance, and production usage requires a payment method on file. There is no permanent free tier.
How does Together's fine-tuning pricing work?
Fine-tuning is priced per 1M training tokens (LoRA vs full-parameter). Standard tier: up to 16B at $0.48 SFT LoRA / $1.20 full; 17–69B at $1.50 / $3.75; 70–100B at $2.90 / $7.25. A specialized per-model tier covers frontier architectures (DeepSeek-R1 $10 SFT LoRA, GLM-5 $40, Qwen3.5-397B $8, Llama 4 Scout $3), ranging roughly $3–$40 SFT LoRA and up to $100 DPO LoRA per 1M tokens. The standard tier has no per-job minimum, while the specialized tier carries per-model minimum charges of $6–$60 (Qwen3-235B is the exception with no minimum).
What is Together's Code Sandbox and how is it priced?
Code Sandbox is a managed code-execution environment for agentic workflows. Billed at $0.0446 per vCPU-hour and $0.0149 per GiB-hour for the sandbox runtime. Code Interpreter (a higher-level managed session API) bills at $0.03/session. Storage attached to sandbox sessions is $0.16/GiB-month.