Together AI boosts Provisioned Throughput capacity, adds Kimi K3
Together AI added Kimi K3 as a third Provisioned Throughput model and raised per-PTU capacity 18-80% for MiniMax M3 and GLM-5.2 while holding the $0.05/PTU-minute sticker price flat.
Provisioned Throughput covered MiniMax M3 (138,840 input / 694,200 cached / 23,140 output TPM/PTU) and GLM-5.2 (35,731 / 192,400 / 9,620 TPM/PTU), both at $0.05/PTU-minute
Provisioned Throughput covers MiniMax M3 (166,667 / 833,333 / 41,667 TPM/PTU, +20-80%), GLM-5.2 (35,714 / 192,308 / 11,364 TPM/PTU, output +18%), and new third model Kimi K3 (16,667 / 166,667 / 3,333 TPM/PTU) — all still at $0.05/PTU-minute
Together’s Provisioned Throughput SKU — launched in July 2026 as a fixed tokens-per-minute reservation billed per PTU-minute — gained its third supported model on 2026-08-12: Kimi K3 joins MiniMax M3 and GLM-5.2 on the on-page PTU sizing calculator and its underlying compute-costs reference table, all still priced at $0.05 per PTU-minute.
The more consequential move is a capacity increase on the two existing models. MiniMax M3’s published per-PTU capacity rose from 138,840 to 166,667 input tokens-per-minute and from 694,200 to 833,333 cached tokens-per-minute (both +20%), and from 23,140 to 41,667 output tokens-per-minute (+80%). GLM-5.2’s output capacity rose from 9,620 to 11,364 tokens-per-minute (+18%), with input and cached capacity effectively unchanged. Because the $0.05/PTU-minute sticker price did not move, a buyer reserving the same number of PTUs now gets meaningfully more throughput for the same bill — an effective price cut on the unit economics of reserved capacity, even though the headline rate card shows no change. Every other compute meter (GPU Clusters, Dedicated Inference, Code Sandbox, Storage, both fine-tuning tiers) held its 2026-07-29 rate for a third consecutive week, and the chat/image/video serverless catalogs continued growing without moving any other previously-published rate.
Together added Kimi K3 as a third Provisioned Throughput model and increased published per-PTU capacity for MiniMax M3 (input 138,840→ 166,667 and cached 694,200→833,333 TPM/PTU, +20% each; output 23,140→ 41,667 TPM/PTU, +80%) and GLM-5.2 (output 9,620→11,364 TPM/PTU, +18%), while holding the $0.05/PTU-minute sticker price flat for all three models — an effective cut in cost per unit of reserved throughput. Together's model-catalog docs also newly confirmed PrismML Ternary Bonsai 27B as fully "Free" (input and output), resolving the ambiguous blank-output cell flagged since 2026-07-21. Every other compute meter (GPU Clusters, Dedicated Inference, Code Sandbox, Storage, both fine-tuning tiers) held its 2026-07-29 rate for a third consecutive week; the specialized fine-tuning tier gained three new per-model rows (Llama 4 Maverick, Qwen3-Coder-480B-A35B-Instruct, Qwen3.5-122B-A10B) and the chat/image/video catalogs continued growing without moving any other previously-published rate.