Ask
Launch

Baseten adds Kimi K3 and GLM-5.2 Fast, cuts GLM-5.2 cache pricing 46%

Baseten pricing

Baseten grew its Model APIs catalog to ten models with new flagship Kimi K3 and GLM-5.2 Fast, and cut GLM-5.2's cache-input rate 46% to $0.14 per 1M tokens.

Before

8 Model APIs SKUs: Inkling, GLM-5.2 ($1.40 in / $0.26 cache / $4.40 out), GLM 4.7, Kimi K2.7 Code, Kimi K2.6, NVIDIA Nemotron 3 Ultra, DeepSeek V4, GPT OSS 120B.

After

10 Model APIs SKUs adding Kimi K3 ($3.00 in / $0.30 cache / $15.00 out) and GLM-5.2 Fast ($2.10 / $0.21 / $6.60); GLM-5.2 cache input cut to $0.14 (input/output unchanged at $1.40/$4.40).

Eight days after pruning its Model APIs catalog from eleven models to eight, Baseten expanded it again — this time by addition rather than retirement. Kimi K3 launched as the platform’s new flagship model at $3.00 input / $0.30 cache input / $15.00 output per 1M tokens, flagged by a homepage “Kimi K3 is here.” banner, and became the most expensive SKU on the rate card by a wide margin (72% above the prior ceiling, DeepSeek V4, on input; over 3x on output). GLM-5.2 Fast joined alongside it at $2.10 / $0.21 / $6.60, a faster and pricier sibling to the existing GLM-5.2 line.

The same capture recorded a price cut on an existing model: GLM-5.2’s own cache-input rate fell from $0.26 to $0.14 per 1M tokens — a 46% reduction — while its $1.40 input and $4.40 output rates held steady. No other pricing surface moved: dedicated deployment GPU/CPU rates, the on-demand Training rate card, and the Basic/Pro/Enterprise tier structure are all unchanged from the prior capture.

From Baseten's pricing timeline
Model APIs Catalog Grows to Ten Models; GLM-5.2 Cache Rate Cut 46%

Eight days after pruning its Model APIs catalog to eight models, Baseten added flagship Kimi K3 ($3.00 input / $0.30 cache input / $15.00 output per 1M tokens) and GLM-5.2 Fast ($2.10 / $0.21 / $6.60), and cut GLM-5.2's own cache-input rate 46% to $0.14 per 1M tokens (from $0.26; its $1.40 input and $4.40 output rates held steady). Dedicated GPU/CPU/Training rate cards and the Basic/Pro/Enterprise tier structure were unchanged.

About Baseten
baseten.co ↗

Baseten runs a pure-usage GPU-minute billing model for dedicated model deployments plus a separate per-token Model APIs catalog — both pay-as-you-go from the Basic tier with no monthly minimum.

Free tier
Yes
Commits
Available
Transparency
public

Baseten pricing history

  1. Aug 2026
    Model APIs Catalog Grows to Fourteen Models
  2. Aug 2026
    Model APIs Catalog Grows to Twelve Models; DeepSeek V4 Renamed Pro
  3. Jul 2026
    Model APIs Catalog Grows to Ten Models; GLM-5.2 Cache Rate Cut 46%
  4. Jul 2026
    Model APIs Catalog Pruned to Eight Models; Inkling Added
  5. Feb 2026
    Cached Input Pricing on Model APIs
Full Baseten timeline

More Baseten activity

All pricing activity