Ask
Launch

Baseten adds GLM-5.3-Flash and DeepSeek V4 Pro 0813 to Model APIs catalog

Baseten pricing

Baseten grew its Model APIs catalog from 12 to 14 models, adding GLM-5.3-Flash ($0.15/$0.03/$0.50 per 1M tokens) and DeepSeek V4 Pro 0813 ($1.32/$0.132/$3.96); no existing SKU's rate changed.

Before

12 Model APIs SKUs, cheapest input rate $0.10 (GPT OSS 120B) to $0.13 (DeepSeek-V4-Flash-0731); DeepSeek V4 Pro at $1.74 in / $0.145 cache / $3.48 out was the only DeepSeek V4-family SKU.

After

14 Model APIs SKUs adding GLM-5.3-Flash ($0.15 in / $0.03 cache / $0.50 out) and DeepSeek V4 Pro 0813 ($1.32 in / $0.132 cache / $3.96 out, a lower-input/higher-output dated variant of DeepSeek V4 Pro); all 12 prior SKUs unchanged.

Three weeks after renaming DeepSeek V4 to DeepSeek V4 Pro alongside the Flash and Inkling-Small additions, Baseten grew its Model APIs catalog again — from twelve models to fourteen. GLM-5.3-Flash launched at $0.15 input / $0.03 cache input / $0.50 output per 1M tokens, promoted by a new homepage banner (“Try the new GLM-5.3 Flash today. Frontier intelligence at a fraction of the cost.”) that replaced the prior “Try the new DeepSeek V4 Flash today” promo — the second consecutive banner cycle built around a new cheap-tier Model API launch rather than a funding milestone.

DeepSeek V4 Pro 0813 joined alongside it as a second, dated DeepSeek V4 variant: its $1.32 input rate is 24% below the original DeepSeek V4 Pro’s $1.74, but its $3.96 output rate is 14% above DeepSeek V4 Pro’s $3.48 — the first Model APIs addition on this rate card to cut one rate while raising another on the same SKU family. No rate on any of the twelve prior SKUs moved, and the dedicated GPU, CPU, Training rate cards and the Basic/Pro/Enterprise tier structure were all unchanged.

From Baseten's pricing timeline
Model APIs Catalog Grows to Fourteen Models

Three weeks after the DeepSeek V4 Pro rename, Baseten added GLM-5.3-Flash ($0.15 input / $0.03 cache input / $0.50 output per 1M tokens) and DeepSeek V4 Pro 0813 ($1.32 / $0.132 / $3.96 — a dated DeepSeek V4 Pro variant with a 24% lower input rate but 14% higher output rate). No rate on any of the twelve prior SKUs changed, and the homepage banner switched from promoting DeepSeek V4 Flash to promoting GLM-5.3-Flash. Dedicated GPU/CPU/Training rate cards and the Basic/Pro/Enterprise tier structure were unchanged.

About Baseten
baseten.co ↗

Baseten runs a pure-usage GPU-minute billing model for dedicated model deployments plus a separate per-token Model APIs catalog — both pay-as-you-go from the Basic tier with no monthly minimum.

Free tier
Yes
Commits
Available
Transparency
public

Baseten pricing history

  1. Aug 2026
    Model APIs Catalog Grows to Fourteen Models
  2. Aug 2026
    Model APIs Catalog Grows to Twelve Models; DeepSeek V4 Renamed Pro
  3. Jul 2026
    Model APIs Catalog Grows to Ten Models; GLM-5.2 Cache Rate Cut 46%
  4. Jul 2026
    Model APIs Catalog Pruned to Eight Models; Inkling Added
  5. Feb 2026
    Cached Input Pricing on Model APIs
Full Baseten timeline

More Baseten activity

All pricing activity