Packaging

Baseten trims Model APIs catalog to eight models and adds Inkling

Baseten pricing

Baseten cut its Model APIs rate card from 11 models to 8, retiring GLM 5.1, GLM 5, Kimi K2.5 and Nemotron 3 Super, and added Inkling at $1.00/$4.05 per 1M tokens.

Before

11 Model APIs SKUs: GLM 5.2, GLM 5.1, GLM 5, GLM 4.7, Kimi K2.7 Code, Kimi K2.6, Kimi K2.5, NVIDIA Nemotron 3 Ultra, NVIDIA Nemotron 3 Super, DeepSeek V4, GPT OSS 120B.

After

8 Model APIs SKUs: Inkling ($1.00 in / $0.17 cache / $4.05 out), GLM 5.2, GLM 4.7, Kimi K2.7 Code, Kimi K2.6, NVIDIA Nemotron 3 Ultra, DeepSeek V4, GPT OSS 120B.

Between the 2026-07-06 and 2026-07-21 captures of baseten.co/pricing, four models disappeared from the Model APIs rate card in a single sweep — GLM 5.1 ($1.30 in / $4.30 out), GLM 5 ($0.95 / $3.15), Kimi K2.5 ($0.60 / $3.00) and NVIDIA Nemotron 3 Super ($0.30 / $0.75) — while Inkling was added at $1.00 input / $0.17 cache input / $4.05 output per 1M tokens.

The retirements pull the cheap end of the catalog upward: with Nemotron 3 Super gone, the lowest published input rate outside GPT OSS 120B ($0.10) is now GLM 4.7 and Nemotron 3 Ultra at $0.60, and the cheapest published cache-input rate rises from $0.06 to $0.12. Baseten documents a deprecation policy for Model APIs, so the sweep reads as scheduled catalog pruning rather than a repricing.

No other pricing surface moved. Dedicated deployment GPU rates (T4 $0.01052/min through B200 $0.16633/min), CPU instance rates, the on-demand Training rate card, and the Basic / Pro / Enterprise tier structure are all unchanged from the prior capture.

From Baseten's pricing timeline
Model APIs Catalog Pruned to Eight Models; Inkling Added

Baseten cut its Model APIs rate card from eleven SKUs to eight, retiring GLM 5.1, GLM 5, Kimi K2.5 and NVIDIA Nemotron 3 Super, and added Inkling at $1.00 input / $0.17 cache input / $4.05 output per 1M tokens. No surviving model's rate changed, but the retirements removed the cheap end of the catalog: the lowest published cache-input rate doubled from $0.06 to $0.12. Dedicated GPU, CPU, Training and Basic/Pro/Enterprise surfaces were unchanged.

About Baseten
baseten.co ↗

Baseten runs a pure-usage GPU-minute billing model for dedicated model deployments plus a separate per-token Model APIs catalog — both pay-as-you-go from the Basic tier with no monthly minimum.

Free tier
Yes
Commits
Available
Transparency
public

Baseten pricing history

  1. Jul 2026
    Model APIs Catalog Pruned to Eight Models; Inkling Added
  2. Feb 2026
    Cached Input Pricing on Model APIs
  3. Sep 2025
    Self-Hosted (BYOC) Deployment Option
  4. Mar 2025
    B200 GPU Availability + Mission Critical SLA Tier
  5. Nov 2024
    H100 GPU Pricing Published at $0.10833/min
Full Baseten timeline

More Baseten activity

All pricing activity