Baseten trims Model APIs catalog to eight models and adds Inkling
Baseten cut its Model APIs rate card from 11 models to 8, retiring GLM 5.1, GLM 5, Kimi K2.5 and Nemotron 3 Super, and added Inkling at $1.00/$4.05 per 1M tokens.
11 Model APIs SKUs: GLM 5.2, GLM 5.1, GLM 5, GLM 4.7, Kimi K2.7 Code, Kimi K2.6, Kimi K2.5, NVIDIA Nemotron 3 Ultra, NVIDIA Nemotron 3 Super, DeepSeek V4, GPT OSS 120B.
8 Model APIs SKUs: Inkling ($1.00 in / $0.17 cache / $4.05 out), GLM 5.2, GLM 4.7, Kimi K2.7 Code, Kimi K2.6, NVIDIA Nemotron 3 Ultra, DeepSeek V4, GPT OSS 120B.
Between the 2026-07-06 and 2026-07-21 captures of baseten.co/pricing, four models disappeared from the Model APIs rate card in a single sweep — GLM 5.1 ($1.30 in / $4.30 out), GLM 5 ($0.95 / $3.15), Kimi K2.5 ($0.60 / $3.00) and NVIDIA Nemotron 3 Super ($0.30 / $0.75) — while Inkling was added at $1.00 input / $0.17 cache input / $4.05 output per 1M tokens.
The retirements pull the cheap end of the catalog upward: with Nemotron 3 Super gone, the lowest published input rate outside GPT OSS 120B ($0.10) is now GLM 4.7 and Nemotron 3 Ultra at $0.60, and the cheapest published cache-input rate rises from $0.06 to $0.12. Baseten documents a deprecation policy for Model APIs, so the sweep reads as scheduled catalog pruning rather than a repricing.
No other pricing surface moved. Dedicated deployment GPU rates (T4 $0.01052/min through B200 $0.16633/min), CPU instance rates, the on-demand Training rate card, and the Basic / Pro / Enterprise tier structure are all unchanged from the prior capture.
Baseten cut its Model APIs rate card from eleven SKUs to eight, retiring GLM 5.1, GLM 5, Kimi K2.5 and NVIDIA Nemotron 3 Super, and added Inkling at $1.00 input / $0.17 cache input / $4.05 output per 1M tokens. No surviving model's rate changed, but the retirements removed the cheap end of the catalog: the lowest published cache-input rate doubled from $0.06 to $0.12. Dedicated GPU, CPU, Training and Basic/Pro/Enterprise surfaces were unchanged.