W&B delists 5 Inference models, adds Nemotron 3.5 Lightning
Weights & Biases removed five models from its Serverless Inference rate card and added one new model, NVIDIA Nemotron 3.5 Lightning, at $0.10 in / $0.05 cached / $0.25 out per 1M tokens — no surviving model's price moved.
32 models on the Serverless Inference rate card (2026-08-11).
28 models on the rate card (2026-08-12): 5 delisted (Qwen3.5 27B, Moonshot Kimi K2.5, Qwen3 235B A22B-2507, Qwen3 235B A22B Thinking-2507, Microsoft Phi 4 Mini 3.8B), 1 added (NVIDIA Nemotron 3.5 Lightning).
After four straight weeks of per-model repricing, Weights & Biases’s Serverless Inference catalog moved again on August 12 — but this time the change was pure roster churn, not a price cut or hike. Five models were delisted from the canonical Token-Based Pricing rate card: Qwen3.5 27B ($0.39 in / $0.08 cached / $3.12 out per 1M tokens), Moonshot AI Kimi K2.5 ($0.60 in / $0.10 cached / $3.00 out), Qwen3 235B A22B-2507 ($0.10 in / $0.10 out), Qwen3 235B A22B Thinking-2507 ($0.10 in / $0.10 out), and Microsoft Phi 4 Mini 3.8B ($0.08 in / $0.35 out). One new model, NVIDIA Nemotron 3.5 Lightning, joined at $0.10 input / $0.05 cached / $0.25 output — landing near the catalog’s cheap end. Net effect: the roster shrank from 32 to 28 models, and every price that survived the cut is byte-identical to the August 11 capture.
The move also flips a running discrepancy between W&B’s two Inference surfaces. The marketing “Available models” gallery had been lagging behind the canonical rate card (missing Z.AI GLM 5, then also dropping Microsoft Phi 4 Mini 3.8B). As of August 12 the gap reversed: the marketing gallery still lists all five delisted models even though the billing-authoritative rate card no longer prices them, so the gallery is now ahead of, not behind, the canonical page. Published cloud tiers (Free $0, Pro from $60/mo, custom Enterprise), storage overage ($0.03/GB), and Weave data ingestion overage ($0.10/MB) were unchanged.
DeepSeek V4-Pro was repriced from $1.74 input / $0.14 cached / $3.48 output to $1.15 input / $0.20 cached / $2.55 output per 1M tokens (roughly -34% input, -27% output, +43% cached), and a new model, DeepSeek V4-Flash-0731, joined the Serverless Inference catalog at $0.13 in / $0.07 cached / $0.28 out (32 models total). Cloud tiers (Free / Pro $60 / Enterprise), storage ($0.03/GB), and Weave ingestion ($0.10/MB) were unchanged.