DeepInfra cuts GLM-5.2 pricing about 20%
DeepInfra cut GLM-5.2 pricing ~20%: $0.93/$3.00 to $0.75/$2.40 per 1M tokens (in/out), while every other featured catalog model held steady.
$0.93 in / $3.00 out per 1M tokens ($0.18 cached)
$0.75 in / $2.40 out per 1M tokens ($0.14 cached)
GLM-5.2 (Z-AI’s flagship long-horizon agentic model, 1M-token context) dropped from $0.93 in / $3.00 out per 1M tokens ($0.18 cached) to $0.75 in / $2.40 out ($0.14 cached) between the 2026-07-29 and 2026-08-04 captures of DeepInfra’s /models catalog — a roughly 20% cut across all three prices. The cut is isolated to this one SKU: every other featured model checked in the same comparison (DeepSeek-V4-Pro, DeepSeek-V4-Flash, Kimi-K2.6/K2.7-Code, Nemotron-3-Ultra, Qwen3-Max, Qwen3.6-35B-A3B, GLM-5.1, MiMo-V2.5-Pro) held its rate exactly. GLM-5.2 pricing is not shown on DeepInfra’s main /pricing page — it currently surfaces only via the /models catalog’s Featured carousel.
GLM-5.2 (Z-AI's flagship long-horizon agentic model, 1M-token context) is cut roughly 20% across all three published rates — $0.93 in / $3.00 out / $0.18 cached to $0.75 / $2.40 / $0.14 per 1M tokens — while every other Featured model checked in the same comparison (DeepSeek-V4-Pro, DeepSeek-V4-Flash, Kimi-K2.6/K2.7-Code, Nemotron-3-Ultra, Qwen3-Max, Qwen3.6-35B-A3B, GLM-5.1, MiMo-V2.5-Pro) holds its rate exactly. GLM-5.2 is not shown on the /pricing page at all — it surfaces only in the /models Featured carousel — making this the first tracked price move that lives entirely outside DeepInfra's main pricing surface (source: deepinfra.com/models 2026-08-04).