DeepSeek's peak/off-peak API pricing goes live
DeepSeek's confirmed time-of-use rate card took effect as scheduled on 2026-08-16, replacing flat per-token API rates with peak/off-peak pricing up to 12x higher, alongside a new experimental vision model on the same rate card.
Flat per-token rates: V4-Flash cache-miss input $0.14/1M, cache-hit input $0.0028/1M, output $0.28/1M; V4-Pro cache-miss input $0.435/1M, cache-hit input $0.003625/1M, output $0.87/1M. Two published models (V4-Flash, V4-Pro).
Peak/off-peak billing live: V4-Flash cache-miss input $0.22/1M off-peak / $0.44/1M peak, cache-hit input $0.007/1M off-peak / $0.014/1M peak, output $0.66/1M off-peak / $1.32/1M peak; V4-Pro cache-miss input $0.66/1M off-peak / $1.32/1M peak, cache-hit input $0.022/1M off-peak / $0.044/1M peak, output $1.98/1M off-peak / $3.96/1M peak. Peak hours 01:00-04:00 and 06:00-10:00 UTC, Monday-Friday. A third model, DeepSeek-V4-Flash-Vision-Exp, was added on the same rate card as V4-Flash.
DeepSeek’s Models & Pricing page now shows only the peak/off-peak rate card first confirmed on 2026-08-14 — the prior flat per-token rates have been fully removed from the page, and a 2026-08-26 capture confirms the cutover took effect as scheduled at 16:00 UTC on 2026-08-16. Off-peak rates run roughly 1.5-2.5x the old flat rates and peak rates run roughly 3-12x, with the steepest jump on cache-hit input: DeepSeek-V4-Pro cache-hit input rises from a flat $0.003625/1M to $0.044/1M at peak, a more than 12x increase. Peak hours are defined as 01:00-04:00 and 06:00-10:00 UTC, Monday through Friday — a weekday restriction not previously documented, meaning the full weekend is billed at the off-peak rate.
The same capture also surfaced a new model, DeepSeek-V4-Flash-Vision-Exp, an experimental vision-capable variant of V4-Flash billed on V4-Flash’s exact rate card (2500-request concurrency, same peak/off-peak dollar figures), with images converted to input tokens by dimension. No grace period or legacy-rate opt-out has been published for the price increase.
DeepSeek replaced its vague "significant price increase" footnote with a concrete time-of-use rate card: peak-hour rates (01:00-04:00 and 06:00-10:00 UTC) roughly 3-11x current per-token prices depending on the line item, with off-peak rates at half of peak, effective 16:00 UTC on 2026-08-16. V4-Flash cache-miss input moves from a flat $0.14 to $0.22 off-peak / $0.44 peak; V4-Pro cache-miss input moves from a flat $0.435 to $0.66 off-peak / $1.32 peak. DeepSeek-V4-Pro also picked up a dated build suffix (-0813) and gained Responses API support, reaching feature parity with V4-Flash.