Novita AI adds Kimi K3 at $3/$15 per M tokens, ends Hy3 free pricing, prunes catalog to 196 models
Novita AI listed Kimi K3 at $3/M input and $15/M output, moved Tencent Hy3 off free pricing to $0.14/$0.58 per M, and cut its model catalog from 221 to 196 SKUs.
221 models listed; Hy3 free ($0/M input, $0/M output); no Kimi K3; RTX 5090 32GB (High frequency) instance at 0.72/hr on-demand, 0.36/hr spot.
196 models listed; Hy3 at $0.14/M input ($0.035/M cache read) and $0.58/M output; Kimi K3 at $3/M input ($0.3/M cache read) and $15/M output on a 1,048,576-token context; RTX 5090 32GB instance at 0.73/hr on-demand, 0.37/hr spot.
Novita AI’s serverless rate card moved in three directions at once between the 2026-07-06 and 2026-07-21 captures.
A new price ceiling. Moonshot’s Kimi K3 enters the catalog at $3 per million input tokens ($0.3/M cache read) and $15 per million output tokens on a 1,048,576-token context window — the most expensive LLM Novita lists, roughly 4x the output rate of the previous Moonshot flagship Kimi K2.7 Code ($0.95/$4). For a platform whose positioning is “cheapest place to run open-weight models,” carrying a $15/M output SKU widens the top of the range considerably.
A free model goes paid. Tencent’s Hy3 was listed at $0/M input and $0/M output on 2026-07-06; it now bills $0.14/M input ($0.035/M cache read) and $0.58/M output. Free listings have been a recurring Novita acquisition lever — several small models were shown Free as far back as the September 2025 rate card — so a 262k-context model graduating out of free is a signal about where the loss-leader line now sits.
A 25-SKU cull. The catalog counter dropped from 221 to 196 models. The delisted SKUs are almost entirely legacy media and audio: GLM Image Generation, GLM Audio-to-Text, GLM Text-to-Speech and GLM Voice Clone, Hunyuan Image 3, Seedream 3.0 Text to Image, Seedance V1 Lite and Pro, Vidu 2.0 and Vidu Q1, Kling v2.1 including the Master variants, MiniMax Video 01 and 02, and the MiniMax speech-02 and speech-2.5-preview voices. Their successors — Seedance 1.5 Pro, Vidu Q2 and Q3, Kling v2.5, v2.6, v3.0 and o1, MiniMax speech-2.6 and 2.8 — all remain listed, so this reads as generational pruning rather than a retreat from media.
Infrastructure pricing was steady. Dedicated endpoints (RTX 4090 $0.61, RTX-5090 $0.73, H100 $1.99, H200 $2.99 per GPU-hour), bare metal (H100 SXM $1.70, B200 SXM $4.77 per GPU/hr), and the Agent Sandbox rate card ($0.0000098/vCPU-second, $0.0000032/GiB-second, $0.00009/GB-hour with the first 60 GB included) were all unchanged. The only compute move was the self-serve RTX 5090 32GB instance, which dropped its “High frequency” label and ticked from 0.72 to 0.73 per hour on-demand (0.36 to 0.37 spot).
Moonshot's Kimi K3 is added at $3/M input ($0.3/M cache read) and $15/M output on a 1,048,576-token context — the most expensive LLM on the rate card. Tencent's Hy3 moves off free pricing to $0.14/M input ($0.035/M cache read) and $0.58/M output. The catalog shrinks from 221 to 196 models as legacy media/audio SKUs are delisted (GLM Image Generation, GLM Audio-to-Text/TTS/Voice-Clone, Hunyuan Image 3, Seedream 3.0, Seedance V1 Lite/Pro, Vidu 2.0/Q1, Kling v2.1, MiniMax speech-02 and speech-2.5 preview). The RTX 5090 32GB instance drops its 'High frequency' label and ticks to 0.73/hr on-demand (0.37 spot) from 0.72/0.36. Dedicated endpoints, bare-metal, and Agent Sandbox rate cards unchanged.