Packaging

Novita AI adds Kimi K3 at $3/$15 per M tokens, ends Hy3 free pricing, prunes catalog to 196 models

Novita AI pricing

Novita AI listed Kimi K3 at $3/M input and $15/M output, moved Tencent Hy3 off free pricing to $0.14/$0.58 per M, and cut its model catalog from 221 to 196 SKUs.

Before

221 models listed; Hy3 free ($0/M input, $0/M output); no Kimi K3; RTX 5090 32GB (High frequency) instance at 0.72/hr on-demand, 0.36/hr spot.

After

196 models listed; Hy3 at $0.14/M input ($0.035/M cache read) and $0.58/M output; Kimi K3 at $3/M input ($0.3/M cache read) and $15/M output on a 1,048,576-token context; RTX 5090 32GB instance at 0.73/hr on-demand, 0.37/hr spot.

Novita AI’s serverless rate card moved in three directions at once between the 2026-07-06 and 2026-07-21 captures.

A new price ceiling. Moonshot’s Kimi K3 enters the catalog at $3 per million input tokens ($0.3/M cache read) and $15 per million output tokens on a 1,048,576-token context window — the most expensive LLM Novita lists, roughly 4x the output rate of the previous Moonshot flagship Kimi K2.7 Code ($0.95/$4). For a platform whose positioning is “cheapest place to run open-weight models,” carrying a $15/M output SKU widens the top of the range considerably.

A free model goes paid. Tencent’s Hy3 was listed at $0/M input and $0/M output on 2026-07-06; it now bills $0.14/M input ($0.035/M cache read) and $0.58/M output. Free listings have been a recurring Novita acquisition lever — several small models were shown Free as far back as the September 2025 rate card — so a 262k-context model graduating out of free is a signal about where the loss-leader line now sits.

A 25-SKU cull. The catalog counter dropped from 221 to 196 models. The delisted SKUs are almost entirely legacy media and audio: GLM Image Generation, GLM Audio-to-Text, GLM Text-to-Speech and GLM Voice Clone, Hunyuan Image 3, Seedream 3.0 Text to Image, Seedance V1 Lite and Pro, Vidu 2.0 and Vidu Q1, Kling v2.1 including the Master variants, MiniMax Video 01 and 02, and the MiniMax speech-02 and speech-2.5-preview voices. Their successors — Seedance 1.5 Pro, Vidu Q2 and Q3, Kling v2.5, v2.6, v3.0 and o1, MiniMax speech-2.6 and 2.8 — all remain listed, so this reads as generational pruning rather than a retreat from media.

Infrastructure pricing was steady. Dedicated endpoints (RTX 4090 $0.61, RTX-5090 $0.73, H100 $1.99, H200 $2.99 per GPU-hour), bare metal (H100 SXM $1.70, B200 SXM $4.77 per GPU/hr), and the Agent Sandbox rate card ($0.0000098/vCPU-second, $0.0000032/GiB-second, $0.00009/GB-hour with the first 60 GB included) were all unchanged. The only compute move was the self-serve RTX 5090 32GB instance, which dropped its “High frequency” label and ticked from 0.72 to 0.73 per hour on-demand (0.36 to 0.37 spot).

From Novita AI's pricing timeline
Kimi K3 lands at $3/$15; Hy3 exits free; catalog pruned 221 → 196

Moonshot's Kimi K3 is added at $3/M input ($0.3/M cache read) and $15/M output on a 1,048,576-token context — the most expensive LLM on the rate card. Tencent's Hy3 moves off free pricing to $0.14/M input ($0.035/M cache read) and $0.58/M output. The catalog shrinks from 221 to 196 models as legacy media/audio SKUs are delisted (GLM Image Generation, GLM Audio-to-Text/TTS/Voice-Clone, Hunyuan Image 3, Seedream 3.0, Seedance V1 Lite/Pro, Vidu 2.0/Q1, Kling v2.1, MiniMax speech-02 and speech-2.5 preview). The RTX 5090 32GB instance drops its 'High frequency' label and ticks to 0.73/hr on-demand (0.37 spot) from 0.72/0.36. Dedicated endpoints, bare-metal, and Agent Sandbox rate cards unchanged.

About Novita AI
novita.ai ↗

Novita AI is a pay-as-you-go AI cloud offering inference across 196 listed models (down from 221 after a 2026-07-21 delisting of legacy media and audio SKUs), on-demand and bare-metal GPUs, and secure per-second agent sandboxes under a single API.

Free tier
Yes
Commits
None
Transparency
public

Novita AI pricing history

  1. Jul 2026
    Kimi K3 lands at $3/$15; Hy3 exits free; catalog pruned 221 → 196
  2. Jul 2026
    Base RTX 4090 instance trimmed; catalog refresh
  3. Jun 2026
    GPU-instance repricing + published sandbox rate card
  4. Jun 2026
    Pricing snapshot — per-token inference, per-hour GPU, per-second sandbox
  5. Oct 2025
    GLM-4.6, Kimi K2, DeepSeek V3.2 Exp added; output-token cuts
Full Novita AI timeline

More Novita AI activity

All pricing activity