AI Summary
About
Novita AI is a pay-as-you-go AI cloud that bundles three products under a single API: serverless model inference across 200+ open models, GPU compute (on-demand, spot, and bare-metal), and a per-second Agent Sandbox runtime for executing AI-generated code, browser workflows, and computer-use tasks. Its positioning line — “200+ models, on-demand GPUs, and secure agent runtimes — unified under one API. Free to start, scales as you grow” — captures the strategy: be the cheapest, broadest place for developers to run open-weight models and the GPUs underneath them.
The company competes with inference aggregators (Together AI, Fireworks, DeepInfra, Replicate) on the model-API side and with GPU clouds (RunPod, Lambda, Vast.ai) on the compute side. Its differentiator is doing both at once and undercutting first-party model APIs: DeepSeek, Qwen, GLM, Kimi, MiniMax, and Llama families are all listed at aggressive per-million-token rates, frequently with cache-read discounts. Novita serves individual developers and prosumers (self-serve, free to start) up through SMBs and enterprises (dedicated endpoints, bare-metal clusters, and a contact-sales Enterprise sandbox tier with a 99.95% SLA).
Pricing transparency is a notable strength: nearly every dimension — token rates for the models currently listed, per-image and per-video-second media rates, on-demand and spot GPU hourly rates, and per-second sandbox examples — is published openly, with “contact sales” reserved for bare-metal nodes beyond H100/B200 and the Enterprise sandbox tier. The catalog churns quickly in both directions: the 2026-07-21 capture added Moonshot’s Kimi K3 at $3/M input and $15/M output while delisting 25 legacy media and audio SKUs, the 2026-07-29 capture pruned the catalog further to 172 models — removing Hunyuan Video Fast, the PixVerse V4.5 family, and the Kling-o1 lineup — while Flux.1 Kontext Pro image-generation pricing rose 10x, from $0.036 to $0.36 per image, and the 2026-08-04 capture grew the catalog back to 177 with new SKUs (Deepseek V3.2, V4 Flash, DeepSeek-OCR 2, PaddleOCR-VL) while removing RTX 6000 Ada 48GB from the self-serve GPU-instance lineup. The 2026-08-14 capture eased the catalog to 174, delisting Ling 3.0 Tiny — a $0 LLM added just three days earlier on 2026-08-11 — and leaving no free LLM on the rate card. The 2026-08-25 capture shows the /en/models catalog counter (All Models tab) at 144 — a different basis than the “174” this page logged for 2026-08-14, not reconciled by this capture — alongside a real repricing (Deepseek V4 Flash 0731 up 214% on input and 371% on output to $0.44/$1.32 per M tokens, with a new DeepSeek V4 Pro 0813 SKU added), RTX 6000 Ada 48GB returning to the self-serve GPU lineup at $0.77/hr in place of the delisted RTX 4090 “High frequency” tier, and a sharp contraction of the Image and Video model catalogs (Z Image Turbo, Seedream 4.0, the FLUX 2 family, Veo 3.1, the Vidu Q2/Q3 lineup, older Kling versions, Wan 2.1/2.2, and Seedance 1.5 Pro were all delisted).
Novita’s archived pricing pages show it started life around 2023–2024 as a credit-funded Stable-Diffusion image-generation API (footer HQ in Singapore), then pivoted through 2025 into a full inference-plus-GPU cloud and relisted its HQ in San Francisco — see Pricing evolution. Public funding is undisclosed (Crunchbase and Tracxn list no disclosed rounds), and community signal is modest rather than viral: the company has no Hacker News thread above single digits and surfaces mostly in r/LocalLLaMA model-availability chatter, so its growth has been distribution-led (cheap open-model inference) rather than launch-hype-led.
Pricing summary : How Novita AI’s pure usage-based AI cloud bills
Novita AI uses a pure usage-based model — free to start, no seats, no monthly minimum — billed across four metering dimensions:
- Per-token model inference (serverless): LLMs are billed per million tokens, split input/output, e.g. Llama 3.1 8B at $0.02/M in · $0.05/M out, DeepSeek V3.1 at $0.27/M in · $1/M out, DeepSeek V4 Pro at $1.6/M in · $3.2/M out, and Kimi K3 at the top of the card at $3/M in · $15/M out, with cache-read rates on many models (often ~50% of input; Kimi K3 reads cache at $0.3/M against its $3/M input rate). Image generation is per image (from $0.018/image, Flux.1 Kontext Dev fast_mode — the cheapest listed image model as of 2026-08-25, after Z Image Turbo and most of the prior Image catalog were delisted), video per video or per second (e.g. $0.084/s Kling v3.0 Standard, no audio), and audio per 1M characters ($15/1M characters Fish Audio TTS). Batch inference carries an introductory 50% discount on input and output tokens for supported models.
- Per-hour GPU instances: On-demand instances now span 0.33/hr/GPU (RTX 4090 24GB) to 3.39/hr/GPU (NVIDIA H100 SXM 80GB), with spot pricing at roughly half on RTX-class cards (RTX 4090 0.33 on-demand vs 0.17 spot; RTX 5090 0.73 vs 0.37) and H100 spot priced much lower, at 1.70/hr — prices shown as Novita renders them, in USD with no
$glyph. Billed per second, scale to zero. (The self-serve/en/gpuslineup was RTX-class only from 2026-06-24 through 2026-08-25; the 2026-08-26 capture shows NVIDIA H100 SXM 80GB and NVIDIA L40S 48GB back on the self-serve page — H100 on-demand 3.39/hr / spot 1.70/hr, L40S on-demand only at 0.55/hr with no spot listed — while RTX 6000 Ada 48GB, which had returned just one day earlier on 2026-08-25, was removed again, and the RTX 5090 “High frequency” variant was dropped, leaving a single RTX 5090 listing at 0.73/0.37.) - Per-GPU-hour dedicated endpoints + bare metal: Isolated dedicated endpoints at $0.61 (RTX 4090) / $1.99 (H100) / $2.99 (H200) per GPU-hour; bare-metal 8-GPU nodes at $1.70/GPU/hr (H100 SXM) and $4.77/GPU/hr (B200 SXM).
- Per-second agent sandbox: Billed on allocated vCPU + memory, e.g. ~$0.0034 for a 5-minute task on 1 vCPU + 512 MiB RAM.
What makes this different: Novita prices the same physical GPU three ways depending on packaging (shared-serverless token rates, on-demand hourly instances, and reserved bare-metal per-GPU-hour), letting a customer move down the cost curve as commitment rises — without ever signing an annual contract.
Pricing by product
Serverless model inference (per-token LLMs)
| Model | Context | Input | Output | Key mechanics |
|---|---|---|---|---|
| Llama 3.1 8B Instruct | 16,384 | $0.02 /M | $0.05 /M | Cheapest flagship-class LLM listed |
| Llama 3.3 70B Instruct | 12,000 | $0.135 /M | $0.4 /M | Popular mid-size open model; short listed context |
| DeepSeek V3.1 | 131,072 | $0.27 /M | $1 /M | Cache read $0.135/M |
| DeepSeek V4 Pro | 1,048,576 | $1.6 /M | $3.2 /M | Flagship; cache read $0.135/M |
| DeepSeek V4 Flash | 1,048,576 | $0.14 /M | $0.28 /M | Low-cost long-context; cache read $0.028/M |
| Kimi K3 | 1,048,576 | $3 /M | $15 /M | Newest and priciest LLM on the card; cache read $0.3/M |
| Qwen3.7-Max | 977,000 | $1.25 /M | $3.75 /M | Cache read $0.25/M; the “Limited Time 50% Off” tag seen on earlier captures is gone as of 2026-08-14 — the rate held steady rather than reverting higher |
| GLM-5.1 | 204,800 | $1.38 /M | $4.4 /M | Cache read $0.26/M |
| Kimi K2.6 | 262,144 | $0.8 /M | $3.4 /M | Cache read $0.16/M |
| Hy3 (Hunyuan) | 262,144 | $0.14 /M | $0.58 /M | Moved off free pricing; cache read $0.035/M |
| MiniMax-M2 | 204,800 | $0.3 /M | $1.2 /M | Cache read $0.03/M |
| OpenAI GPT OSS 120B | 131,072 | $0.05 /M | $0.25 /M | Open-weight GPT OSS |
Catalog lists 174 models across LLM, Image, Audio, Video, Embedding, Reranker, and Vision as of the 2026-08-14 capture — down from 177 at the 2026-08-04 and 2026-08-11 captures, which had grown from 172 at 2026-07-29 with net-new SKUs including Deepseek V3.2 (non-Exp, $0.269/M in · $0.4/M out), Deepseek V4 Flash / V4 Flash 0731 ($0.14/M in · $0.28/M out), DeepSeek-OCR / DeepSeek-OCR 2 ($0.03/M flat), and PaddleOCR-VL ($0.02/M flat). Some models (Qwen3 Max, MiniMax M3) still show “Tiered pricing” instead of a flat rate on the pricing page — though the separate models catalog surfaces effective per-token rates for both (Qwen3 Max $2.11/M in · $8.45/M out; MiniMax M3 $0.3/M in · $1.2/M out) — and Qwen3 Omni 30B A3B Instruct / Thinking still show only “Omnimodal” with no rate on the pricing page itself. The 2026-08-11 capture saw three models that carried the “Time Limited Free” tag (Macaron V1 Venti, Macaron V1 Tall, Ling 3.0 Flash) convert to paid per-token pricing — Macaron V1 Venti $1.5/M in ($0.3/M cache read) · $4.5/M out, Macaron V1 Tall $0.45/M in ($0.08/M cache read) · $2.6/M out, Ling 3.0 Flash $0.06/M in ($0.012/M cache read) · $0.18/M out — while a new model, Ling 3.0 Tiny, was added at $0/M flat under the same “Time Limited Free” tag, taking over the free-launch slot (the same free-then-convert pattern Hy3 went through on 2026-07-21). That free slot didn’t last the week: the 2026-08-14 capture shows Ling 3.0 Tiny delisted outright rather than converted to a paid rate, leaving no free LLM on the rate card. Separately, Qwen3.7-Max’s “Limited Time 50% Off” tag (visible in earlier captures) no longer appears next to its $1.25/M in · $3.75/M out rate, which held steady rather than reverting to a higher price. Embeddings are billed per million tokens: BAAI BGE-M3 $0.01/M, Qwen3 Embedding 0.6B and 8B $0.07/M; rerankers are also per-million-token: baai/bge-reranker-v2-m3 $0.01/M, Qwen3 Reranker 8B $0.05/M. Batch inference carries an introductory 50% discount on input/output tokens for supported models. The 2026-08-25 capture shows the /en/models catalog counter (All Models tab) at 144 — a different count basis than the “174” figure logged for 2026-08-14, which this capture cannot retroactively reconcile; the discrepancy suggests either a counting-method change or a labeling error in a prior capture, and future runs should treat the models-page counter as the source of truth. The 2026-08-25 capture also shows a real repricing: Deepseek V4 Flash 0731 jumped to $0.44/M in ($0.028/M cache read) · $1.32/M out, up from $0.14/M in · $0.28/M out on 2026-08-14 (input +214%, output +371%), while the un-dated Deepseek V4 Flash ($0.14/$0.28) and Deepseek V4 Pro ($1.6/$3.2) rows held steady. A new flagship, DeepSeek V4 Pro 0813, was added at $1.32/M in ($0.132/M cache read) · $3.96/M out. In the inclusionai group, Ling-2.6-flash, Ling-2.6-1T, and Ring-2.6-1T were delisted, replaced by a new Ling 3.0 Flash Fast at the same $0.06/M in ($0.012/M cache read) · $0.18/M out rate as the existing Ling 3.0 Flash.
Serverless media inference (per-image / per-video / per-character)
| Modality | Example API | Price | Key mechanics |
|---|---|---|---|
| Image | Qwen-Image Text to Image | $0.02 /image | Base text-to-image; edit variant same rate |
| Image | Flux.1 Kontext Pro | $0.36 /image | Repriced 10x versus the 2026-07-21 capture; now pricier than Kontext Max ($0.072) |
| Image | Flux.1 Kontext Dev (fast_mode) | $0.018 /image | Cheapest listed image model as of 2026-08-25 — Z Image Turbo, Seedream 4.0, Image Eraser/Remove Background/Upscaler, and the FLUX 2 family (Dev/Flex/Pro) were all delisted from the Image catalog on 2026-08-25, which now carries just 5 SKUs |
| Video | Heygen Video-translate | $0.0375 /video | Flat per-video rate; Veo 3.1, the Vidu Q2/Q3 lineup, Wan 2.1/2.2, and Seedance 1.5 Pro were all delisted from the Video catalog on 2026-08-25 |
| Video | Kling v3.0 Standard (5s, No Audio) | $0.084 /s | Kling v3.0 replaced the entire delisted Kling V1.6/V2.5 Turbo/V2.6 Pro lineup on 2026-08-25 |
| Video | Kling v3.0 Pro (per second) | $0.112–$0.168 /s | Audio variant is the higher rate |
| Audio | Fish Audio Text to Speech | $15 /1M characters | Per-character TTS |
| Audio | MiniMax speech-2.6-hd | $100 /1M characters | Premium HD voice ($60 for turbo) |
| Audio | MiniMax Voice-Cloning | $1.5 /voice | Per cloned voice; Voice Design $3/voice |
| AI Search | EXA neuralSearch | $0.007 /request | deepSearch $0.012, deepReasoningSearch $0.015 |
| AI Search | Tavily basicSearch | $0.008 /request | advancedSearch $0.016; extract from $0.0016/url |
GPU instances (on-demand vs spot, per hour)
| GPU | VRAM | On-Demand | Spot | Key mechanics |
|---|---|---|---|---|
| RTX 4090 24GB | 24 GB | 0.33 /hr/GPU | 0.17 /hr/GPU | Cheapest on-demand instance; spot ~48% off |
| NVIDIA L40S 48GB | 48 GB | 0.55 /hr/GPU | unknown (page shows ”—”, no spot rate listed) | Returned to self-serve 2026-08-26, after being pulled from this page on 2026-06-24 |
| RTX 5090 32GB | 32 GB | 0.73 /hr/GPU | 0.37 /hr/GPU | Sole RTX 5090 listing as of 2026-08-26 — the “High frequency” variant (was 0.72/hr) was delisted the same day |
| NVIDIA H100 SXM 80GB | 80 GB | 3.39 /hr/GPU | 1.70 /hr/GPU | Returned to self-serve 2026-08-26, after being pulled from this page on 2026-06-24; priciest self-serve on-demand instance, but its spot rate matches the bare-metal on-demand rate |
GPU instance rates above are shown exactly as Novita renders them on /en/gpus (USD per GPU per hour, no $ glyph on the source page). The self-serve instance lineup regained NVIDIA H100 SXM and L40S on 2026-08-26 — RTX 6000 Ada 48GB (itself only re-added the day before, 2026-08-25) and the RTX 5090 “High frequency” tier were both removed in the same capture. H100 capacity is also sold via dedicated endpoints ($1.99/GPU-hr) and bare-metal nodes ($1.70/GPU/hr). Novita advertises spot instances as “Save up to 50% on Costs,” alongside “Scales to zero in 30s,” “Unlimited concurrency,” and per-second billing.
Dedicated endpoints (isolated GPUs, per GPU-hour)
| GPU | VRAM | Price / GPU-hour | Key mechanics |
|---|---|---|---|
| RTX 4090 | 24 GB | $0.61 | Per-second billing on running replicas |
| RTX-5090 | 32 GB | $0.73 | Newest Blackwell consumer GPU as an endpoint |
| NVIDIA H100 | 80 GB | $1.99 | Guaranteed performance, no sharing |
| NVIDIA H200 | 141 GB | $2.99 | Marked “Popular”; autoscaling + scale-to-zero |
Dedicated image endpoints are sold as monthly subscriptions: Standard $559/month and Pro $1,199/month (both: exclusive high-performance GPU, unlimited images, 24/7 support, load 500 models), via contact sales.
Bare-metal GPU servers (8 GPUs per node, per GPU/hr)
| GPU | Config | Price / GPU/hr | Key mechanics |
|---|---|---|---|
| H100 SXM | 8× per node, 80 GB HBM3 | $1.70 | ”Best value”; NVLink 900 GB/s + RDMA |
| B200 SXM | 8× per node, 192 GB HBM3e | $4.77 | ”Top performance”; NVLink 5th-gen 1.8 TB/s |
| H200 SXM | 8× per node, 141 GB HBM3e | Custom | Large-context / KV-cache workloads; contact us |
| RTX 5090 | 8× per node, 32 GB GDDR7 | Custom | Cost-efficient inference; contact us |
| RTX 4090 | 8× per node, 24 GB GDDR6X | Custom | Broadest software compatibility; contact us |
Agent Sandbox (per-second on vCPU + memory)
Novita now publishes the underlying per-unit sandbox rate card (previously only example costs were shown):
| Billing basis | Unit price | Key mechanics |
|---|---|---|
| vCPU | $0.0000098 / vCPU-second | Calculated per allocated vCPU-second |
| Memory | $0.0000032 / GiB-second | Calculated per allocated GiB-second |
| Storage | $0.00009 / GB-hour | Measured hourly, charged daily; first 60 GB included |
| Configuration | Workload example | Estimated cost | Key mechanics |
|---|---|---|---|
| 1 vCPU + 512 MiB RAM | Short-lived agent task (5 min) | ~$0.0034 | Billed per second |
| 2 vCPU + 1 GiB RAM | Code execution job (1 hr) | ~$0.0821 | No plans, no lock-ins |
| 8 vCPU + 8 GiB RAM | Multi-agent / RL workload (1 hr) | ~$0.3744 | Sub-200ms startup |
| Enterprise | Higher limits, custom regions | Custom (contact sales) | 99.95% uptime SLA; unlimited concurrency |
New sandbox users get $100 in free credits (valid 90 days, no card required). The free tier allows 5 concurrent sandboxes, 1-hour max sessions, and 2 vCPU / 4 GB RAM per sandbox; topping the balance above $0 auto-unlocks the paid tier (100 concurrent, 24-hour sessions, up to 8 vCPU / 8 GB RAM).
Sales motions across products: PLG / self-serve for serverless inference, GPU instances, dedicated endpoints, and standard sandbox usage; sales-led for bare-metal nodes beyond H100/B200, image endpoint subscriptions, and the Enterprise sandbox tier.
Hidden costs : What inference + GPU bills actually look like at volume
The “free to start” headline hides how quickly a production workload can compound across token, GPU, and sandbox dimensions. Two representative archetypes:
A startup running a DeepSeek-V3.1 chat product
| Line item | Monthly cost |
|---|---|
| 300M input tokens @ $0.27/M | $81 |
| 120M output tokens @ $1/M | $120 |
| 50M cache-read tokens @ $0.135/M | $6.75 |
| 1× dedicated H100 endpoint, ~200 hrs active @ $1.99/GPU-hr | $398 |
| Total | ~$605.75 |
The dedicated endpoint — not the tokens — dominates the bill once you reserve isolated capacity, so teams must weigh shared serverless token rates against the latency guarantees of a dedicated GPU.
An agent platform running browser + code sandboxes
| Line item | Monthly cost |
|---|---|
| 50,000 short coding tasks @ ~$0.0034 (1 vCPU/512MiB, 5 min) | $170 |
| 5,000 hour-long multi-agent jobs @ ~$0.3744 (8 vCPU/8GiB) | $1,872 |
| Spot RTX 4090 for model serving, ~300 hrs @ 0.17/hr | $51 |
| Total | ~$2,093 |
Per-second sandbox pricing looks trivial per task, but at agent-platform concurrency the long-running 8-vCPU jobs are the cost driver — exactly the line a usage-based pricing buyer should model before committing.
Want to estimate your own Novita AI bill? Use the Novita AI pricing calculator to model your monthly cost based on token volume, GPU hours, and sandbox seconds.
Pricing evolution : From a credit-funded image API to a full AI cloud stack
Novita’s archived pricing pages tell a clean origin story: it started in 2023–2024 as a Stable-Diffusion image-generation API billed out of a prepaid credit balance (USDT/Stripe top-ups, “1/10 of DALL-E2 and MJ”), then pivoted through 2025 into a broad model-inference + GPU cloud, adding LLMs, Audio, GPUs, Dedicated Endpoints, and finally a per-second Agent Sandbox — moving its listed HQ from Singapore to San Francisco along the way.
Cadence
| Quarter | Price changes | Product / SKU additions | Notes |
|---|---|---|---|
| 2024 Q1 | 0 (baseline) | 0 | Credit-funded image API; text-to-image $0.0015/image (512×512), upscale from $0.0021/image; Singapore HQ. |
| 2024 Q2 | 0 | 2 | LLM API and Audio (Text to Speech, Voice Cloning) added to the catalog; image rates unchanged. |
| 2025 Q1 | 1 | 1 | Redesign to “Model APIs / GPUs” tabs; HQ moves to San Francisco; $10,000 startup-credits program; image shown at $0.003/image (1024×1024); Hunyuan/Wan video rates listed per-second. |
| 2025 Q3 | 1 | 1 | Agent Sandbox launches; page splits into Serverless / Dedicated / GPUs / Sandbox tabs; image back to $0.001/image (512×512); full LLM rate card published (DeepSeek V3.1 $0.27/$1, Llama 3.1 8B $0.02/$0.05). |
| 2025 Q4 | 1 | 1 | GLM-4.6, Kimi K2 and DeepSeek V3.2 Exp added; V3.2 Exp output cut to $0.41/M vs V3.1’s $1/M. |
| 2026 Q2 | 1 | 1 | GPU-instance repricing: RTX 4090 on-demand cut to $0.35/hr (from $0.67), spot to $0.18 (from $0.34); lineup narrowed to RTX-class (L40S + H100 instances removed); RTX 6000 Ada added. RTX-5090 dedicated endpoint ($0.73) added; Agent Sandbox per-unit rate card published + $100 / 90-day free credits. |
| 2026 Q3 | 11 | 8 add / 75 delist | RTX 4090 instance trimmed again to 0.33/hr (0.17 spot); Kimi K3 added at $3/$15 per M — a new price ceiling; Hy3 moves off Free to $0.14/$0.58; RTX 5090 (HF) instance ticks to 0.73/hr (0.37 spot) then reverts to 0.72/hr (0.36 spot) with the “High frequency” label restored; catalog pruned 221 → 196 → 172 then rebuilt to 177 as legacy media/video/audio SKUs are delisted (Google’s Veo 3.1 added) and new LLM/OCR SKUs land (Deepseek V3.2, DeepSeek-OCR 2); Flux.1 Kontext Pro image price jumps 10x to $0.36/image, now pricier than Kontext Max; RTX 6000 Ada 48GB is delisted from self-serve GPUs and replaced by a new base RTX 5090 32GB on-demand tier at $0.73/hr — priced above the adjacent HF variant; three Time Limited Free models (Macaron V1 Venti, Macaron V1 Tall, Ling 3.0 Flash) graduate to paid pricing on 2026-08-11 as Ling 3.0 Tiny takes the $0 slot, then is delisted outright three days later on 2026-08-14 — the shortest-lived free listing tracked — easing the catalog to 174; on 2026-08-25, Deepseek V4 Flash 0731 repriced sharply (input +214% to $0.44/M, output +371% to $1.32/M) as a new DeepSeek V4 Pro 0813 SKU landed at $1.32/$3.96, RTX 6000 Ada 48GB returned to self-serve GPUs at $0.77/hr in place of the delisted RTX 4090 “High frequency” tier, and the Image catalog was cut 14 → 5 SKUs while Video dropped Veo 3.1, Vidu Q2/Q3, older Kling versions, and Wan 2.1/2.2 in favor of Kling v3.0 and Wan 2.5–2.7; one day later, on 2026-08-26, the self-serve GPU lineup flipped again — H100 SXM 80GB ($3.39/hr on-demand, $1.70/hr spot) and L40S 48GB ($0.55/hr) returned for the first time since 2026-06-24, while the just-restored RTX 6000 Ada 48GB and the RTX 5090 “High frequency” variant were both delisted in the same pass. |
Tracked range: 2024-02-25 → 2026-08-26 via Wayback + live capture. Quarters not listed (2024 Q3–Q4, 2025 Q2, 2026 Q1) showed no material pricing or packaging change in the archived snapshots reviewed.
Notable changes
- 2024-02-25 — Earliest archived surface: a credit-funded image-generation API (txt2img $0.0015/image, USDT/Stripe top-ups), tagline “1/10 of DALL-E2 and MJ,” Singapore HQ.
- 2024-05-21 — LLM API and Audio (TTS, Voice Cloning) added to the product catalog.
- 2025-02-10 — Major redesign to Model APIs + GPUs tabs; listed HQ moves to San Francisco; $10,000 startup-credits program launches.
- 2025-08-04 — Agent Sandbox launches; pricing splits into four endpoint tabs; batch inference advertised at an introductory 50% token discount.
- 2025-09-11 — Full serverless LLM rate card published (DeepSeek V3.1 $0.27/$1, Llama 3.1 8B $0.02/$0.05, GLM-4.5 $0.6/$2.2), with several small models listed Free.
- 2025-10-03 — GLM-4.6, Kimi K2, DeepSeek V3.2 Exp added; DeepSeek V3.2 Exp output token cut to $0.41/M.
- 2026-04-21 — Per the Novita blog, Kimi K2.6 launched at $0.95/$4.00 per M tokens; by the 2026-06-02 capture the same family is listed at $0.8/$3.4 — a downward revision within ~6 weeks, consistent with Novita’s habit of trimming rates after launch.
- 2026-06-24 — GPU-instance repricing and lineup change (live capture): RTX 4090 on-demand drops to $0.35/hr ($0.18 spot) from $0.67/$0.34; the
/en/gpuspage is narrowed to RTX-class GPUs (RTX 4090, RTX 4090 HF, RTX 5090 HF, new RTX 6000 Ada 48GB at $0.77), with L40S and H100 SXM instances removed — H100 now sold only via dedicated endpoints ($1.99/GPU-hr) and bare-metal ($1.70/GPU/hr). Dedicated endpoints add RTX-5090 at $0.73/GPU-hr, and the Agent Sandbox begins publishing its per-unit rate card ($0.0000098/vCPU-second, $0.0000032/GiB-second, $0.00009/GB-hour, first 60 GB free) alongside a $100 / 90-day free-credit offer. - 2026-07-06 — Second, much smaller instance trim: RTX 4090 24GB goes to 0.33/hr on-demand (0.17 spot) from $0.35/$0.18, while the rest of the RTX lineup holds. The serverless catalog absorbs a wave of new flagships (GLM 5.2 $1.4/$4.4, GLM-4.7 $0.6/$2.2, Kimi K2.7 Code $0.95/$4, MiniMax M2.7/M2.5) with no rate change on the incumbent SKUs — Novita adding at the top of the card rather than repricing beneath it.
- 2026-07-21 — Three moves in one capture. Kimi K3 lands at $3/M input · $15/M output (cache read $0.3/M, 1,048,576-token context), roughly 4x the output rate of the previous Moonshot flagship and the highest price Novita has ever listed for an LLM. Tencent’s Hy3 exits free pricing to $0.14/M in · $0.58/M out, emptying the free-LLM shelf entirely — only MOSS TTS v1.5 remains a Free listing, and it is an audio model. The catalog is pruned 221 → 196 as 25 legacy media/audio SKUs are delisted (GLM Image Generation, GLM Audio-to-Text/TTS/Voice-Clone, Hunyuan Image 3, Seedream 3.0, Seedance V1 Lite/Pro, Vidu 2.0/Q1, Kling v2.1, MiniMax Video 01/02, MiniMax speech-02 and speech-2.5-preview), with their successors all still listed. Compute barely moved: the RTX 5090 32GB instance ticked 0.72 → 0.73 on-demand (0.36 → 0.37 spot) and dropped its “High frequency” label, while dedicated endpoints, bare metal, and the Agent Sandbox rate card were byte-identical to the prior capture.
- 2026-07-29 — A straight repricing, not just a catalog edit: Flux.1 Kontext Pro image-generation pricing rises 10x, from $0.036 to $0.36 per image, confirmed against three independent prior captures (2026-07-06 and 2026-07-21’s main pricing page and models catalog) — it is now pricier than the ostensibly higher-tier Flux.1 Kontext Max, which held at $0.072. In the same capture the catalog is pruned again, 196 → 172, as Hunyuan Video Fast, the PixVerse V4.5 family, and the entire Kling-o1 lineup are delisted while Google’s Veo 3.1 is added (from $0.10/video at 720p). The RTX 5090 32GB instance reverts to 0.72/hr on-demand (0.36 spot) with the “High frequency” label restored, undoing the 2026-07-21 tick.
- 2026-08-04 — Self-serve GPU-instance lineup swap: RTX 6000 Ada 48GB (was $0.77/hr on-demand, $0.39 spot) is delisted from
/en/gpus, replaced by a new base RTX 5090 32GB on-demand tier at $0.73/hr ($0.37 spot) — priced a cent above the existing “High frequency” RTX 5090 variant ($0.72/hr), inverting the RTX 4090 pair, where the HF variant costs roughly double the base rate. Dedicated endpoints, bare-metal, and the Agent Sandbox rate card are unchanged; the model catalog grows 172 → 177 with additions including Deepseek V3.2, Deepseek V4 Flash, DeepSeek-OCR 2, and PaddleOCR-VL. - 2026-08-11 — Three Time Limited Free models graduate to paid pricing: Macaron V1 Venti ($1.5/M in · $4.5/M out), Macaron V1 Tall ($0.45/M in · $2.6/M out), and Ling 3.0 Flash ($0.06/M in · $0.18/M out) — while a new model, Ling 3.0 Tiny, takes over the $0 free-launch slot under the same tag. Catalog holds at 177; GPU, dedicated-endpoint, bare-metal, and Agent Sandbox rate cards unchanged.
- 2026-08-14 — Ling 3.0 Tiny is delisted outright, just three days after launch — not repriced the way Hy3 was on 2026-07-21, simply removed from both the pricing page and the model catalog — leaving no free LLM on the rate card and easing the overall catalog from 177 to 174 models. Qwen3.7-Max’s “Limited Time 50% Off” tag also disappears from the rate table, though its $1.25/M in · $3.75/M out rate holds steady rather than reverting higher.
- 2026-08-25 — The sharpest single-SKU repricing since Flux.1 Kontext Pro’s 10x jump: Deepseek V4 Flash 0731 rises from $0.14/M in · $0.28/M out to $0.44/M in · $1.32/M out (input +214%, output +371%), while a new flagship, DeepSeek V4 Pro 0813, is added at $1.32/M in · $3.96/M out. On the hardware side, RTX 6000 Ada 48GB returns to the self-serve
/en/gpuslineup at $0.77/hr — the same GPU delisted from that exact page on 2026-08-04 — swapped in for the delisted RTX 4090 24GB “High frequency” tier. The media catalog contracts hard: Image drops from 14 to 5 SKUs (Z Image Turbo, Seedream 4.0, Image Eraser/Remove Background/Upscaler, and the FLUX 2 Dev/Flex/Pro family all delisted) and Video sheds Veo 3.1, the Vidu Q2/Q3 lineup, older Kling versions (V1.6/V2.5 Turbo/V2.6 Pro), Wan 2.1/2.2, and Seedance 1.5 Pro in favor of Kling v3.0 and Wan 2.5–2.7. The/en/modelscatalog counter separately reads 144, versus the 174 this page logged for 2026-08-14 — a count-basis discrepancy this capture flags but does not resolve. - 2026-08-26 — The self-serve
/en/gpuslineup turns over again after just one day: NVIDIA H100 SXM 80GB ($3.39/hr on-demand, $1.70/hr spot) and NVIDIA L40S 48GB ($0.55/hr on-demand, no spot rate shown) return to self-serve for the first time since being pulled on 2026-06-24, while RTX 6000 Ada 48GB — re-added only the day before, on 2026-08-25 — is delisted again, and the RTX 5090 “High frequency” variant ($0.72/hr) is also removed, leaving a single RTX 5090 32GB listing at $0.73/$0.37. Self-serve H100 on-demand ($3.39/hr) now prices above both the dedicated-endpoint ($1.99/hr) and bare-metal ($1.70/hr) rate for the same GPU, while H100 spot ($1.70/hr) happens to land exactly on the bare-metal on-demand rate. Serverless model rates, dedicated endpoints, bare-metal, and Agent Sandbox were all unchanged from 2026-08-25.
The model-API pivot in detail
The most consequential change is not a single price move but a product-category pivot. In early 2024 Novita’s entire pricing page was image-generation APIs billed from a prepaid credit wallet — the same packaging a hobbyist Stable-Diffusion service would use. By late 2025 the page had become a four-tab AI-cloud rate card spanning per-token LLM inference, per-hour GPUs, per-GPU-hour dedicated endpoints, and per-second agent sandboxes. The credit-wallet UX was replaced by pure pay-as-you-go metering, and the headline shifted from “cheaper than DALL-E” to “200+ models under one API.” Novita rode the open-weight model wave — DeepSeek, Qwen, GLM, Kimi — from a niche media tool into a general-purpose inference cloud in under two years.
The 2026 Q3 captures show what the second act looks like. For two years the catalog only grew, and the cheap end did the marketing — free small models, sub-cent images, a $0.02/M Llama. Between 2026-07-06 and 2026-07-21 both ends moved at once: Kimi K3 pushed the ceiling to $15/M output, roughly 300x the cheapest listed output rate, while the free-LLM shelf emptied and 25 superseded media SKUs were struck off. Eight days later, on 2026-07-29, the pattern extended from catalog to price: Flux.1 Kontext Pro — an established, already-paid image model, not a new listing or a graduating freebie — was repriced 10x in a single capture, while another 25 media/video SKUs were pruned. Novita is no longer just racing to the bottom on open-weight tokens; it is stretching the rate card wide enough to carry frontier-class models, letting the loss-leaders lapse once they have done their acquisition job, and now repricing established SKUs sharply when it chooses to. The 2026-08-04 capture shows the same churn instinct applied to hardware SKUs, not just models: RTX 6000 Ada was pulled from the self-serve GPU lineup entirely and replaced by a new RTX 5090 tier priced above the variant sitting right next to it — packaging edits on this rate card are no longer confined to the token side of the business. One week later the free-LLM shelf itself became the object lesson: Ling 3.0 Tiny took over the $0 slot on 2026-08-11 only to be delisted outright on 2026-08-14, without ever converting to a paid rate the way Hy3 did — the shortest tenure of any listing tracked on this card, and a sign that even acquisition SKUs are treated as disposable rather than earning a grace period. Eleven days after that, on 2026-08-25, the sharp-repricing pattern crossed from images into tokens for the first time: Deepseek V4 Flash 0731 — a paid, established SKU, not a launch or a graduate — jumped 214% on input and 371% on output, the same no-notice mechanic that hit Flux.1 Kontext Pro a month earlier, just applied to an LLM instead of an image model. The same capture also proved the hardware churn isn’t one-directional: RTX 6000 Ada, delisted from self-serve GPUs on 2026-08-04, came back three weeks later at an unchanged rate — swapped in for a different tier this time — while the Image and Video catalogs were each cut by more than half. Three axes (LLM price, GPU lineup, media catalog) all moved in the same capture, which is the clearest evidence yet that Novita treats every row on the rate card, paid or free, hardware or token, as editable without warning. The very next capture, 2026-08-26, compressed that cadence from weeks to a single day: RTX 6000 Ada was delisted again just 24 hours after its 2026-08-25 return, in the same pass that finally reversed the biggest packaging call in this page’s history — H100 and L40S coming back to self-serve after 63 days off the page. Self-serve is no longer the strictly-cheapest rung either: H100 on-demand at $3.39/hr now costs more than the same GPU as a dedicated endpoint ($1.99) or bare-metal ($1.70), so “self-serve → dedicated → bare-metal” has stopped being a clean ascending ladder for at least one SKU.
What’s unique : One GPU, multiple prices, zero contracts
1. The same GPU is sold at multiple price points by packaging — and self-serve isn’t always the cheapest of the three. An NVIDIA H100 is $1.99/GPU-hr as a dedicated inference endpoint and $1.70/GPU/hr on an 8-GPU bare-metal node; an RTX 4090 is 0.33/hr as an on-demand instance, 0.17/hr spot, and $0.61/GPU-hr as a dedicated endpoint. Customers self-select down the cost curve as their commitment and isolation needs rise — with no annual contract required at any step. The 2026-06-24 repricing split the ladder by GPU class for a while: the laddered self-serve product became the RTX 4090 (instance → spot → dedicated), with data-center H100/H200 capacity reachable only through dedicated endpoints and bare-metal after Novita pulled L40S and H100 off the self-serve instance page entirely. That split held for 63 days — the 2026-08-26 capture put H100 SXM 80GB and L40S 48GB back on self-serve, but at $3.39/hr and $0.55/hr on-demand respectively, which makes self-serve H100 pricier than the same GPU as a dedicated endpoint ($1.99) or bare-metal ($1.70): the usual “commit more, pay less” logic inverts for at least this one SKU.
2. Four metering dimensions under one API. Tokens, GPU hours, GPU-hours-per-replica, and sandbox vCPU-seconds are all billed independently, so a single customer can mix shared inference, reserved compute, and ephemeral agent runtimes on one account and one invoice.
3. Per-second billing reaches all the way to agent sandboxes. Most inference clouds stop at per-hour GPU billing; Novita extends pure-usage granularity to per-second vCPU+memory for agent execution, quoting a 5-minute task at fractions of a cent.
4. Aggressive undercutting of first-party model APIs. With 174 listed models and cache-read discounts (~50% of input on many), Novita positions as the cheap default for open-weight DeepSeek, Qwen, GLM, Kimi, and Llama inference rather than going to each model maker directly — leaning into the token-cost deflation that keeps compressing per-million-token rates.
5. A rate card that now spans three orders of magnitude — and individual rows can move just as sharply. After the 2026-07-21 capture the same API serves Llama 3.1 8B at $0.05/M output and Kimi K3 at $15/M output — a 300x spread on one invoice, with no plan boundary between them. That range is the point: cheap models are the on-ramp, and a buyer whose quality bar rises can swap a model string rather than a vendor. The flip side is that a single mis-set model default is now a 300x cost event, which is why per-model spend caps matter more here than on a narrower catalog — and the 2026-07-29 10x repricing of Flux.1 Kontext Pro (from $0.036 to $0.36/image) shows the risk isn’t confined to switching models: an unchanged model string can absorb a full order-of-magnitude cost swing overnight. The 2026-08-25 capture confirms it isn’t confined to images either — Deepseek V4 Flash 0731, an already-paid LLM, jumped $0.14→$0.44/M input and $0.28→$1.32/M output (+214%/+371%) with no plan change and no new model ID involved.
Strengths & weaknesses
| Strengths | Weaknesses |
|---|---|
| Radical price transparency — 174 model rates + GPU + sandbox examples all public | Pricing surface is sprawling; comparing the “same” GPU across products is confusing |
| Pure usage, free to start, no monthly minimum or seat fees | Some flagship models show only “Tiered pricing” with no public rate |
| Spot GPUs at ~50% of on-demand, and a deep self-serve repricing (RTX 4090 on-demand cut ~48% across 2026-06-24 and 2026-07-06 to 0.33/hr) keep the entry rung aggressively cheap | No committed-use discounts or annual commit tier published |
| Per-second granularity down to agent sandboxes, now with a published per-unit rate card + $100/90-day starter credits | Bare-metal beyond H100/B200 and Enterprise sandbox are contact-sales only; self-serve GPU instances were RTX-class only for 63 days (2026-06-24 to 2026-08-25) and, now that H100/L40S are back as of 2026-08-26, self-serve H100 on-demand ($3.39/hr) actually costs more than the same GPU via dedicated endpoint ($1.99) or bare-metal ($1.70) |
| One API + one invoice across inference, GPUs, and agent runtimes | No public uptime SLA except the Enterprise sandbox tier (99.95%) |
| Rapid catalog refresh — new flagship models added within days of release (DeepSeek V4, Kimi K2.6, GLM-4.7, Kimi K3 on 2026-07-21; Veo 3.1 on 2026-07-29; DeepSeek V4 Pro 0813 on 2026-08-25) | Catalog churns out as fast as it churns in (50 media/audio SKUs delisted across 2026-07-21 and 2026-07-29, Image/Video cut by more than half again on 2026-08-25, Hy3 pulled off free pricing) and existing rows can reprice without notice — Flux.1 Kontext Pro jumped 10x on 2026-07-29 and Deepseek V4 Flash 0731 jumped 214%/371% on 2026-08-25, neither flagged on the rate card, and a $0 launch listing can vanish entirely within days: Ling 3.0 Tiny lasted just three days (2026-08-11 to 2026-08-14) before being delisted outright rather than repriced |
| Cheap on-ramp preserved even as the ceiling rises — Llama 3.1 8B still $0.02/$0.05 alongside Kimi K3 at $3/$15 | Undisclosed funding and thin public/community signal (no HN thread above single-digit points) make durability harder to assess |
Billing UX : Tabbed pricing surface, spot toggle, and sandbox estimator
- Tabbed pricing page — the
/en/pricingsurface splits into Serverless Endpoints, Dedicated Endpoints, Agent Sandbox, and GPUs tabs, each rendering its own rate tables. - Modality filters + model search on serverless rates — an All / LLM / Image / Audio / Video / AI Search / Cache filter row plus a “Search Model” box narrow the rate table (144 models on the
/en/modelscatalog counter as of 2026-08-25) to a single modality or model. - Batch-inference discount banner — a persistent line above the rate table advertises the introductory 50% discount on input and output tokens for supported models, with a “Learn More” link.
- On-Demand vs Spot toggle — the GPU Instance page shows on-demand and spot rates side by side per GPU, surfacing the ~50% spot saving.
- Pricing Calculator links — image and video rate tables explicitly point to a Pricing Calculator for dimension-dependent estimates (“varies based on image dimensions, inference steps, and upscaling”).
- Sandbox rate card + cost estimator — the Agent Sandbox page now publishes the per-unit billing basis ($0.0000098/vCPU-second, $0.0000032/GiB-second, $0.00009/GB-hour with the first 60 GB included) alongside sample configurations with workload examples and estimated per-task costs (e.g. ~$0.0034 / ~$0.0821 / ~$0.3744), plus a $100 / 90-day free-credit start.
- Scale-to-zero controls — dedicated endpoints advertise “per-second billing on running replicas only. Scale to zero, pay zero.”
Strategic wins : Why Novita’s packaging decisions work
1. Bundling inference and GPUs captures the whole workload
By selling both the model API and the GPU underneath it, Novita captures customers at whichever layer they prefer to operate — and can route them up or down the stack as needs change. This is the same full-funnel logic that makes hybrid platforms sticky, applied to pure usage.
2. GPU laddering by class monetizes commitment without contracts
Offering the same GPU at instance, dedicated, and bare-metal prices lets customers trade flexibility for cost on their own terms. The 2026-06-24 repricing sharpened that ladder along GPU class rather than blurring it, and the 2026-07-06 trim pushed it one notch further: the RTX 4090 now carries the full self-serve rung (on-demand 0.33 → spot 0.17 → dedicated $0.61), while for 63 days H100/H200 data-center capacity was sold only as dedicated endpoints or bare-metal. By cutting the RTX 4090 on-demand rate ~48% and pulling H100 off the self-serve instance page, Novita kept a deliberately cheap on-ramp for prosumers while routing serious-capacity buyers straight to the reserved products — earning more from each segment without an annual commitment. The 2026-08-26 capture complicates that logic: H100 is back as a self-serve instance at $3.39/hr on-demand, priced above the dedicated ($1.99) and bare-metal ($1.70) rates for the same GPU, so the “cheap on-ramp, reserved for volume” segmentation now coexists with a self-serve tier that’s the most expensive way to buy H100 rather than the cheapest.
3. Per-second sandbox pricing lands the agent wave early
Pricing agent execution per vCPU-second positions Novita squarely in front of the 2026 surge in coding agents, browser automation, and RL environments — a use case many GPU clouds price too coarsely to win, even as agentic workflows become a cost monster for buyers who don’t meter them tightly. The 2026-06-24 move doubled down here: the sandbox went from publishing only illustrative task costs to a full per-unit rate card ($0.0000098/vCPU-second, $0.0000032/GiB-second, $0.00009/GB-hour, first 60 GB free), and paired it with a $100 / 90-day no-card free-credit offer — a transparency-plus-acquisition play aimed at winning agent builders before they standardize on a competing runtime.
4. Free listings are treated as a launch ramp, not a permanent tier
Novita has used $0 model listings as an acquisition lever since at least the September 2025 rate card, and on 2026-07-21 it collected: Hy3 moved from Free to $0.14/M in · $0.58/M out, leaving no free LLM on the card. The mechanic is worth copying — a free listing buys trial traffic and benchmark coverage for a newly supported model, then converts to a paid rate once the model has an installed base, without ever promising a free tier that would have to be honored forever. Because the free-to-paid step lands per model rather than per account, it also avoids the classic freemium cliff: nobody’s account changes state, and a user on a graduating model is one string swap away from a still-cheap alternative such as Llama 3.1 8B at $0.02/$0.05. The cost is that “free” here is a temporary property of a SKU, so any team that built on a $0 listing should assume it is a trial rate with an expiry it will not be told about in advance. The 2026-08-11 capture ran the mechanic again — Macaron V1 Venti, Macaron V1 Tall, and Ling 3.0 Flash all graduated off Free to paid per-token rates as a new small model, Ling 3.0 Tiny, took the $0 slot — but 2026-08-14 showed a harder edge to the pattern: Ling 3.0 Tiny wasn’t repriced, it was delisted outright after just three days, the shortest tenure of any free listing tracked on this card. A team building on a free SKU here should now assume not just an eventual price, but a possible removal with no conversion step at all.
Areas to improve : Where the pricing surface creates friction
1. Consolidate the “same GPU, multiple prices” confusion — now worse after the lineup split
A buyer comparing H100 options sees $1.99 (dedicated endpoint) and $1.70 (bare-metal) across separate pages — and from 2026-06-24 to 2026-08-25 the self-serve instance page didn’t even list H100, only RTX-class GPUs — with no single comparison view. The lineup split actually raised the navigation burden: a buyer who landed on /en/gpus expecting a datacenter GPU had to discover that H100/H200 lived on entirely different products. The 2026-08-04 capture added a subtler version of the same problem within a single GPU class: RTX 6000 Ada 48GB was dropped from /en/gpus outright, and the base RTX 5090 32GB tier that effectively replaced it is priced at $0.73/hr — a cent above the adjacent “High frequency” RTX 5090 variant at $0.72/hr, inverting the RTX 4090 pair, where the HF variant costs roughly double the base rate. A buyer skimming the table for the “premium” option has no label-based way to tell that the plain-named tier now costs more. The 2026-08-25 capture showed the lineup isn’t even stable in one direction: RTX 6000 Ada 48GB returned to /en/gpus at its old $0.77/hr rate, swapped in for the delisted RTX 4090 24GB “High frequency” tier — and the very next day, 2026-08-26, it was delisted again, this time to make room for H100 SXM 80GB and L40S 48GB returning to self-serve after 63 days off the page. That’s four GPU-class edits to the same page in three months (2026-06-24, 2026-08-04, 2026-08-25, 2026-08-26), the last two on consecutive days, each changing which row is the priciest self-serve option with no changelog marking any of them — and the H100 return makes the comparison problem worse, not better: a buyer now sees H100 at $3.39/hr on /en/gpus, $1.99/hr as a dedicated endpoint, and $1.70/hr on bare-metal, with the self-serve price the highest of the three and no page explaining why the “least committed” option costs the most. A unified “choose your GPU packaging” matrix — which class is available as instance vs dedicated vs bare-metal, with break-even hours — would convert better than forcing the buyer to reconcile separate tabs, and would blunt the bill-shock and cost-unpredictability risk that scares finance teams away from raw usage pricing.
2. Publish rates for “Tiered pricing” models
Flagship models like Qwen3 Max and MiniMax-M3 show only “Tiered pricing” with no number, undercutting the otherwise-excellent transparency. Even a starting rate or a published tier table would remove a sales-friction point for the most in-demand models.
3. Offer a committed-use discount tier
Novita has no public annual-commit or reserved-capacity discount outside bare-metal. A self-serve committed-use option (prepay credits for a discount) would give predictable workloads a reason to consolidate spend on Novita, echoing how credit-pool models reward commitment.
4. Publish a deprecation window for delisted models, graduating free rates — and sudden repricings
The 2026-07-21 capture removed 25 media and audio SKUs and moved Hy3 from Free to $0.14/$0.58 with nothing on the rate card signalling either change in advance. On 2026-07-29 the pattern repeated and widened: 25 more media/video SKUs were delisted, and — more consequential for anyone already paying — Flux.1 Kontext Pro’s per-image rate rose 10x, from $0.036 to $0.36, with no changelog entry distinguishing it from routine catalog churn. Each move is individually defensible — the delisted models mostly have listed successors, free listings were never sold as permanent, and $0.036 vs $0.36 is easy to miss in a decimal-dense rate table — but a buyer whose product hardcodes Kling v2.1, a $0 model, or Flux.1 Kontext Pro as a default finds out by way of a failed request or a 10x line item, not a warning. The 2026-08-14 capture adds a new wrinkle: Ling 3.0 Tiny, a free listing added just three days earlier on 2026-08-11, was pulled outright rather than converted to a paid rate — so even the graduation pattern buyers might have learned to expect (free, then priced) isn’t guaranteed; a listing can simply disappear. The 2026-08-25 capture confirms this isn’t just an image or free-tier problem: Deepseek V4 Flash 0731 — an established, already-paid LLM SKU — repriced 214%/371% (to $0.44/M in · $1.32/M out) in the same capture that cut the Image catalog from 14 to 5 SKUs and restructured Video, with no distinction anywhere on the page between “new model added,” “model delisted,” and “existing paid rate tripled.” A dated “sunset” column on the rate table, or a documented repricing-notice window (even 7 days) for existing paid SKUs, would cost Novita nothing and would remove the one thing that makes an otherwise very transparent surface feel unpredictable to plan against.
Monetization stack & signals : how Novita AI builds & buys its revenue engine
2 signal roles
No monetization vendor is sourceable — the read is the GTM build-out below. Novita's first quota-carrying AE, paired with a presales Solutions Engineer, layers an enterprise sales-led motion onto its self-serve, public-rate-card inference cloud.
- Account Executive Deal desk seen Apr 1, 2026
Novita's first dedicated quota-carrying sales hire — a full-cycle AE closing CTO/VP-Eng deals — layers an enterprise sales-led motion onto its self-serve, published-rate-card core. The deal desk and negotiated contracts this implies sit alongside, not instead of, the public per-token/per-GPU-hour metering.
“Own the entire sales process from prospecting and lead generation to negotiation and closing. Conduct deep-dive discovery calls with technical buyers (CTOs, VP of Engineering, Lead AI Scientists).”
- Solutions Engineer (AI Cloud Infrastructure) Customer success seen Aug 27, 2025
A presales SE paired with the AE confirms the enterprise GTM build-out: POC-led technical selling of GPU/inference workloads, with the SE owning post-sale onboarding. The pairing is the canonical two-role enterprise motion bolted onto a PLG inference cloud.
“Partner with Account Executives to deeply understand customer needs... Design, manage, and execute successful POCs, proving the value and performance of our platform... provide initial onboarding support and architectural guidance.”
Signals reviewed · derived from public job posts
Job postings fill and close over time — once a posting is filled we keep it as a dated citation (the quoted evidence remains); use View open roles for current listings.
Key takeaways
- Sell the same resource at multiple price points by packaging — but check that the ladder still ascends with commitment. Novita’s instance/dedicated/bare-metal packaging shows how to monetize commitment without contracts, and how trimming a self-serve lineup (H100/H200 reserved for dedicated and bare-metal only, 2026-06-24 to 2026-08-25) can steer high-end demand toward higher-commitment products. The 2026-08-26 reversal is the cautionary half of the lesson: once H100 came back to self-serve at $3.39/hr, it priced above the dedicated ($1.99) and bare-metal ($1.70) versions of the same GPU — a packaging ladder only works as a demand-shaping tool if the self-serve rung stays the cheapest rung.
- Transparency is a competitive weapon in usage-based markets. Publishing 174 model rates plus GPU and sandbox examples lets developers self-qualify and reduces sales friction versus gated competitors.
- Extend metering granularity to match the workload. Per-second sandbox billing fits agent execution far better than per-hour GPU billing — the unit should mirror how the customer actually consumes.
- Cache-read discounts reshape effective token cost. Listing cache-read rates (~50% of input) on many models materially changes real bills and should be modeled, not ignored.
- “Free to start” still needs a cost model — and a free listing is a launch tactic, not a tier. The headline hides four compounding dimensions, so the buyers who win model token, GPU, and sandbox spend together before scaling. Hy3’s move from $0/M to $0.14/$0.58 on 2026-07-21 emptied Novita’s free-LLM shelf on the same day 25 SKUs were delisted: price the paid fallback before you ship on a $0 or long-tail model. The 2026-07-29 capture showed the same discipline applies to paid SKUs too — Flux.1 Kontext Pro’s per-image rate rose 10x with no changelog entry, so treat any published per-model rate as provisional and keep your own dated pricing evidence rather than trusting a stable number to stay stable. Novita restocked the shelf on 2026-08-11 with Ling 3.0 Tiny, then emptied it again on 2026-08-14 by delisting the model outright rather than repricing it — the fastest free-to-gone cycle tracked yet, and proof the shelf can go empty even without a graduation step to warn you. On 2026-08-25 the same no-warning repricing hit a paid LLM directly — Deepseek V4 Flash 0731 rose 214%/371% — confirming the risk sits on the token side of the business, not just images and free tiers, so treat every dimension (tokens, images, video, GPUs) as capable of moving without notice.
UBP implications
- Multi-dimensional metering is becoming table stakes for AI clouds. Novita bills on tokens, GPU hours, replica-hours, and vCPU-seconds simultaneously — a sign that single-metric usage pricing is insufficient for stacked AI infrastructure.
- Packaging, not just price, is the lever — and which SKUs you withdraw (or reprice) is part of it. The same silicon at multiple prices proves that how usage is packaged (shared vs dedicated vs reserved) is itself a pricing dimension; Novita’s 2026-06-24 move to pull H100 off self-serve while cutting the RTX 4090 ~48% shows that removing a packaging option can be as deliberate a lever as adding one. The 2026-07-21 capture generalizes it: one model added, 25 delisted, and the last free LLM rate retired — all without touching a plan boundary. The 2026-07-29 capture pushed the same lever again, and this time on price rather than availability: one model added (Veo 3.1), 25 more delisted, and an already-paid SKU (Flux.1 Kontext Pro) repriced 10x — still without touching a plan boundary. The 2026-08-04 capture applied the same lever to hardware: RTX 6000 Ada was withdrawn from self-serve GPUs and a new RTX 5090 base tier landed priced above its own “High frequency” sibling — proof the packaging-as-lever pattern isn’t confined to the model catalog. The 2026-08-11 and 2026-08-14 captures completed the loop on the model side: a free listing (Ling 3.0 Tiny) was added, then withdrawn entirely three days later with no conversion step — confirming that even acquisition-oriented $0 SKUs are edited on the same no-notice cadence as paid rows. The 2026-08-25 capture shows the lever pulled on three axes in a single pass: RTX 6000 Ada was re-added to self-serve GPUs (reversing the 2026-08-04 withdrawal) in place of a different tier, an already-paid LLM (Deepseek V4 Flash 0731) was repriced 214%/371%, and the Image and Video catalogs were each cut by roughly half — packaging, price, and catalog breadth all moving together, none of it flagged in advance. The very next capture, 2026-08-26, shows the lever can flip within 24 hours: RTX 6000 Ada was withdrawn again, and Novita reversed its own biggest packaging call — pulling H100 and L40S back onto self-serve after 63 days off the page — in the same pass, at a self-serve H100 rate ($3.39/hr) that now sits above the dedicated and bare-metal prices for the identical GPU. Where seat-based vendors must run a migration to change what a customer pays, a pure-usage catalog reprices by editing a row, which makes the listing, not the contract, the unit of pricing change — and means buyers should monitor rate cards the way they’d monitor a renewal.
- Per-second billing is the new frontier for agent economies. As AI agents proliferate, the vendors that meter execution at second-level granularity will out-price those still selling hourly compute blocks.
Sources
- Novita AI pricing page (accessed 2026-08-11)
- Novita AI GPU Instance pricing (accessed 2026-08-11)
- Novita AI Bare Metal GPU servers (accessed 2026-08-11)
- Novita AI Agent Sandbox (accessed 2026-08-11)
- Novita AI model catalog (accessed 2026-08-11)
- Novita AI docs — Agent Sandbox pricing (accessed 2026-08-27)
- Novita AI docs — Dedicated Endpoint billing (accessed 2026-08-27)
- Novita AI docs — GPU Instance pricing and storage rates (accessed 2026-08-14)
- Novita AI docs — LLM billing (accessed 2026-08-14)
- Novita AI public model-list API (per-token rates) (accessed 2026-08-27)
- Novita AI Bare Metal GPU product page (independent verification) (accessed 2026-08-27)
- Novita AI LLM API product page (accessed 2026-08-27)
- Novita AI Dedicated Endpoints product page (accessed 2026-08-27)
Bottom line
Novita AI is one of the most transparent pure-usage AI clouds in the corpus: 174 model rates, on-demand and spot GPUs, bare-metal nodes, and per-second agent sandboxes (now with a published per-unit rate card) are all open, with GPUs laddered across price points by packaging so customers buy exactly the commitment they want. The 2026-06-24 repricing sharpened that ladder by class — a deeply cut RTX-class self-serve on-ramp (RTX 4090 on-demand ~48% cheaper, at 0.33/hr after a further 2026-07-06 trim) with H100/H200 reserved for dedicated and bare-metal. The 2026-07-21 capture stretched the other axis: Kimi K3 at $3/$15 per M set a new ceiling roughly 300x the cheapest listed output rate, Hy3 graduated off free pricing, and 25 superseded media SKUs were struck from the catalog. Eight days later, on 2026-07-29, the same instinct turned on an already-paid SKU: Flux.1 Kontext Pro’s image rate rose 10x to $0.36, and the catalog was pruned again to 172 as Veo 3.1 arrived and two dozen more legacy video SKUs retired. Six days after that, on 2026-08-04, the catalog rebuilt to 177 models and the edit-without-notice habit showed up on the hardware side too: RTX 6000 Ada was pulled from self-serve GPUs and replaced by a new RTX 5090 tier priced a cent above its own “High frequency” sibling. A week later, the free-LLM shelf itself became disposable: three Time Limited Free models graduated to paid rates on 2026-08-11 as Ling 3.0 Tiny took the $0 slot, only for Ling 3.0 Tiny to be delisted outright on 2026-08-14 — never repriced, just removed — the shortest-lived free listing tracked on this card and the reason the catalog eased to 174 models with no free LLM on it. Eleven days later, on 2026-08-25, the no-notice repricing habit reached the token side directly for the first time: Deepseek V4 Flash 0731, an established paid LLM, jumped 214% on input and 371% on output to $0.44/$1.32 per M tokens, in the same capture that saw RTX 6000 Ada return to self-serve GPUs (reversing its own 2026-08-04 delisting) and the Image and Video catalogs each cut by more than half. One day after that, on 2026-08-26, the GPU lineup turned over again — RTX 6000 Ada delisted a second time, and H100 SXM 80GB plus L40S 48GB restored to self-serve after 63 days off the page, this time at a self-serve H100 rate ($3.39/hr) that costs more than buying the same GPU as a dedicated endpoint ($1.99) or bare-metal node ($1.70). The strategy is coherent — cheap on-ramp, frontier-class top end, loss-leaders retired once they’ve paid for themselves — but the buyer-side cost is now three-sided: a sprawling multi-tab surface that asks you to reconcile token, GPU, and sandbox math yourself; a rate card that adds, removes, and reprices SKUs — on any dimension, including established paid tokens — without advance notice; and, as of 2026-08-26, a self-serve GPU ladder that no longer guarantees the least-committed option is the cheapest one. Worth planning against with a model default you’ve priced and a fallback you’ve tested.
Want to compare Novita AI against other AI-infrastructure pricing? Browse the pricing blueprint.
Pricing timeline : Major events on a vertical axis
Each milestone below corresponds to a public pricing change, product launch, or material adjustment. Major events use a filled marker; minor adjustments use a faded one.
H100 SXM and L40S return to self-serve GPU instances; RTX 6000 Ada and RTX 5090 HF variant removed
The /en/gpus self-serve lineup regained two data-center-class GPUs pulled on 2026-06-24: NVIDIA H100 SXM 80GB (new — $3.39/hr on-demand, $1.70/hr spot) and NVIDIA L40S 48GB (new — $0.55/hr on-demand, no spot rate shown). In the same capture, RTX 6000 Ada 48GB — itself only re-added the day before, 2026-08-25, at $0.77/hr — was delisted again, and the RTX 5090 'High frequency' variant ($0.72/hr) was removed, leaving a single RTX 5090 32GB listing at $0.73/hr on-demand / $0.37/hr spot. Serverless model rates, dedicated endpoints, bare-metal, and Agent Sandbox rate cards were all unchanged.
Deepseek V4 Flash 0731 repriced sharply; RTX 6000 Ada returns; Image/Video catalog cut hard
Deepseek V4 Flash 0731 jumped from $0.14/$0.28 to $0.44/$1.32 per M tokens (input +214%, output +371%) as a new DeepSeek V4 Pro 0813 SKU ($1.32/$3.96) landed alongside it. On the self-serve GPU page, RTX 6000 Ada 48GB returned at $0.77/hr in place of the delisted RTX 4090 'High frequency' tier ($0.69/hr). The Image catalog shrank from 14 to 5 SKUs (Z Image Turbo, Seedream 4.0, Image Eraser/Remove Background/Upscaler, and the FLUX 2 family all delisted) and the Video catalog dropped Veo 3.1, the Vidu Q2/Q3 lineup, older Kling versions (V1.6/V2.5 Turbo/V2.6 Pro), Wan 2.1/2.2, and Seedance 1.5 Pro in favor of the newer Kling v3.0 and Wan 2.5–2.7 lines. The `/en/models` catalog counter reads 144, versus 174 logged on 2026-08-14 (methodology not reconciled by this capture).
Ling 3.0 Tiny delisted after 3 days, emptying the free-LLM shelf
Ling 3.0 Tiny — the free LLM added 2026-08-11 — was removed outright from the pricing page and model catalog rather than converted to a paid rate, the shortest-lived free listing tracked on this rate card. The catalog eased from 177 to 174 models; Qwen3.7-Max's 'Limited Time 50% Off' tag also disappeared from the rate table though its $1.25/$3.75 rate held steady. All other tracked per-token, GPU, dedicated-endpoint, bare-metal, and Agent Sandbox rates were unchanged.
Three free LLMs graduate to paid pricing; Ling 3.0 Tiny takes the $0 slot
Macaron V1 Venti, Macaron V1 Tall, and Ling 3.0 Flash moved off their introductory Time Limited Free tag to paid per-token rates (Macaron V1 Venti $1.5/M in · $4.5/M out; Macaron V1 Tall $0.45/M in · $2.6/M out; Ling 3.0 Flash $0.06/M in · $0.18/M out), while a new small model, Ling 3.0 Tiny, was added as the new $0 free-launch listing under the same tag. Catalog held at 177 models; GPU, dedicated-endpoint, bare-metal, and Agent Sandbox rates were unchanged.
RTX 6000 Ada dropped from self-serve GPUs; base RTX 5090 tier added at $0.73/hr
The /en/gpus self-serve lineup lost RTX 6000 Ada 48GB (was $0.77/hr on-demand, $0.39 spot) and gained a new base RTX 5090 32GB on-demand tier at $0.73/hr ($0.37 spot) — priced a cent above the existing 'High frequency' RTX 5090 variant at $0.72/hr, an inversion of the RTX 4090 pair where the HF variant costs roughly double the base rate. Dedicated endpoints, bare-metal, and Agent Sandbox rate cards held steady; the model catalog grew from 172 to 177 with additions including Deepseek V3.2, Deepseek V4 Flash, DeepSeek-OCR 2, and PaddleOCR-VL.
Flux.1 Kontext Pro image price jumps 10x; catalog pruned 196 → 172
Flux.1 Kontext Pro image-generation pricing rose 10x, from $0.036 to $0.36/image (confirmed against three independent prior captures), while Kontext Max held at $0.072. The model catalog shrank from 196 to 172 as Hunyuan Video Fast, the PixVerse V4.5 family, and the entire Kling-o1 lineup were delisted; Google's Veo 3.1 was added.
Kimi K3 lands at $3/$15; Hy3 exits free; catalog pruned 221 → 196
Moonshot's Kimi K3 is added at $3/M input ($0.3/M cache read) and $15/M output on a 1,048,576-token context — the most expensive LLM on the rate card. Tencent's Hy3 moves off free pricing to $0.14/M input ($0.035/M cache read) and $0.58/M output. The catalog shrinks from 221 to 196 models as legacy media/audio SKUs are delisted (GLM Image Generation, GLM Audio-to-Text/TTS/Voice-Clone, Hunyuan Image 3, Seedream 3.0, Seedance V1 Lite/Pro, Vidu 2.0/Q1, Kling v2.1, MiniMax speech-02 and speech-2.5 preview). The RTX 5090 32GB instance drops its 'High frequency' label and ticks to 0.73/hr on-demand (0.37 spot) from 0.72/0.36. Dedicated endpoints, bare-metal, and Agent Sandbox rate cards unchanged.
Base RTX 4090 instance trimmed; catalog refresh
Self-serve RTX 4090 24GB instance nudged down to $0.33/hr on-demand ($0.17 spot) from $0.35/$0.18; RTX 4090 HF ($0.69/$0.35), RTX 5090 HF ($0.72/$0.36) and RTX 6000 Ada ($0.77/$0.39) held. Serverless catalog refreshed with new flagships (GLM 5.2 $1.4/$4.4, GLM-4.7 $0.6/$2.2, Kimi K2.7 Code $0.95/$4, MiniMax M2.7/M2.5) while referenced SKUs (DeepSeek V3.1/V4 Pro, Llama 3.1 8B, Qwen3.7-Max, GLM-5.1, Kimi K2.6, GPT OSS 120B) held rates. Dedicated endpoints, bare-metal, and Agent Sandbox rate cards unchanged.
GPU-instance repricing + published sandbox rate card
Self-serve GPU instances (/en/gpus) cut and re-lined-up to RTX-class only: RTX 4090 24GB drops to $0.35/hr on-demand ($0.18 spot) from $0.67/$0.34, with RTX 4090 HF $0.69, RTX 5090 HF $0.72, and a new RTX 6000 Ada 48GB at $0.77 — L40S and H100 SXM instances removed (H100 now sold via dedicated endpoints $1.99/GPU-hr and bare-metal $1.70/GPU/hr). Dedicated endpoints add RTX-5090 at $0.73/GPU-hr. Agent Sandbox now publishes per-unit rates ($0.0000098/vCPU-second, $0.0000032/GiB-second, $0.00009/GB-hour with first 60 GB free) and a $100 / 90-day free-credit offer.
Pricing snapshot — per-token inference, per-hour GPU, per-second sandbox
Novita lists 226 models on per-token/per-image/per-video pricing; on-demand GPUs from $0.55/hr; bare-metal H100 $1.70 and B200 $4.77 per GPU/hr; dedicated endpoints H100 $1.99 / H200 $2.99 per GPU-hour; agent sandbox billed per-second on vCPU + memory.
GLM-4.6, Kimi K2, DeepSeek V3.2 Exp added; output-token cuts
Catalog expands with zai-org/glm-4.6 ($0.6/$2.2), moonshotai/kimi-k2-instruct ($0.57/$2.3) and deepseek-v3.2-exp at $0.27/$0.41 — a sharp output-token cut versus V3.1's $1. DeepSeek V3.1 input/output held at $0.27/$1. Wayback web.archive.org/web/20251003030153/novita.ai/pricing.
Full serverless LLM rate card published
Serverless Endpoints tab renders the full per-million-token catalog: DeepSeek V3.1 $0.27/$1, Llama 3.1 8B $0.02/$0.05, Llama 3.3 70B $0.13/$0.39, GLM-4.5 $0.6/$2.2, Qwen3-Coder-480B $0.29/$1.2, with several small models (Llama 3.2 1B, Qwen3 4B, Gemma 3 1B) listed Free. Wayback web.archive.org/web/20250911221805/novita.ai/pricing.
Agent Sandbox launches; pricing splits into four endpoint tabs
Pricing page restructured into Serverless Endpoints / Dedicated Endpoints / GPUs / Agent Sandbox tabs — the per-second Agent Sandbox product is now live. Batch inference advertised at an introductory 50% token discount. Image rate back to $0.001/image (512×512, 5 steps); MiniMax speech-02-hd $80/1M characters, Voice-Cloning $2.4/voice. Wayback web.archive.org/web/20250804005313/novita.ai/pricing.
Redesign to Model APIs + GPUs; HQ moves to San Francisco
Pricing page redesigned to a light-theme 'Model APIs / GPUs' tabbed layout; footer HQ changes to 156 2nd Street, San Francisco. A '$10,000 in credits' startup program launches. Image rate shown at $0.003/image (1024×1024, 20 steps); Text to Speech $15/1M characters. Wayback web.archive.org/web/20250210165912/novita.ai/pricing.
LLM API and Audio added alongside image/video
An 'LLMs' tab and an LLM API product appear in the catalog; Audio (Text to Speech, Voice Cloning) is added. Still credit-based with the same image rates (txt2img $0.0015/image). Wayback web.archive.org/web/20240521203043/novita.ai/pricing.
Credit-funded image-generation API (Singapore)
Earliest archived pricing surface: Novita was a Stable-Diffusion image API billed from a prepaid USDT/Stripe credit balance, tagline 'The price is only 1/10 of DALL-E2 and MJ.' Text-to-image quoted at $0.0015/image (512×512, 20 steps), upscale from $0.0021/image. Footer listed a Singapore HQ (14 Robinson Road). Wayback web.archive.org/web/20240225192128/novita.ai/pricing.
- · Novita publishes per-second billing for both GPU instances and agent sandboxes — a 5-minute coding-agent task on 1 vCPU + 512 MiB RAM is quoted at roughly $0.0034.
- · The same NVIDIA H100 now appears at three different prices depending on product: $3.39/hr on-demand as a self-serve instance, $1.99/GPU-hour as a dedicated endpoint, and $1.70/GPU/hr on an 8-GPU bare-metal node — and its self-serve spot rate ($1.70/hr) happens to land exactly on the bare-metal on-demand rate.
- · Novita lists 174 models on its catalog and undercuts first-party APIs — DeepSeek V3.1 runs $0.27 input / $1 output per million tokens versus DeepSeek's own rates.
Questions & answers
- How does Novita AI pricing work?
- Novita is pure pay-as-you-go. Model inference is billed per million tokens (LLMs), per image, per video-second, or per 1M characters (audio); GPUs are billed per hour (on-demand or spot); and agent sandboxes are billed per second on vCPU and memory. There is no monthly minimum and you can start for free.
- How much does an NVIDIA H100 cost on Novita AI?
- It depends on the product. A dedicated inference endpoint on H100 80GB is $1.99 per GPU-hour, and an 8-GPU bare-metal H100 SXM node is $1.70 per GPU/hr. As of the 2026-08-26 capture, H100 SXM 80GB is also back on Novita's self-serve GPU-instance page at $3.39/hr on-demand ($1.70/hr spot) — having been RTX-class only (RTX 4090 from $0.33/hr on-demand, $0.17/hr spot) since a 2026-06-24 repricing removed it.
- Does Novita AI have a free tier?
- Yes, in the sense that there is no subscription — Novita advertises 'Free to start, scales as you grow', so you sign up and pay only for the tokens, images, GPU hours, or sandbox seconds you consume, and new Agent Sandbox users get $100 in credits valid for 90 days with no card required. There is no free LLM on the rate card as of the 2026-08-14 capture: Tencent's Hy3 was the last long-standing $0 language model, moving to $0.14/M input and $0.58/M output on 2026-07-21; a replacement free listing, Ling 3.0 Tiny, briefly held the $0 slot from 2026-08-11 but was delisted outright by 2026-08-14.
- What is the cheapest LLM on Novita AI?
- Among flagship-class models, Llama 3.1 8B Instruct is among the cheapest at $0.02/M input and $0.05/M output. Larger models like DeepSeek V4 Pro run $1.6/M input and $3.2/M output, and the priciest listed LLM is Kimi K3 at $3/M input and $15/M output. The cheapest option in a given modality can change between captures: Novita's catalog fell from 221 models (2026-07-06) to 196 (2026-07-21) to 172 (2026-07-29) as legacy media, video, and audio SKUs were delisted and Tencent's Hy3 moved off free pricing, then grew back to 177 (2026-08-04, held at 2026-08-11) as new LLM and OCR SKUs (Deepseek V3.2, DeepSeek-OCR 2, PaddleOCR-VL) landed, then eased to 174 (2026-08-14) as the brief Ling 3.0 Tiny free listing was delisted outright — and on 2026-07-29 Flux.1 Kontext Pro's image rate itself rose 10x with no changelog notice — so teams calling long-tail, $0, or even established paid models should keep a tested paid fallback and re-check rates rather than assume a published price is fixed.
- How is the Novita Agent Sandbox billed?
- The Agent Sandbox is billed per second based on the vCPU and memory you allocate, with no plans or lock-ins. Example quotes include ~$0.0034 for a 5-minute task on 1 vCPU + 512 MiB RAM and ~$0.3744 for a 1-hour multi-agent workload on 8 vCPU + 8 GiB RAM.
- Does Novita AI offer bare-metal GPU servers?
- Yes. Bare-metal nodes ship 8 GPUs per node with zero virtualization overhead and contractual SLAs. Published rates are $1.70/GPU/hr for H100 SXM and $4.77/GPU/hr for B200 SXM; H200, RTX 5090, and RTX 4090 nodes are quoted via contact sales.