Hyperbolic drops its Serverless Inference marketing page and shrinks the live model rate card
Hyperbolic's dedicated Serverless Inference pricing page now 404s; its docs show only two priced models ($0.40/M each) and mark the entire image-model catalog plus several text/vision models for sunset.
hyperbolic.ai/inference listed a 5-tier per-model rate card from $0.10/M (small Llama) to $4.00/M (Llama-3.1-405B), with image and VLM modality tabs alongside it.
hyperbolic.ai/inference returns HTTP 404 and is unlinked from site nav. Docs (hyperbolic.ai/docs/inference/*) show only 2 priced Instruct SKUs — Llama 3.3 70B and Qwen3-Coder 480B, both $0.40/M tokens — against unchanged 'from' floors ($0.10/M text, $0.0025/image, $0.15/M VLM). Every listed image model (FLUX.1-dev, SDXL 1.0/Turbo, SD 1.5/2, Segmind SD 1B) is now flagged 'Sunset', as are several text and vision-language SKUs (GPT-OSS 120B/20B, Qwen3-Next 80B, Llama-3.1-405B BASE, Qwen2.5-VL-72B/7B, Pixtral 12B, Nemotron Nano 12B VL). GPU marketplace rates are unchanged (H100 SXM $3.19, H200 $3.99, B200 $5.99/GPU-hr).
Hyperbolic’s standalone Serverless Inference marketing page (hyperbolic.ai/inference), which previously hosted the product’s full per-model rate card, now returns a genuine HTTP 404 and is no longer linked from the site’s navigation or homepage. Per-model pricing has moved to Hyperbolic’s Mintlify docs, but the live catalog shown there is much narrower than what the marketing page used to publish: the “Available Models” table for chat completions lists just two SKUs — Llama 3.3 70B and Qwen3-Coder 480B, both priced at $0.40 per million tokens with tool-calling support — against broader “from” claims of $0.10/M for text, $0.0025/image for image generation, and $0.15/M for vision-language models.
The docs also disclose a wave of upcoming model deprecations: every image-generation model Hyperbolic lists (FLUX.1-dev, SDXL 1.0, SDXL Turbo, Stable Diffusion 1.5/2, Segmind SD 1B) is now flagged “Sunset” with no replacement model named, and several text and vision-language SKUs — GPT-OSS 120B/20B, Qwen3-Next 80B (Thinking and Instruct), the Llama-3.1-405B BASE completions model, Qwen2.5-VL-72B/7B-Instruct, Pixtral 12B, and NVIDIA Nemotron Nano 12B VL — are being retired too. GPU marketplace pricing is unaffected: on-demand starting rates remain H100 SXM $3.19, H200 $3.99, and B200 $5.99 per GPU-hour, unchanged from the prior capture.
On-demand starting rates move again: H100 SXM $2.89 to $3.19/GPU-hr (+10.4%), H200 $3.49 to $3.99 (+14.3%); B200 holds at $5.99. The marketplace's 'starting at' claim, previously an unverifiable $0.20/GPU/hr floor with no matching card, now equals the H100 SXM rate exactly. Serverless inference rates and account-tier mechanics are unchanged.