SambaNova cuts gemma to /bin/bash.22 and adds cached-input pricing
SambaNova Cloud trimmed its rate card to six models, cut gemma-4-31B-it from /bin/bash.38/.15 to /bin/bash.22//bin/bash.59, and added a Cached Input Tokens column led by MiniMax-M2.7 at /bin/bash.06/1M cached.
gemma-4-31B-it /bin/bash.38 in / .15 out; no cached-input rate; DeepSeek-V3.1-cb (/bin/bash.15//bin/bash.75) and DeepSeek-R1-Distill-Llama-70B (/bin/bash.70/.40) on the card
gemma-4-31B-it /bin/bash.22 in / /bin/bash.59 out; new Cached Input Tokens column (MiniMax-M2.7 /bin/bash.06/1M cached); DeepSeek-V3.1-cb and DeepSeek-R1-Distill dropped; DeepSeek-V3.1/V3.2 stay .00/.50
SambaNova Cloud’s public per-1M-token rate card was refreshed in July 2026. The list of self-serve models was trimmed to six, and gemma-4-31B-it dropped roughly 42% on input (to /bin/bash.22) and 49% on output (to /bin/bash.59), matching gpt-oss-120b as the cheapest models on the card.
The bigger structural change is a new Cached Input Tokens column: MiniMax-M2.7 is the first model to carry a cached-input rate, at /bin/bash.06/1M versus /bin/bash.60/1M uncached — a 10x discount for cached prompt tokens. Two models that anchored the old low end (DeepSeek-V3.1-cb at /bin/bash.15//bin/bash.75) and the reasoning tier (DeepSeek-R1-Distill-Llama-70B at /bin/bash.70/.40) were removed from the public card, while the DeepSeek-V3.1/V3.2 frontier tier holds at .00/.50.
July 2026 rate card trimmed to six models: gemma-4-31B-it cut from $0.38/$1.15 to $0.22/$0.59; DeepSeek-V3.1-cb and DeepSeek-R1-Distill dropped; a new Cached Input Tokens column debuts with MiniMax-M2.7 at $0.06/1M cached. DeepSeek-V3.1/V3.2 stay $3.00/$4.50, Llama-3.3-70B $0.60/$1.20. SambaNova closed a $1B round at an $11B valuation (General Atlantic) on July 8, 2026.