What is it
Data Platform Pricing is pricing for data platforms — scraping, enrichment, search API, and knowledge-graph vendors.
These are the companies that sell (or store, or search) the structured data AI systems need for grounding, enrichment, retrieval, and training. The corpus tracks 33 of them, and the category spans five overlapping sub-types: web-extraction and proxy platforms (Bright Data, Oxylabs, ScraperAPI, Firecrawl, Apify, Browse AI); search and knowledge APIs (SerpApi, Linkup, Diffbot, Nomic); GTM enrichment and analytics (Clay, Julius AI, Powerdrill, Rows); vector and search databases (Pinecone, Weaviate, Qdrant, Chroma, Milvus, turbopuffer, Upstash); and the training-data and infrastructure layer underneath (Scale AI, Labelbox, Snorkel AI, Mercor, Unstructured, OpenMeter).
What unites the category is a structural fact: the cost driver is volume of data, not headcount. A single account can run millions of automated requests, store terabytes of vectors, or label hundreds of thousands of rows, so seat-based pricing has no relationship to the cost of service. Every self-serve company on this page therefore meters something — requests, bandwidth, records, credits, stored dimensions, or query bytes — even when it wraps that meter in a monthly subscription tier.
The category is also consolidating and repricing fast: ScraperAPI was acquired into SaaS.group, the proxy and scraping vendors are converging on near-identical rate cards, and prices are trending down — the opposite of most software. With the meter itself commoditized, differentiation has moved to the generosity of the included volume, the steepness of the overage curve, and how honestly the unit math is exposed on the public pricing page. For the strategy behind these choices, see our introduction to usage-based pricing and the guide to choosing the right usage metric.
How it works
Data platforms pick a meter that tracks their dominant cost, then wrap it in either pure pay-as-you-go or a subscription tier that pre-buys volume. Six common meters and their canonical examples:
| Billing unit | What it counts | Example rate | Companies |
|---|---|---|---|
| Bandwidth (per GB) | Data transferred through proxies/browsers | $8 → $2.50/GB residential (Bright Data) | Bright Data, Oxylabs, Apify |
| Requests / results (per 1K) | Successful API responses | $0.50–$1.15 / 1K results (Oxylabs) | Oxylabs, Bright Data, SerpApi |
| Credits (per pool) | Abstracted unit, ≈1 page or 1 call | $0.001/credit (Diffbot Startup) | Firecrawl, ScraperAPI, Diffbot, Clay |
| Records (per 1K) | Rows of ready-made data | $2.50 / 1K records (Bright Data datasets) | Bright Data, Labelbox |
| Stored dimensions / vectors | Vector index size held | $0.00465 / 1M dims·mo (Weaviate) | Weaviate, Pinecone, turbopuffer |
| Query / storage | Bytes scanned + GB-month stored | $1/PB queried (turbopuffer) | turbopuffer, Upstash, Pinecone |
Worked example — credit-pool scraping. Firecrawl’s Standard plan is $83/mo for 100,000 credits, where 1 credit ≈ 1 scraped page (Search costs 2 credits per 10 results). A pipeline scraping 95,000 pages/month stays inside the pool and pays the flat $83. Cross to 110,000 pages and the overflow auto-recharges at the plan’s per-credit rate — the model behaves like a subscription until you exceed the pool, then like pure usage. Its full ladder runs Free ($0, 1,000 credits), Hobby ($16, 5,000), Standard ($83, 100,000), Growth ($333, 500,000), and Scale ($599, 1,000,000). You can model your own pool with the Firecrawl pricing calculator.
Worked example — per-GB proxy. Oxylabs residential proxies start at $6/GB on the $30/mo Starter (5GB) and decline to $2.50/GB on the $2,500/mo Corporate tier; its cheapest datacenter traffic runs $0.59/GB — roughly a 10x spread driven by IP type, not just volume. Moving 200GB of residential traffic costs on the order of $500–$1,200 depending on the committed tier. The unit math is nearly identical to Bright Data’s $8 → $2.50/GB residential ladder, which is why included-volume and overage steepness, not headline rate, decide the deal.
Worked example — vector storage with a floor. Pinecone bills per read unit, per write unit, and per GB of storage: on the Standard plan (a $50/mo minimum) that’s roughly $16–18 per million read units, $4–4.50 per million write units, and ~$0.33/GB-month; Enterprise raises the minimum to $500/mo and the rates to ~$24–27/M RU. turbopuffer takes the purest version of this — writes, queries (now $1/PB after the February 2026 cut from $5/PB), and per-GB-month storage — but every tier carries a monthly minimum of $64, $256, or ≥$4,096. A workload billing $40 of metered usage on the $64 tier still pays $64; the floor is the real entry price. See usage invoicing and billing cycles for how minimums and overage are reconciled on the invoice.
Companies using this
The corpus tracks 33 data platforms spanning web extraction, search and knowledge APIs, GTM enrichment and analytics, vector and search databases, and the training-data and metering infrastructure underneath. The table below lists each company with its pricing model, billing units, free-tier status, and last-verified date.
Patterns observed
The meter follows the cost driver, and often changes per product. Bright Data is the clearest case: bandwidth-heavy residential proxies bill per-GB ($8 → $2.50), static ISP and datacenter proxies bill per-IP ($1.80 → $1.30 and $1.40 → $0.90), scraping and SERP APIs bill per 1,000 requests, and ready-made datasets bill per 1,000 records — four billing units under one roof, each tracking the dominant cost of its product line. Oxylabs mirrors this exactly with per-GB and per-IP proxies plus success-based per-1K-results scraper APIs (where 5xx/6xx system errors aren’t charged at all). The lesson is that “data platform pricing” is rarely a single model; it’s a portfolio of meters, one per product line.
Credits are the abstraction of choice for scraping and enrichment. Firecrawl (1 credit ≈ 1 page), ScraperAPI (a credit multiplier — 1 credit for a plain page, 10 for JS rendering or premium proxies, 75 for ultra-premium plus render), Diffbot ($0.001/credit on prepaid pools), and Clay (an Actions capacity tier plus a Data Credits pool) all use a credit layer. Credits let these vendors expose many features at different internal rates without forcing the buyer to learn each one, and let the vendor change underlying economics without renegotiating. Clay pushes this furthest: a single plan name like “Launch” spans $185/mo to $2,125+/mo depending purely on where the Data-Credits slider sits.
Free volume is table stakes, but only on the legible, cheap-to-serve meters. Diffbot gives 10,000 credits/month free forever, Firecrawl 1,000 credits, SerpApi a 250-search free plan, Qdrant a permanent 1GB free cluster, and Linkup a recurring $20 balance (~4,000 searches). The free tier sits on the request, credit, or small-cluster meter where marginal cost is low. Raw bandwidth and bulk datasets — where Bright Data charges from the first GB — and query-heavy infra rarely get a free allowance, because the cost of serving is real and immediate. turbopuffer has no free tier at all; its cheapest entry point is a $64/month usage minimum.
Subscription tiers are increasingly just pre-bought usage. ScraperAPI’s ladder ($49/mo for 100K credits up to $1,975/mo for 21.5M), Firecrawl’s six tiers, and Apify’s plans (Free, Starter $29, Scale $199, Business $999, each bundling a matching dollar amount of prepaid compute-unit usage) all present as flat monthly plans but are sized by an included usage allotment with metered overage beyond it. This hybrid framing — predictable floor, metered ceiling — is now the default for the category, blending the budgeting comfort of subscription with the cost-alignment of usage.
Meters are getting cheaper over time — the opposite of most software. Residential-proxy rates roughly halved from 2022 to 2026 across Oxylabs and Bright Data; Apify cut its Scale plan from $499 to $199 and dropped compute-unit rates ~20–25% in September 2025; Upstash replaced its old $280/mo Redis Pro plans with $10/mo Fixed plans (a ~28x drop); and turbopuffer cut queries 80%. The one clean counter-move is Bright Data’s dataset base rate, which rose from ~$1/1k to $2.50/1k records as the buyer mix shifted from scrapers to AI training labs — a demand signal, not an inflation one.
Counterexamples & variants
Vector databases price nothing like a scraper — and don’t even agree with each other. Pinecone, Weaviate, Qdrant, and turbopuffer all store and serve vectors, but their meters diverge sharply. Pinecone Serverless charges per read unit, write unit, and GB above a monthly minimum. Weaviate Cloud charges per million vector dimensions stored (from $0.00465/1M on Flex down to $0.002718/1M on Premium). Qdrant Cloud charges $0 for queries entirely — you pay for the cluster’s RAM and CPU, not the searches you run, which inverts Pinecone’s per-query logic. And turbopuffer meters bytes scanned per query ($1/PB) with a tier minimum up to ≥$4,096. Treating these as interchangeable would mislead on both the meter and the entry price: a bursty, low-query workload is cheapest on Pinecone Serverless, while a steady high-query index is often cheapest on a fixed Qdrant cluster.
The human-data marketplaces and labeling engines have little or no public meter. Scale AI, Mercor, micro1, and Snorkel AI sell data-engine work, human-labeling, and RL-environment partnerships to frontier AI labs, and they are almost entirely sales-quoted. Scale exposes only a thin self-serve Data Engine (free trial on the first 1,000 labeling units and 10,000 images, then pay-as-you-go via card, per-unit rates unpublished); the real business runs on committed annual contracts. Mercor discloses only the hourly expert pay, not the buyer take-rate. Labelbox is the exception that proves the rule: it normalizes usage into a Labelbox Unit (LBU) at $0.10/LBU on Starter, but its managed Alignerr human-data services remain sales-quoted. These vendors sit in the data-platform category because the product is data — but the pricing motion is enterprise sales-led, the opposite of the self-serve per-request norm.
Open-core inverts the frame entirely. OpenMeter is the metering-and-billing platform other usage-based vendors run on, and it’s priced open-core: a free Apache-2.0 self-hosted edition plus a contact-sales managed cloud. There is no per-event public rate — the free tier is the entire software, and revenue comes from the managed and enterprise upsell. The same pattern appears under the vector databases: Qdrant (Apache-2.0), Weaviate (BSD-3), Milvus, and Chroma all ship a fully free self-hostable engine and monetize only the managed cloud. In these variants the “data” being platformed is the customer’s own events or embeddings, and the pricing logic — free software, paid convenience — inverts the metered-from-the-first-unit norm of the scraping side of the category.
What this means for buyers vs vendors
For buyers
Forecast on your actual cost driver, not the headline tier. A scraping pipeline should model credits-per-page or successful-results-per-1K; a proxy workload should model GB-per-month and confirm which IP type it needs (residential vs datacenter is a ~10x price gap on Oxylabs); a vector store should model reads, writes, and stored GB or dimensions plus the tier minimum. The headline rates across Bright Data and Oxylabs are near-identical, so the deal is decided by included volume and overage steepness — get both in writing.
Watch for floors and hidden multipliers. Tier minimums on the infra products (turbopuffer, Pinecone) mean low-usage months still cost the floor price, so the minimum — not the per-unit rate — is your real entry cost. And on ScraperAPI, a “100,000-credit” plan can be anywhere from 1,333 to 100,000 actual scrapes because JS rendering and premium proxies cost 10–75 credits each, with concurrent-thread caps throttling throughput regardless of balance. Model the multiplier for your real target mix, not the plain-page rate.
Use the free allowances to benchmark real per-job cost on your own workload before committing — every legible meter in the category carries one. And remember the direction of travel: because these meters tend to fall, avoid long lock-ins on infra where Upstash-style repricing (a ~28x entry-price drop) or turbopuffer-style query cuts could strand you above the new rate card.
For vendors
Pick the meter that tracks your dominant marginal cost, and don’t be afraid to run multiple meters if you sell multiple products — Bright Data’s four-unit portfolio proves buyers tolerate this when each meter is intuitive for its product line. Use a credit abstraction when you need to expose many features at different internal rates or expect to change economics, as Firecrawl and Diffbot do — but keep the credit-to-output mapping legible (1 credit ≈ 1 page) or buyers will distrust it, and publish the multiplier table if you use one, the way ScraperAPI discloses its 1/10/75-credit tiers.
Put a free allowance on your cheapest-to-serve meter to drive self-serve acquisition, and reserve commitment minimums for infra products — vector, cache, and query stores — where buyers expect them. If you’re on the open-core path like Qdrant, Weaviate, or OpenMeter, make the free engine genuinely production-grade and monetize convenience, SLA, and support rather than gating core features.
Finally, plan for a falling meter. Residential proxies, compute units, and vector queries have all cheapened across the corpus, so design your rate card to survive a price cut — lead with included volume and differentiated support rather than a headline unit price you’ll have to walk down. Study OpenMeter if you’re building the metering itself; for the broader strategy, our usage-based pricing introduction and billing-cycle guide cover floor reconciliation and overage design.
| Company | Product | Pricing model | Billing units | Free tier | Verified |
|---|---|---|---|---|---|
| 6sense | ABM and B2B revenue-intelligence platform — predictive account scoring, buyer intent data, and AI sales/marketing workflows | No | 2026-07-14 | ||
| Anyscale | Managed Ray platform for distributed AI training, inference, and batch processing (RayTurbo, Anyscale Compute Units) | Yes | 2026-05-29 | ||
| Apify | Apify Platform — web scraping and browser-automation cloud with an Actors marketplace | Yes | 2026-06-03 | ||
| Bright Data | Web data platform — proxy networks, scraping APIs, a managed scraping browser, SERP and unlocker APIs, ready-made datasets, and eCommerce insights | Yes | 2026-07-14 | ||
| Browse AI | No-code web scraping and website-monitoring platform that turns any site into a structured dataset or API | Yes | 2026-06-04 | ||
| Chroma | Open-source vector database + Chroma Cloud | Yes | 2026-06-09 | ||
| Clay | AI-powered GTM data-enrichment and outbound platform billed on Actions plus Data Credits | Yes | 2026-07-06 | ||
| Comet | AI/ML observability and experiment-tracking platform — Opik (LLM/agent observability) and Comet MLOps (experiment tracking) | Yes | 2026-06-02 | ||
| Databricks (Mosaic AI) | Mosaic AI — enterprise GenAI & ML on the Data Intelligence Platform | Yes | 2026-06-15 | ||
| Diffbot | Web-extraction APIs (Extract, Crawl, Natural Language) plus a Knowledge Graph, metered on monthly credits | Yes | 2026-06-04 | ||
| Firecrawl | Web-scraping and data-extraction API for AI agents — scrape, crawl, map, search, and extract pages into clean markdown/JSON | Yes | 2026-06-30 | ||
| Julius AI | Julius AI — AI data-analyst chat & notebooks | Yes | 2026-06-08 | ||
| Labelbox | AI training-data platform (data labeling, curation & model evaluation) | Yes | 2026-06-15 | ||
| Linkup | Web search API for AI agents — Search, Fetch, and async Research endpoints with grounded, structured results | Yes | 2026-07-14 | ||
| LlamaIndex | RAG/agent orchestration framework + LlamaCloud document parsing | Yes | 2026-06-10 | ||
| Mercor | AI talent marketplace + enterprise data partnerships for frontier AI labs | No | 2026-07-14 | ||
| micro1 | Human-data engine, RL environments, and agent evaluation for frontier AI labs | No | 2026-07-14 | ||
| Milvus | Vector database (OSS) + Zilliz Cloud (managed) | Yes | 2026-06-09 | ||
| Nomic | Nomic Platform (AEC agentic workflows) + Atlas data-exploration app + Nomic Embed embedding/Developer API | Yes | 2026-06-04 | ||
| OpenMeter | Open-source usage metering and billing platform for AI, agentic, and developer tools | Yes | 2026-06-03 | ||
| Oxylabs | Web data collection: residential, datacenter, ISP & mobile proxies plus Web Scraper API and Web Unblocker | Yes | 2026-07-06 | ||
| Pinecone | Managed vector database (serverless) | Yes | 2026-06-09 | ||
| Powerdrill | AI-native data analytics platform that turns spreadsheets, PDFs, and databases into insights via specialized data agents | Yes | 2026-07-14 | ||
| Qdrant | Open-source vector database + Qdrant Cloud | Yes | 2026-06-09 | ||
| Rows | Rows AI spreadsheet | Yes | 2026-06-08 | ||
| Scale AI | Data engine, GenAI platform & contributor marketplace | No | 2026-06-15 | ||
| ScraperAPI | Web scraping API that handles proxies, browsers, and CAPTCHAs behind a single endpoint | No | 2026-06-04 | ||
| SerpApi | Real-time search-results API (Google, Bing, and other engines) | Yes | 2026-06-04 | ||
| Snorkel AI | Programmatic AI data development platform & expert data | No | 2026-06-15 | ||
| Snowflake Cortex | AI functions and model APIs on Snowflake | Yes | 2026-07-06 | ||
| turbopuffer | Serverless vector and full-text search database on object storage | No | 2026-07-14 | ||
| Unstructured | Document ingestion / ETL API | Yes | 2026-07-14 | ||
| Upstash | Upstash (Redis, Vector, QStash, Search, Workflow) | Yes | 2026-07-14 | ||
| Weaviate | AI-native vector database (open-source core + Weaviate Cloud managed serverless, dedicated/Enterprise Cloud, BYOC) | Yes | 2026-07-06 | ||
| ZoomInfo | GTM / sales-intelligence platform (contact + company data, intent, and the ZoomInfo Copilot AI GTM assistant) | No | 2026-07-06 |
Explore this theme in the knowledge graph
FAQ
What is data platform pricing?
Data platform pricing is the set of usage-based models used by scraping, enrichment, search-API, knowledge-graph, and vector-database vendors. Companies typically meter per request, per GB of bandwidth, per record, per credit, or per stored dimension, with tiered volume discounts and a free or trial allowance.
Why do data platforms charge per request or per GB instead of per seat?
The cost driver for a data platform is the volume of data extracted, stored, or delivered — not the number of users. A single account can run millions of automated requests, so per-request, per-GB, and per-record meters align the bill with the vendor's actual compute, proxy, and bandwidth cost.
Which billing unit is most common for web scraping APIs?
Per-request or per-credit billing dominates scraping APIs. ScraperAPI sells monthly API-credit pools with a credit multiplier (1 credit for a plain page, up to 75 for ultra-premium plus rendering), Firecrawl sells credit pools where 1 credit ≈ 1 page, and Bright Data and Oxylabs bill scraping endpoints per 1,000 successful results and proxies per GB or per IP.
Do data platforms offer free tiers?
Most do. Diffbot includes 10,000 credits per month free forever, Firecrawl gives 1,000 credits, SerpApi offers a 250-search free plan, Qdrant runs a permanent free cluster, and Linkup refills a $20 balance every month — though raw proxy bandwidth and bulk datasets are usually paid from the first GB.
How do vector databases price differently from web scrapers?
Vector databases meter storage and query volume rather than extraction volume. Pinecone bills per read unit, write unit, and GB of storage above a monthly minimum; Weaviate bills per million vector dimensions stored (from $0.00465/1M); and turbopuffer bills per write, per query ($1/PB), and per GB-month with a tier-scaled minimum. Scrapers meter requests, credits, or bandwidth instead.
Are AI data-labeling platforms priced like data APIs?
Only the self-serve slice is. Labelbox meters a normalized Labelbox Unit (LBU) at $0.10/LBU on its Starter tier, while Scale AI and Mercor sell human-data and RL-environment work almost entirely through enterprise sales-quoted contracts with no public per-unit rate.
Related product categories
- AI Coding Product PricingPricing for products whose primary surface is AI-assisted coding — IDEs, completion engines, and review agents.
- Developer Tools PricingPricing models used by tools sold to developers — IDEs, CLIs, libraries, voice-to-code, and adjacent products.
- AI Platform PricingPricing for general-purpose AI platforms — model APIs, inference services, and multi-model hosting providers.
- AI Infrastructure & Cloud PricingPricing for AI compute infrastructure — GPU clouds, serverless inference, and training platforms.
- Vertical SaaS PricingPricing for vertical SaaS products — AI software purpose-built for a specific industry (legal, healthcare, sales, marketing).
- PaaS PricingPricing for platform-as-a-service products that abstract away the underlying infrastructure and bill for higher-level units.
- Customer Service Platform PricingPricing for customer service software platforms — ticketing, chat, automation, and AI agent products.
- Horizontal SaaS PricingPricing for horizontal AI SaaS — productivity and workflow products sold across industries rather than to one vertical.
- Observability Platform PricingPricing for LLM and ML observability platforms — tracing, evaluation, and monitoring of model behavior in production.
- Fintech AI PricingPricing for AI-era fintech products — billing infrastructure, accounting automation, and financial operations platforms.
- Security AI PricingPricing for AI-powered security products — covering code security, voice fraud detection, SOC automation, and threat analysis.
- LLM Observability PricingPricing for platforms purpose-built to observe, debug, and optimize LLM application behavior — logging prompts, responses, latency, and cost.