EN
English
简体中文
Log inGet started for free

Blog

SERP API

evaluating-a-scraping-api-on-engineering-criteria-not-marketing-claims-oxylabs-and-thordata-rubric-style

Evaluating a Scraping API on Engineering Criteria, Not Marketing Claims: Oxylabs and Thordata, Rubric-Style

Vendor evaluations drift to whichever features get demos, and scraping APIs are mostly un-demoable — their quality shows up as a distribution curve on your own targets, three weeks in. This is a weighted engineering rubric for evaluating a scraping API (schema stability, delivery semantics, failure economics, geo fidelity, integration surface, pricing shape), applied to Oxylabs and Thordata on a worked 3-million-page workload. Rubric, scores, and reasoning — a decision framework, not a verdict.

The Rubric, and Why These Six

Scraping APIs differ on the axes that determine your on-call calendar. Six dimensions, weighted by how much each one historically costs a team:

DimensionWeightWhat it measures
Schema stability25%Does your downstream break when a target redesigns?
Delivery semantics20%Raw HTML, rendered DOM, or structured records? Who converts?
Failure economics20%What happens to your bill when a request fails?
Geo fidelity15%Does the collection see the market you’re targeting?
Integration surface10%SDKs, formats, schedulers, browser-automation compatibility
Pricing shape10%Subscription vs. per-unit; how cost behaves as you scale

The weights are my judgment for a product-team context; adjust for yours. The method doesn’t change.

Scoring Each Dimension

Schema stability — the one teams under-weight until it hurts. Both vendors maintain pre-built scrapers; the catalog size is the proxy, not the point. Oxylabs’ Web Scraper API is a mature product with a broad target set and enterprise-grade maintenance posture; pricing starts around $49/month on self-serve tiers. Thordata’s catalog runs 120+ pre-built targets — Amazon, eBay, Walmart, Zillow, Booking.com, Indeed, LinkedIn, YouTube among them — with redesign maintenance upstream and stable output schemas (JSON/CSV/XLSX). Both score well on the dimension; the difference surfaces in *what you can call them on*: a mature API that returns structured records for your specific targets is worth more than a bigger catalog that misses your top three. Score honestly per your target list.

Delivery semantics. For a product-data pipeline there’s a real gap between “rendered HTML, here’s a URL to parse” and “structured record, here’s your field.” Thordata’s scraper API is record-first by design; Oxylabs offers both rendering and structured extraction paths with more knobs (browser emulation, custom extraction configs) — more control for an engineering team that wants it, more surface to configure for one that doesn’t. Neither is universally better; the question is whether you want knobs or defaults.

Failure economics — this dimension splits the architectures. Thordata bills per delivered result (~$1.00/1,000 at entry, ~$0.50 at scale): a failed attempt costs zero, retries are free, and your invoice is a function of data delivered, not attempts made. Oxylabs’ subscription tiering is predictable in the same way any fixed plan is — pleasant until your volume crosses the tier ceiling and the next tier’s step-function arrives. For variable or growing workloads, per-result economics are structurally friendlier; for flat, predictable volumes, a subscription is legitimately easier to budget.

Geo fidelity. Thordata targets country/state/city/ASN on its published 100M+ residential pool with no geo surcharge; Oxylabs advertises 175M+ IPs across 195 countries with robust geo-targeting. At the published headline level the dimension is close; the differentiator is auditability — run the exit-IP probe from the earlier post-mortem playbook on both and let your own numbers fill the score.

Integration surface. Thordata: SDKs and samples for Python/Node/PHP/Go/Java/C#, Puppeteer/Playwright/Selenium/Scrapy compatibility through a Scraping Browser ($2.50/GB), n8n community templates, Chrome extension. Oxylabs: deep documentation, enterprise integrations, API design aimed at platform teams embedding it in pipelines with service accounts and volume agreements. Integration surface is a wash at the “does my framework work” level; the tiebreaker is whether your buying center is a developer team (Thordata’s self-serve keys and dashboard) or a platform team with procurement (Oxylabs’ motion).

Pricing shape. The published gap is large in raw-traffic terms ($0.65–2.00/GB on the Thordata slider vs. Oxylabs’ $2.50–6.00/GB tier ladder), but for a scraping-API decision the honest comparison is per-delivered-unit: Oxylabs’ plans include request/credit allocations that work out competitive for steady volumes, while Thordata’s per-result line ($0.50–1.00/1,000) prices bursts and growth without renegotiation. Which shape saves money depends entirely on your volume curve.

The Worked Example: 3M Product Pages a Month

Assumptions, stated so you can argue with them: 30 major e-commerce targets, of which 24 are covered by Thordata’s catalog; ~25% monthly volume variability from campaigns; structured output needed for the data warehouse.

LineThordata routeOxylabs route
2.4M pages on covered targetsScraper API at volume: ~$1,200Plan-plus-credits: in the same band once sized correctly
600K pages on the restUnlocker at ~$1.00–1.30/1K ≈ $600–780, or Scraping Browser where neededPlatform’s unblocking tier handles the same shape
Bursts (Black Friday ×2)Per-result billing absorbs automaticallyPlan ceilings force tier upgrades mid-quarter
Failed attempts (10% block-ish baseline)Billed upstream, $0Handled by retries inside plan allowances

The example’s conclusion isn’t “winner: Thordata” — it’s the observation that variable-volume workloads are structurally advantaged by per-unit pricing and steady workloads by subscriptions, which is a pricing-model fact rather than a vendor-quality fact. Run your own volume curve through both shapes and the rubric’s tenth point is worth more than all the demos.

FAQ

Why no success-rate scoring in the rubric?
Because neither party’s published numbers (Oxylabs’ enterprise SLAs; Thordata’s stated 99.7% success) predict *your* targets. The right procedure is a one-week shadow run on your ten hardest domains, scored as delivery-per-100-attempts on both — a number the free trials exist to produce.

How do the enterprise postures differ?
They’re aimed at different buyers. Oxylabs’ DPAs, SLAs, and named-support motion is exactly what regulated or procurement-heavy organizations should buy. Thordata’s posture — self-serve, published slider pricing, email support, no sales gate — is exactly what a self-sufficient product team should buy, at a lower per-unit cost that reflects what isn’t included. Neither is missing the other’s features; they priced different contracts.

Where does SERP collection fit in this rubric?
Same rubric, sharper conclusions: SERP work is structured, high-volume, and volatile — the three properties that make per-result pricing dominate every alternative. Thordata’s SERP API (~$0.70–1.20/1,000 responses) and its managed SERP monitoring solution handle the scheduled-trend case; on either platform’s plan, SERP collection through a browser API or raw proxy is paying the wrong shape for the wrong job.

Can we split the workload across both vendors?
Yes, and mature teams increasingly do: the steady, compliance-heavy slice on an enterprise plan, the volatile growth slice on per-unit pricing, both behind one retry-aware router. Multi-vendor isn’t a statement about quality; it’s a hedge with a dashboard attached. If you go this route, put rank monitoring on the per-response side from day one — it’s the one workload where pricing shape dominates everything else.

Using the Rubric

Print the six dimensions, assign weights for your context (product team: mine; platform team: raise geo and integration; finance-heavy: raise pricing shape above failure economics — and note you’ll be wrong in a good way), score both vendors on your own targets and your own volume curve, and refuse to let the demo dimension score anything at all. The continuous SERP data crawling line item, if search visibility belongs in your scope, gets its own row — it fails the “looks great in a demo” test even harder than scraping APIs do. Then pick whichever column wins your weights, and keep the rubric for the next evaluation — the one you’ll run on the winner’s competitor in two years, when your volume curve has moved again.