Directory · Scraping stack

Web Scraping Tools & Scraping Proxies

Fourteen reviewed tools — crawling frameworks, headless browsers, unlocker APIs, managed platforms and AI scrapers — with real 2026 pricing, measured behaviour, and the proxy type each one needs behind it.

How to choose proxies for web scraping

The tool you pick decides how data comes out; the proxies decide whether it comes out at all. Almost every failed scraping project we review has the same shape — a capable crawler pointed at a hostile target through the wrong exit IPs. Start on datacenter proxies for scraping, promote a domain to rotating residential proxies when its block rate climbs, and escalate only the stubborn remainder to an unlocker API.

Below, each tool is graded on ease, scale ceiling, anti-bot resilience, documentation and value, with a full review covering pricing, benchmarks, proxy wiring code and alternatives. Together they cover every mainstream approach to building a scraping proxy network: self-hosted crawlers, headless browsers, managed scraping APIs and AI-native extraction.

Tools reviewed

15

Cheapest viable stack

$0 + $1/GB

Unlocker cost range

$0.20 – $3 / 1k

Typical success lift

+18 pts on residential

Crawling frameworks

Self-hosted crawling engines. Cheapest cost per page, and you own the proxy contract, the retries and the monitoring.

Scrapy logo

Scrapy

9.1/10

Crawling framework · Zyte / open source

The mature asynchronous Python crawling framework — middlewares, pipelines, AutoThrottle and a proxy layer that scales to millions of pages.

Price from
Free
Platforms
Python 3.9+ · Linux · macOS · Windows
Best for
Large structured crawls on stable HTML
Proxies
HTTP, HTTPS and SOCKS5 via downloader middleware; rotating gateways work out of the box
  • + Twisted async engine with per-domain concurrency
  • + AutoThrottle, retry and robots middleware
  • + Item pipelines to any store
  • + Huge middleware ecosystem
Crawlee logo

Crawlee

9.0/10

Crawling framework · Apify

Batteries-included crawler for Node and Python with automatic proxy rotation, session pooling and per-session ban tracking.

Price from
Free
Platforms
Node.js · Python
Best for
Production crawlers built fast, with ban handling included
Proxies
First-class ProxyConfiguration with rotation, tiers and automatic retirement of banned sessions
  • + SessionPool with automatic ban detection
  • + Tiered proxy configuration (cheap first, expensive on retry)
  • + Unified request queue across HTTP and browser crawlers
  • + Adaptive crawler picks HTTP vs browser per page

Headless browsers

Real browser engines for JavaScript-rendered targets. Expensive per page — attach the proxy per context and block heavy assets.

Playwright logo

Playwright

9.3/10

Headless browser · Microsoft

Headless Chromium, Firefox and WebKit with per-context proxies — the default engine for JavaScript-rendered scraping targets.

Price from
Free
Platforms
Node · Python · .NET · Java
Best for
Dynamic, JS-rendered and login-gated pages
Proxies
Per-browser and per-context HTTP/SOCKS5 proxies with username auth
  • + Per-context proxy rotation
  • + Network interception and request blocking
  • + Auto-wait removes flaky sleeps
  • + Codegen and trace viewer for debugging
Puppeteer logo

Puppeteer

8.5/10

Headless browser · Chrome DevTools team

Chrome DevTools Protocol automation for Node, with the largest stealth-plugin ecosystem in scraping.

Price from
Free
Platforms
Node.js (Chrome & Firefox)
Best for
Chrome-only automation and stealth-plugin workflows
Proxies
Per-browser proxy via --proxy-server; per-request rotation via a proxy-chain bridge
  • + puppeteer-extra-plugin-stealth
  • + Raw CDP access for deep control
  • + proxy-chain for authenticated rotation
  • + Enormous community and snippet base
Selenium logo

Selenium

7.8/10

Headless browser · Selenium project (SeleniumHQ)

The WebDriver standard — every language binding, every browser, and Grid for horizontal scale.

Price from
Free
Platforms
Java · Python · C# · Ruby · JS
Best for
Legacy stacks, cross-browser and antidetect-browser control
Proxies
Per-driver HTTP/SOCKS5 proxy via Options or a Selenium Wire interceptor
  • + W3C WebDriver standard
  • + Grid for distributed sessions
  • + First-class antidetect browser support
  • + Selenium Wire for request-level proxying

Unlocker / scraping APIs

Managed endpoints that replace your proxy layer and bill per successful response. Route only the blocked tail through them.

Bright Data Web Unlocker logo

Bright Data Web Unlocker

8.9/10

Unlocker / scraping API · Bright Data

A single proxy endpoint that solves challenges, rotates fingerprints and retries until it returns clean HTML — billed per successful request.

Price from
≈ $1.50 / 1k successful requests
Platforms
Any HTTP client · Playwright/Puppeteer via Scraping Browser
Best for
Hardened anti-bot targets where success rate matters more than unit cost
Proxies
Is the proxy — 150M+ residential IPs behind an automated unlock layer
  • + Automatic CAPTCHA and challenge handling
  • + Fingerprint and header rotation managed for you
  • + Pay only for successful responses
  • + Country, state and city targeting
Zyte API logo

Zyte API

8.6/10

Unlocker / scraping API · Zyte (creators of Scrapy)

Rendering, ban handling and automatic extraction as one API from the team that maintains Scrapy.

Price from
from ≈ $0.20 / 1k requests
Platforms
HTTP API · scrapy-zyte-api · Python/Node SDKs
Best for
Scrapy-native stacks that need ban handling without changing framework
Proxies
Managed proxy selection with automatic escalation from datacenter to residential
  • + Automatic ban detection and proxy escalation
  • + Browser rendering toggled per request
  • + AI extraction for products, articles and job posts
  • + Native scrapy-zyte-api integration
ScraperAPI logo

ScraperAPI

8.2/10

Unlocker / scraping API · ScraperAPI

Simple per-credit scraping API with proxy rotation, rendering and structured endpoints for Amazon, Google and Walmart.

Price from
$49 / mo (100k credits)
Platforms
HTTP API · proxy port · SDKs
Best for
Small and mid-size teams that want one endpoint and no infrastructure
Proxies
40M+ IP pool with automatic rotation; premium residential behind a credit multiplier
  • + Proxy-port mode — no code change needed
  • + Structured endpoints for Amazon, Google, Walmart
  • + Async batch job endpoint
  • + Geotargeting on paid plans
Oxylabs Web Scraper API logo

Oxylabs Web Scraper API

8.8/10

Unlocker / scraping API · Oxylabs

Enterprise scraper APIs with AI parsing over a 175M+ IP pool, returning structured JSON for SERP, e-commerce and real estate.

Price from
from $49 / mo
Platforms
HTTP API (realtime, push-pull, proxy endpoint)
Best for
Structured commercial data at enterprise volume
Proxies
Managed 175M+ residential and datacenter pool with country/city targeting
  • + AI-powered adaptive parser
  • + Dedicated parsers for Amazon, Google, Walmart, Best Buy
  • + Push-pull mode for very large batches
  • + Localised results down to city level

Managed platforms

Hosted infrastructure with compute, proxies and storage on one invoice. Fastest route to a maintained data feed.

Apify logo

Apify

8.7/10

Managed platform · Apify Technologies

Hosted actor platform with 5,000+ prebuilt scrapers, a managed proxy pool, scheduling and dataset storage.

Price from
$0 (free tier) · from $49/mo
Platforms
Cloud · Node/Python SDK · REST API
Best for
Buy-don't-build scraping and scheduled data feeds
Proxies
Built-in datacenter and residential pools plus bring-your-own proxy URLs
  • + 5,000+ ready actors in the store
  • + Managed residential and datacenter proxies
  • + Cron scheduling and webhooks
  • + Datasets, key-value stores and API export
Browserbase & Stagehand logo

Browserbase & Stagehand

8.3/10

Managed platform · Browserbase

Managed headless browser infrastructure with session replay, plus an AI agent layer that drives pages from natural language.

Price from
free tier · from $39 / mo
Platforms
Cloud · Playwright/Puppeteer CDP · Stagehand SDK
Best for
AI browser agents and teams that refuse to run a browser fleet
Proxies
Built-in residential proxies per session, or bring your own
  • + Connect existing Playwright code over CDP
  • + Live view and session replay for debugging
  • + Stagehand act/extract/observe primitives
  • + Captcha solving and stealth mode built in

AI scrapers

LLM-native extraction that survives redesigns because you describe the data instead of where it sits in the DOM.

Firecrawl logo

Firecrawl

8.7/10

AI scraper · Firecrawl (Mendable)

Turns any site into clean, LLM-ready markdown with crawl, scrape, map and extract endpoints.

Price from
free tier · from $16 / mo
Platforms
HTTP API · Python/Node SDK · LangChain & LlamaIndex integrations
Best for
RAG ingestion and LLM pipelines
Proxies
Managed proxies with a stealth mode tier for protected pages
  • + Markdown output tuned for LLM context windows
  • + /map returns every URL on a domain in seconds
  • + Schema-based /extract with JSON output
  • + Self-hostable open-source core
Crawl4AI logo

Crawl4AI

8.4/10

AI scraper · Open source (unclecode)

Fast asynchronous open-source crawler built to feed retrieval pipelines with chunked, cleaned, LLM-ready content.

Price from
Free
Platforms
Python 3.10+ · Docker
Best for
Self-hosted RAG ingestion with your own proxies
Proxies
Per-run proxy config plus a rotating proxy strategy, HTTP and SOCKS5
  • + Fit-markdown filtering removes boilerplate
  • + Built-in chunking strategies for embeddings
  • + LLM and CSS extraction strategies side by side
  • + Docker API server with a job queue
ScrapeGraphAI logo

ScrapeGraphAI

8.0/10

AI scraper · ScrapeGraphAI

Prompt-defined extraction — describe the data you want and an LLM builds the scraping graph instead of you writing selectors.

Price from
Free (OSS) · API from $20 / mo
Platforms
Python · Node SDK · HTTP API
Best for
Long-tail sites where maintaining selectors is not worth it
Proxies
Proxy settings per graph config; works with rotating residential gateways
  • + SmartScraperGraph from a plain-language prompt
  • + Works with OpenAI, Anthropic, Gemini or local Ollama
  • + Search-and-scrape graph combines SERP with extraction
  • + Schema output via Pydantic models
Diffbot logo

Diffbot

8.0/10

AI scraper · Diffbot

Computer-vision page extraction plus a knowledge graph of billions of entities — no rules, no selectors, no maintenance.

Price from
from $299 / mo
Platforms
HTTP API · Crawlbot · KG query language
Best for
Enterprise entity data and zero-maintenance extraction
Proxies
Fully managed — no proxy configuration exposed
  • + Automatic page-type classification
  • + Article, product, discussion and image APIs
  • + Knowledge Graph with billions of entities
  • + Crawlbot for whole-domain jobs

Head-to-head comparisons

Choosing between two tools? Each matchup compares throughput, anti-bot success, published pricing and the proxy tier both tools expect, then states a decision rule.

See all scraping tool comparisons

Scraping proxies comparison — which type for which crawl

WorkloadProxy typeTypical priceTools that fit
Large static HTML crawlsDatacenter$0.30 – $2 / IP / moScrapy, Crawlee
Consumer sites & marketplacesRotating residential$1 – $8 / GBCrawlee, Playwright, Apify
Logged-in dashboardsISP / static residential$1.50 – $6 / IP / moPlaywright, Selenium
App-only endpointsMobile 4G/5G$4 – $20 / GBPlaywright, Browserbase
Hardened anti-bot targetsUnlocker API$0.20 – $3 / 1k requestsWeb Unlocker, Zyte API, ScraperAPI
LLM & RAG ingestionRotating residential$1 – $8 / GBFirecrawl, Crawl4AI, ScrapeGraphAI

Scraping proxy FAQs

What are the best proxies for web scraping in 2026?+

Rotating residential proxies for consumer-facing sites, datacenter proxies for tolerant or internal targets, ISP proxies for logged-in sessions and mobile proxies for app-only endpoints. Most production stacks run a tiered scraping proxy network that starts cheap and escalates per domain.

Do I need residential proxies for scraping, or are datacenter proxies enough?+

Always test datacenter first — they cost a fraction as much and add far less latency. Move a domain to residential proxies for scraping only when your measured block rate on that domain passes a few percent.

Which scraping tool should I choose?+

Scrapy or Crawlee for large structured crawls, Playwright for JavaScript-rendered pages, an unlocker API such as Bright Data Web Unlocker or Zyte API for hardened anti-bot targets, and Firecrawl or Crawl4AI when the output feeds an LLM.

How many rotating proxies do I need for scraping?+

With a rotating gateway you buy concurrency and bandwidth, not IP counts. Divide requests per hour by the target's per-IP rate limit to size concurrency, then add a 20–60% retry factor.

What is a scraping proxy network?+

The routing layer between your crawler and the web: a pool of exit IPs, session management, rotation rules and ban detection. Tool choice determines how you attach it; pool quality determines your success rate.

Is web scraping legal?+

Collecting publicly accessible data is broadly accepted in most jurisdictions, but terms of service, rate limits, authentication walls and data-protection law still apply. Take legal advice for commercial programmes.

Related free proxy tools

All tools →

Trusted partners