Web Scraping at Scale
Reviewed September 2026. This page is the maintained owner for reliable, ethical collection at scale and now includes the useful reliability guidance from the former Web Scraping Best Practices article.
Scale the control plane before volume
Put approved work in a durable queue and let stateless workers claim idempotent jobs. Enforce per-origin concurrency and rate limits centrally, set finite connect and response timeouts, retry only transient failures with jitter, and move exhausted jobs to a reviewable dead-letter queue.
Keep fetching separate from parsing and storage. A worker should be safe to stop and replay without duplicating records or losing provenance.
Design the queue and scheduler
A scheduler decides what is due; a durable queue holds work; workers claim, execute, and acknowledge jobs. Give a job a stable ID, source URL, collection policy, parser version, attempt count, deadline, and idempotency key. Lease jobs for a finite period so a crashed worker can be recovered. Make the storage write idempotent because a worker can complete after its lease expires.
{
"jobId": "source:product:123:2026-09-21",
"origin": "https://source.example",
"parserVersion": "product-v4",
"attempt": 0,
"maxAttempts": 3,
"deadline": "2026-09-21T23:00:00.000Z"
}
Apply a token bucket or equivalent rate limit keyed by origin before jobs reach workers. Keep a global limit as well, so an autoscaling event cannot multiply traffic. A scheduler should delay work when the origin budget is exhausted instead of creating a backlog of immediately failing requests.
Source and ethics review
- Document the business purpose, permitted sources, required fields, and retention.
- Review access controls, terms, privacy, copyright, and the Robots Exclusion Protocol.
- Identify the collector where appropriate and provide a contact path.
- Treat throttling or denial as a signal to slow down, stop, or seek permission.
Reliable job design
Give every job a stable identifier, source URL, attempt count, parser version, and deadline. Use a uniqueness constraint or idempotency key at the write boundary. A retry must not create a second record or repeat a side effect.
Split queues or limits by origin so a slow or failing site cannot consume every worker. Apply backpressure when storage or downstream processing falls behind.
Retry as a bounded state machine
Classify connection timeouts, resets, selected server errors, malformed documents, validation failures, and access-policy responses separately. Retry only failures that the source agreement and operational policy regard as transient. Use a finite retry budget and exponential delay with jitter. Authentication failures, an explicit denial, or a challenge that requires human permission should stop the job rather than trigger more traffic.
Send exhausted jobs to a dead-letter queue that retains only the metadata needed to diagnose them. Provide a reviewed replay action, not an automatic infinite loop. Track attempts and error class at the job level so an incident can distinguish a provider outage from a parser release.
Data quality
Define required fields, types, units, identifiers, freshness, and acceptable missingness before collection. Preserve raw observations or hashes, validate raw and normalized records separately, and quarantine failures rather than silently coercing them.
- Test parsers against saved representative fixtures.
- Alert on structural drift, missing identifiers, unexpected nulls, and duplicate spikes.
- Version parsing and normalization rules.
- Sample records for human review and record the result.
HTTP before browsers
Use ordinary HTTP collection for static HTML or approved APIs. Browser workers require more CPU, memory, browser-version maintenance, and debugging, so reserve them for content that genuinely depends on client execution. Keep browser and HTTP work in separate pools with separate limits.
Observability
- Queue age, scheduled versus completed jobs, and dead-letter volume.
- Connection, TLS, HTTP, and application error classes by origin.
- Latency percentiles, timeouts, retries, bytes, and usable-result cost.
- Parse coverage, schema failures, duplicate rate, and freshness.
- Worker saturation, memory, restarts, and deployment version.
Keep credentials, cookies, authorization headers, and sensitive payloads out of logs. Use correlation identifiers and short retention.
Use traces to answer operational questions
Attach a correlation ID to the scheduler event, queue message, worker attempt, normalized record, and log entry. A useful trace can answer: why was this URL scheduled, which worker version processed it, what exit or route was used where relevant, which parser produced the record, and why did a retry occur? Keep raw content and browser artifacts under a separate approved retention policy because they may contain personal data.
| Signal | Question it answers | Typical action |
|---|---|---|
| Queue age | Can the system meet freshness targets? | Reduce intake or add capacity within global limits |
| Origin error rate | Is a source unavailable or rejecting work? | Back off, pause, or seek permission |
| Schema failure rate | Did content or parsing change? | Quarantine and update a fixture |
| Duplicate rate | Are jobs or entity rules repeating output? | Inspect idempotency and entity keys |
Capacity planning
Load-test with a controlled server or saved response set, then choose worker counts from measured latency, resource use, destination limits, and downstream capacity. Autoscaling should respect a global ceiling and per-origin limits; adding workers must not multiply traffic beyond approved bounds.
Build a capacity model from measurements
Estimate the arrival rate, median and upper-percentile service time, worker concurrency, response size, browser share, storage throughput, and permitted per-origin rate. Test HTTP and browser pools separately. A system that can fetch quickly may still be limited by parsing CPU, database writes, or a freshness window. Increase one constraint at a time on a controlled server or fixture set and document the limit found.
Incident and change handling
Pause a source when error rates, schema failures, or denial signals cross a reviewed threshold. Preserve enough evidence to diagnose the failure, update a fixture, patch and test the parser, then replay only idempotent jobs. Every source should have an owner and a stop switch.