Web Scraping Data Quality Guide
Reviewed September 2026. Quality targets should be defined and measured for each dataset rather than borrowed from generic benchmarks.
Write a data contract
Define required fields, types, identifiers, units, allowed values, freshness, uniqueness, and acceptable missingness before collection. Include the source URL, collection time, parser version, and validation result so every record has provenance.
Separate the stages
Keep raw capture, parsing, normalization, validation, and publication distinct. Preserve raw observations or integrity hashes under an approved retention policy. A layout change should fail validation or enter quarantine, not silently become valid business data.
Quality dimensions
- Completeness: required records and fields are present.
- Accuracy: values match a verified source sample.
- Consistency: equivalent values use the same representation.
- Conformity: types, formats, units, and ranges follow the contract.
- Uniqueness: duplicate entities and observations are controlled.
- Timeliness: data arrives within the documented freshness target.
- Integrity: relationships and identifiers remain valid.
Validation controls
- Reject or quarantine missing identifiers and impossible values.
- Normalize currency, time zones, units, and Unicode explicitly.
- Use stable entity keys and document fuzzy-matching thresholds.
- Compare counts and distributions with a reviewed baseline.
- Sample records for human verification after every material parser change.
Test for structural drift
Save representative source fixtures and expected parsed records. Test selectors, transformations, edge cases, and empty states. Alert on coverage loss, unexpected nulls, new categories, and abrupt distribution changes. Version the parser so downstream users can explain a change.
Monitor production quality
Track source availability, collection delay, parse coverage, schema failures, missingness, duplicate rate, quarantine volume, manual-review results, and freshness. Break these down by source and parser version; a global average can hide one broken feed.
Handle corrections
Keep a clear path to correct or replay affected observations. Determine which parser version and time range produced bad data, stop publication, patch against a fixture, revalidate, and replay idempotently. Tell downstream users when a correction changes material results.
Governance
Collect only approved data, apply privacy and retention rules, restrict access, and delete data when its purpose expires. Review terms, copyright, robots guidance, and access controls for each source. Quality does not make an unauthorized dataset acceptable.
For the surrounding collection architecture, see web scraping at scale. Reliable data comes from an explicit contract, observable validation, and a tested correction process.