JavaScript Web Scraping Guide 2026
Reviewed September 2026. Confirm runtime and library versions against their official documentation, and collect only data you are authorized to use.
Choose the simplest fetch path
Use ordinary HTTP requests for server-rendered HTML or JSON. Add browser automation only when the required, permitted content depends on client-side execution. This keeps resource use, failure handling, and tests much simpler.
Before collecting, document the source, purpose, allowed fields, rate and concurrency limits, retention, and deletion path. Review terms, access controls, privacy, copyright, and the Robots Exclusion Protocol. robots.txt is not authorization by itself.
HTTP collection with bounded concurrency
const controller = new AbortController()
const timeout = setTimeout(() => controller.abort(), 15_000)
try {
const response = await fetch(url, {
signal: controller.signal,
headers: { 'user-agent': 'ExampleResearchBot/1.0 [email protected]' }
})
if (!response.ok) throw new Error('HTTP ' + response.status)
const html = await response.text()
} finally {
clearTimeout(timeout)
}
Use a queue with a per-origin concurrency cap. Retry only transient network failures, 429 responses with an appropriate delay, and selected 5xx responses. Bound attempts and add jitter. Do not retry authentication failures or explicit access denials through new identities.
Parse and validate separately
Keep selectors and transformations outside the fetcher. Parse saved fixtures in unit tests and validate the result against a schema before storage.
import * as cheerio from 'cheerio'
function parseProduct(html, sourceUrl) {
const $ = cheerio.load(html)
return {
sourceUrl,
name: $('h1').first().text().trim() || null,
priceText: $('[data-price]').first().text().trim() || null
}
}
Preserve the raw observation or a content hash, source URL, collection time, parser version, and validation result. Quarantine a record with missing identifiers or ambiguous values instead of turning an empty selector into valid business data.
When a browser is justified
Playwright or Puppeteer can render approved JavaScript applications, interact with controls, and capture traces. Block unnecessary resources only when doing so cannot change the data being measured. Use fresh contexts for isolation and close pages and browsers in a finally block.
import { chromium } from 'playwright'
const browser = await chromium.launch()
try {
const context = await browser.newContext()
const page = await context.newPage()
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 30_000 })
const heading = await page.getByRole('heading').first().textContent()
} finally {
await browser.close()
}
Pin the browser package and review its official release notes. Do not add stealth patches, CAPTCHA solving, or fingerprint spoofing to overcome a site's controls.
Observability
- Queue age and jobs completed.
- Status and application error classes per origin.
- Latency percentiles, timeouts, retries, and bytes.
- Parse coverage, schema failures, duplicate rate, and freshness.
- Parser version and structural-drift alerts.
Keep secrets, authorization headers, cookies, and personal payloads out of logs. Use correlation identifiers and short retention.
Production checklist
- Pin Node.js and dependency versions; consult the official Node.js API documentation.
- Use explicit timeouts, bounded queues, and graceful shutdown.
- Test fetch and parser behavior independently with deterministic fixtures.
- Make jobs idempotent and storage writes safe to replay.
- Stop or seek permission when a destination rejects the workload.
This architecture makes a scraper easier to update because a network change, page-layout change, or schema change has a clear boundary and a measurable failure mode.