🎉 We just launched long anticipated residential proxies & pay-per-GB!
Pay once, switch between multiple proxy providers.
50% discount for a limited time. See more →

JavaScript Web Scraping Guide 2026

published 2025-02-27
by Amanda Williams
9,338 views

Reviewed September 2026. Confirm runtime and library versions against their official documentation, and collect only data you are authorized to use.

Choose the simplest fetch path

Use ordinary HTTP requests for server-rendered HTML or JSON. Add browser automation only when the required, permitted content depends on client-side execution. This keeps resource use, failure handling, and tests much simpler.

Before collecting, document the source, purpose, allowed fields, rate and concurrency limits, retention, and deletion path. Review terms, access controls, privacy, copyright, and the Robots Exclusion Protocol. robots.txt is not authorization by itself.

HTTP collection with bounded concurrency

const controller = new AbortController()
const timeout = setTimeout(() => controller.abort(), 15_000)

try {
  const response = await fetch(url, {
    signal: controller.signal,
    headers: { 'user-agent': 'ExampleResearchBot/1.0 [email protected]' }
  })
  if (!response.ok) throw new Error('HTTP ' + response.status)
  const html = await response.text()
} finally {
  clearTimeout(timeout)
}

Use a queue with a per-origin concurrency cap. Retry only transient network failures, 429 responses with an appropriate delay, and selected 5xx responses. Bound attempts and add jitter. Do not retry authentication failures or explicit access denials through new identities.

Parse and validate separately

Keep selectors and transformations outside the fetcher. Parse saved fixtures in unit tests and validate the result against a schema before storage.

import * as cheerio from 'cheerio'

function parseProduct(html, sourceUrl) {
  const $ = cheerio.load(html)
  return {
    sourceUrl,
    name: $('h1').first().text().trim() || null,
    priceText: $('[data-price]').first().text().trim() || null
  }
}

Preserve the raw observation or a content hash, source URL, collection time, parser version, and validation result. Quarantine a record with missing identifiers or ambiguous values instead of turning an empty selector into valid business data.

When a browser is justified

Playwright or Puppeteer can render approved JavaScript applications, interact with controls, and capture traces. Block unnecessary resources only when doing so cannot change the data being measured. Use fresh contexts for isolation and close pages and browsers in a finally block.

import { chromium } from 'playwright'

const browser = await chromium.launch()
try {
  const context = await browser.newContext()
  const page = await context.newPage()
  await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 30_000 })
  const heading = await page.getByRole('heading').first().textContent()
} finally {
  await browser.close()
}

Pin the browser package and review its official release notes. Do not add stealth patches, CAPTCHA solving, or fingerprint spoofing to overcome a site's controls.

Observability

  • Queue age and jobs completed.
  • Status and application error classes per origin.
  • Latency percentiles, timeouts, retries, and bytes.
  • Parse coverage, schema failures, duplicate rate, and freshness.
  • Parser version and structural-drift alerts.

Keep secrets, authorization headers, cookies, and personal payloads out of logs. Use correlation identifiers and short retention.

Production checklist

  • Pin Node.js and dependency versions; consult the official Node.js API documentation.
  • Use explicit timeouts, bounded queues, and graceful shutdown.
  • Test fetch and parser behavior independently with deterministic fixtures.
  • Make jobs idempotent and storage writes safe to replay.
  • Stop or seek permission when a destination rejects the workload.

This architecture makes a scraper easier to update because a network change, page-layout change, or schema change has a clear boundary and a measurable failure mode.

Amanda Williams
Amanda is a content marketing professional at litport.net who helps our customers to find the best proxy solutions for their business goals. 10+ years of work with privacy tools and MS degree in Computer Science make her really unique part of our team.
Don't miss our other articles!
We post frequently about different topics around proxy servers. Mobile, datacenter, residential, manuals and tutorials, use cases, and many other interesting stuff.