Chatonio

This is taking longer than usual.

Chatonio

How to create a Webscraper integration (Chromium)

April 25, 2026 62 viewsIntegrations

A webscraper tool launches a headless Chromium browser via Playwright, lets the page’s JavaScript run, and then hands the rendered HTML to the AI. Use this when the data only exists in the DOM after scripts have executed — React/Vue/Svelte apps, infinite-scroll pages, or anything that fetches its content client-side.

This is the slow path. Each call takes 2–8 seconds, costs more compute, and runs in an isolated browser environment that handles only a few at a time — so Chromium’s memory appetite never starves the rest of your project. Use it only when the parser genuinely won’t work.

When to use the scraper instead of the parser

  1. Hit View Source on the page in your browser.
  2. Search the source for the data you actually need.
  3. Found it? → use Webparser, it’s ten times faster.
  4. Empty page or just <div id="root"></div>? → you need the scraper.

Step 1 — Add a scrape tool with JS rendering

Same flow as the webparser — Admin → Projects → [project] → Integrations → Add tool — with two changes:

  • Type — Web Page Scrape.
  • Render JavaScript — turn it on. This is the only difference from a parser tool.
  • Allowed domains — still required. Strict allowlist: list exact hostnames, or add a wildcard entry like *.example.com — it covers example.com itself plus every subdomain.

The How to call it step for a scraper, with Render JavaScript ticked

Step 2 — CSS selector strategy

With a JavaScript app the page looks empty for a moment after load — then the framework paints the real content. Pass a CSS selector for the element you actually need; Playwright will wait for that selector to appear before grabbing the HTML. This is the difference between getting the data and getting the loading skeleton.

[data-testid="invoice-total"], .order-history-list

If your target loads progressively (lazy-loaded list, etc.), pick a selector that only exists once the data is in.

The Exclude field works here exactly as it does for the parser: one selector per line, removed from the page before the main selector runs. JavaScript apps tend to be generous with chrome — sticky headers, cookie bars, chat widgets, “you may also like” rails — so it usually earns its keep faster here than on a server-rendered page. See How to create a Webparser integration for the details.

Step 3 — Timeouts

The hard ceiling per call is 120 seconds, and slow scrapes start being wound down at around 90 — if your page genuinely takes longer than that, model the data behind a real API instead.

Custom HTTP headers are not configurable from the dashboard. A tool has no header field in the UI, so a header- or cookie-based session cannot be handed to the browser this way. If a page needs a particular header, contact support — or use a real API integration, where auth is configurable.

Step 4 — Test scrape

Use the same Test button as the parser — the result is identical, including the count of matched and excluded elements. The first call may be slower than steady-state because a fresh browser session has to spin up.

A scraper test result taking four seconds, against the parser’s fraction of one

Production limits and gotchas

  • Concurrency is small and fixed. Scrape calls are handled by a dedicated worker that keeps only a few browser sessions alive at once, and those processes are recycled regularly to keep memory in check. Heavy traffic queues rather than scaling out.
  • Output cap is still 16 KB. Use a tight CSS selector to keep the AI’s context lean.
  • SSRF protection is on. Loopback, RFC 1918 ranges, and link-local are rejected before the browser launches.
  • Anti-bot pages will fight back. Cloudflare Turnstile, hCaptcha, and similar will render their challenge in place of your data — you can’t scrape past them. Reach out to the site owner for an API instead.
  • Don’t scrape your own internal apps just because it’s easier. A real API integration with bearer auth is faster, cheaper, more reliable, and survives UI redesigns.
  • Requests come from a fixed IP — 64.7.198.218. If your WAF challenges unknown addresses, allowlist it: the headless browser cannot solve a challenge any more than it can solve a captcha.

Quick decision table

  • Server-rendered HTML → Webparser.
  • JS app, data appears after load → Webscraper.
  • Login-gated, no public data → API integration.
  • Static FAQ knowledge → KB article, no integration needed.

Was this article helpful?