Chatonio

This is taking longer than usual.

Chatonio

How to create a Webparser integration

April 25, 2026 61 viewsIntegrations

A webparser tool lets the AI fetch and read a web page on demand using a plain HTTP request — the same way curl would. It’s the right choice when the page you need is server-rendered and the data is already in the HTML that comes back. Fast (200–800 ms typical) and cheap — no headless browser has to start up.

Webparser vs Webscraper — which one?

  • Webparser (this article) — plain HTTP fetch. Use when the data is visible in View Source. Examples: documentation pages, status pages, simple pricing tables, blog posts.
  • Webscraper — full headless Chromium with JavaScript execution. Use when the page is a JavaScript app and the data only appears after scripts run. See How to create a Webscraper integration.

Try the parser first. If View Source on the target page shows the data you need, the parser is fine and ten times faster than the scraper.

Step 1 — Add a scrape-type tool

Webparser and webscraper share one tool type called scrape; the difference is the render_js flag. Open Admin → Projects → [project] → Integrations, pick (or create) any integration, then Add tool and set:

  • Tool type — Web Page Scrape.
  • Function name — fetch_status_page or similar.
  • Description — one sentence describing when to call it. The AI reads this to decide.
  • Render JavaScript — leave it off.
  • Allowed domains — comma-separated list of hostnames the tool may fetch from, e.g. status.example.com, docs.example.com. This is enforced — anything outside the list is rejected before a request goes out.

The What it does step with Tool type set to Web Page Scrape

Step 2 — Narrow down with a CSS selector (optional)

By default the tool returns the cleaned text of the whole page (capped at 16 KB). For long pages you can pass a CSS selector so only the matching element’s text is returned:

.status-summary, main article, #pricing-table

Anything CSS can match works, and several selectors separated by commas are matched independently — their text is joined, so .summary, .details returns both blocks. If the selector matches nothing, the tool does not fall back to the full page — the AI receives a “No content found matching selector” message instead. (One convenience retry exists: a bare word like order-status is automatically retried as the class selector .order-status.) Always use the Test button after setting a selector — its preview shows exactly what the AI will see, and reports how many elements matched, so a selector miss is immediately visible before going live.

Excluding parts of the page

A selector chooses which elements to read, not which parts to skip inside them. So if you select an article wrapper that happens to contain a nav bar, a cookie banner or a “related posts” block, all of that text comes along too. Put those in the Exclude field, one selector per line:

.site-nav
.cookie-banner
.related-posts

The How to call it step for a parser: allowed domains, CSS selector, Exclude, Render JavaScript off

Excluded elements are removed from the page before the main selector runs, so the block you selected keeps its structure and simply loses the noise. One per line rather than comma-separated, because a single CSS selector may itself contain commas.

You may be tempted to do this with :not() inside the main selector instead. It is supported, but it works on which elements get picked rather than which descendants survive inside them — so it flattens the block into fragments and makes nested elements repeat their text. The Exclude field is the one that does what you mean.

One thing to watch: if an exclude rule matches an ancestor of what you were selecting, it removes your target as well. Chatonio says so explicitly when that happens, instead of just reporting an empty result.

Step 3 — Test it

Use the Test button on the tool row. Paste a real URL within your allowed-domains list. Alongside the extracted text you get the number of elements the selector matched and how many the exclude rules removed, which is what makes tuning a selector a matter of trial rather than guesswork.

A parser test result showing how many elements the selector matched and the excludes removed

POST/api/projects/{project_id}/integrations/{integration_id}/tools/{tool_id}/test-scrape/

Returns up to 4,000 characters of preview so you can confirm what the AI will see. This is the older scrape-only route and still works; the tool Test button now uses the shared /tools/{tool_id}/test/ endpoint, which serves both scrape and API tools.

Request

Body
{
  "url": "https://status.example.com/incidents/latest"
}

Responses

200Success
{
  "success": true,
  "detail": "Fetched in 412 ms",
  "preview": "All systems operational..."
}
400Domain blocked
{
  "success": false,
  "detail": "Domain not in allowed list"
}

How it shows up in conversations

The AI calls the parser on its own when a customer’s question lines up with the tool’s description — you don’t prompt for it. The fetched text is piped straight into the model’s context, so the response is always grounded in what the page actually said at fetch time, not stale training data.

Limits and tips

  • Output is capped at 16 KB. Very long pages will be truncated — a CSS selector helps a lot.
  • Hard timeout is 15 seconds; the worker will abort and return an error to the AI rather than hang.
  • Public hosts only — both http:// and https:// URLs are accepted (any other scheme is rejected), and the SSRF guard blocks loopback / private-network addresses outright.
  • Custom HTTP headers are not configurable from the dashboard. The fetch goes out with its own User-Agent and a tool has no header field in the UI. If a page needs a particular header, contact support — or model the data behind a real API integration, where auth is configurable.
  • If the page is gated by login, this tool won’t see it. Model the data behind a real API integration instead.
  • Requests come from a fixed address, 64.7.198.218. If your site sits behind a WAF or rate-limiter that challenges unknown IPs, allowlist it.

Was this article helpful?