Web Scraping

Crawl Sitemap

Crawl an entire website's sitemap and return all discovered page URLs.

get/web/scrape/sitemap

Query parameters

domainstring required

Domain to build a sitemap for

maxLinksinteger

Maximum number of links to return from the sitemap crawl. Defaults to 10,000. Minimum is 1, maximum is 100,000.

urlRegexstring
Example:^https?://[^/]+/blog/

Optional RE2-compatible regex pattern. Only URLs matching this pattern are returned and counted against maxLinks.

headersHeaderMap

Optional outbound HTTP headers forwarded only to the target URL, sent as deep-object query params such as headers[X-Custom]=value. When provided, caching is bypassed: the result is neither read from nor written to cache.

timeoutMSinteger

Optional timeout in milliseconds for the request. If the request takes longer than this value, it will be aborted with a 408 status code. Maximum allowed value is 300000ms (5 minutes).

Optional timeout in milliseconds for the request. If the request takes longer than this value, it will be aborted with a 408 status code. Maximum allowed value is 300000ms (5 minutes).

Response

Successful response

successtrue required

Indicates success

domainstring required

The normalized domain that was crawled

urlsstring[] required

Array of discovered page URLs from the sitemap (max 500)

Changes