---
title: "Fetch and extract content from URLs"
method: POST
path: "/"
---

# Fetch and extract content from URLs

`POST /`

Fetches web pages, renders JavaScript-heavy pages when needed, and returns clean extracted content in your preferred format. Submit up to 10 URLs, get back structured content. Per-URL failures appear in `errors[]` and do not fail the entire request.

**Per-URL error codes** (in `errors[].error`):
- `target_http_error` — target server returned a non-2xx HTTP status other than 404/410; the raw status code is in `errors[].status`
- `page_not_found` — target URL returned HTTP 404 or 410; the raw status code is in `errors[].status`
- `target_unreachable` — connection refused, TLS failure, DNS failure, or other network error
- `timeout` — request timed out
- `proxy_error` — proxy tunnel failure
- `bot_blocked` — bot-challenge page detected (Cloudflare, etc.)
- `empty_content` — page loaded but no extractable text was found
- `invalid_url` — malformed URL or SSRF-blocked address
- `invalid_redirect_url` — redirect target rejected before fetch
- `conditional_unsupported` — conditional requests (`if_none_match` / `if_modified_since`) are supported on the fast path only; this URL requires browser rendering
- `selector_not_matched` — no elements matching any `include_selectors` entry remained after `exclude_selectors` was applied; the error carries `unmatched_selectors` plus `candidate_selectors` retry hints (a partial miss is not an error — it's reported on the result's `unmatched_selectors`)
- `selector_unsupported` — `include_selectors` / `exclude_selectors` sent for a URL that resolves to a direct PDF/CSV download (no HTML to scope)

## Request body

- object
  - `urls` string[], required — Array of URLs to fetch (1-10). All URLs are fetched in parallel. Each URL is processed independently — if one fails, others still return successfully. Errors are reported per-URL in the errors array.
  - `purpose` string — Why these URLs are being fetched — the underlying goal or task the content will be used for. Used to better tailor fetching and extraction to your intent.
  - `format` 'markdown' | 'html' | 'json' — Output format for extracted content. "markdown" (default) is ideal for LLM consumption. "html" returns cleaned semantic HTML. "json" returns a structured document tree.
  - `links` boolean — Extract all outbound links (<a href>) from each page. Useful for discovering related pages or navigating to specific content. Links are returned as absolute URLs in the links array of each result.
  - `image_links` boolean — Extract all image URLs (<img src>) from each page. Useful for finding visual content or media assets. Image links are returned as absolute URLs in the image_links array of each result.
  - `ttl` integer — Caller freshness tolerance in seconds for the cached entry. Omit (default) for unlimited tolerance — any cached entry is acceptable. Set to 0 to prefer a live fetch; a cached entry is still served if the origin's Cache-Control: max-age covers its age, or the host is in the small allowlist of operator-pinned never-expire domains. Set to N > 0 to accept a cached entry whose age is below N; the upstream Cache-Control: max-age and the never-expire allowlist may extend (never shorten) this tolerance.
  - `per_url_timeout_ms` integer — Wall-clock timeout budget in milliseconds applied independently to each URL. If one URL exceeds this budget, it returns a per-URL timeout error while other URLs in the same request continue.
  - `if_none_match` string — ETag validator from a prior fetch of this URL, forwarded verbatim as the If-None-Match header on the origin request. Only valid with a single URL — combining with a batch of URLs returns a 400. tf-fetch does not persist validators; the caller owns replaying them.
  - `if_modified_since` string — Last-Modified validator from a prior fetch of this URL, forwarded verbatim as the If-Modified-Since header on the origin request. Only valid with a single URL — combining with a batch of URLs returns a 400. tf-fetch does not persist validators; the caller owns replaying them.
  - `include_etag_and_last_modified` boolean — Opt-in to receiving `etag` / `last_modified` validators (and `not_modified` detection) on each result. Defaults to false — tf-fetch omits these fields unless requested. Independent of `if_none_match` / `if_modified_since`: works with a single URL or a batch.
  - `include_selectors` string[] — Array of CSS selectors (1-20 entries, each 1-1000 characters) that scope extracted content (`text`, `links`, `image_links`) to elements matching ANY entry, concatenated in document order. Tag selectors cover semantic sections (`main`, `article`, `nav`); entries may themselves use CSS comma-grouping. Selected content is returned verbatim in the requested format (scripts/styles stripped) — automatic boilerplate removal is bypassed. Page-level metadata (`title`, `description`, `language`, `author`, `published_date`) still comes from the full document. When some entries match and others do not, the URL still succeeds and the misses are reported in the result's `unmatched_selectors`. When no entry matches anything, that URL fails with the per-URL error code `selector_not_matched` (carrying `unmatched_selectors` and `candidate_selectors` retry hints) — never a silent full-page fallback. URLs that resolve to direct PDF/CSV downloads fail with `selector_unsupported`. Invalid CSS selector syntax is rejected with a 422. Applied post-fetch: caching and routing are unchanged.
  - `exclude_selectors` string[] — Array of CSS selectors (1-20 entries, each 1-1000 characters) for elements to remove before extraction — applied before `include_selectors` scopes what remains, so it also prunes inside selected regions. Entries may themselves use CSS comma-grouping. Entries that match nothing are a no-op, never an error, but URLs that resolve to direct PDF/CSV downloads fail with `selector_unsupported`. Invalid CSS selector syntax is rejected with a 422. Applied post-fetch: caching and routing are unchanged.

## Response `200`

Fetch completed. Check `errors[]` for any per-URL failures.

- object — Fetch response with results and errors
  - `results` object[], required — Successfully fetched URLs
    - `url` string, required — Original requested URL
    - `final_url` string, nullable, required — Final URL after redirects
    - `title` string, nullable, required — Page title
    - `description` string, nullable, required — Page meta description
    - `language` string, nullable, required — Detected language code
    - `format` 'markdown' | 'html' | 'json', required — Format of the extracted text
    - `text` union, required — Extracted content (string for markdown/html, object for json). Null if extraction failed.
      - string
      - unknown
      - unknown
    - `author` string, nullable, required — Page author if available
    - `published_date` string, nullable, required — Published date if available
    - `links` string[] — Extracted links (only if links=true in request)
    - `image_links` string[] — Extracted image links (only if image_links=true in request)
    - `latency_ms` number, nullable — Fetch latency for this URL in milliseconds if available from tf-fetch
    - `not_modified` boolean — Present and `true` only when the origin returned 304 Not Modified for your `if_none_match`/`if_modified_since`; omitted otherwise.
    - `etag` string — The origin's current ETag validator, present only when the origin sent one and you set `include_etag_and_last_modified: true`; omitted otherwise. Store it and replay it as `if_none_match`.
    - `last_modified` string — The origin's current Last-Modified validator, present only when the origin sent one and you set `include_etag_and_last_modified: true`; omitted otherwise. Store it and replay it as `if_modified_since`.
    - `unmatched_selectors` string[] — Partial-miss report for `include_selectors`: the entries that matched no elements on this page, present only when some (but not all) entries matched — the result still carries the matched content. Omitted when every entry matched. A total miss is not a result at all: it is reported as a `selector_not_matched` entry in `errors[]`.
  - `errors` object[], required — URLs that failed to fetch
    - `url` string, required — URL that failed to fetch
    - `error` string, required — Error code describing the failure. Possible values: `target_http_error` (target server returned a non-2xx HTTP status other than 404/410 — see `status` field), `page_not_found` (target returned HTTP 404 or 410), `target_unreachable` (connection refused, TLS failure, or network error), `timeout` (request timed out), `proxy_error` (proxy tunnel failure), `bot_blocked` (bot-challenge page detected), `empty_content` (page loaded but no extractable text), `invalid_url` (malformed URL or SSRF-blocked address), `invalid_redirect_url` (redirect target rejected before fetch), `conditional_unsupported` (conditional requests are supported on the fast path only; this URL requires browser rendering), `selector_not_matched` (no elements matching any `include_selectors` entry remained after `exclude_selectors` was applied — a total miss; carries `unmatched_selectors` and `candidate_selectors` retry hints. A partial miss is NOT an error: it is reported via the result's `unmatched_selectors`), `selector_unsupported` (`include_selectors`/`exclude_selectors` sent for a URL that resolves to a direct PDF/CSV download — no HTML to scope).
    - `status` integer — Upstream HTTP status code. Only present when `error` is `target_http_error` or `page_not_found`.
    - `unmatched_selectors` string[] — The `include_selectors` entries that matched no elements on the fetched page. Only present when `error` is `selector_not_matched`.
    - `candidate_selectors` string[] — Cheap retry hints (max 10), only present when `error` is `selector_not_matched`: landmark tags present on the page (e.g. `main`, `article`, `nav`) plus `#id` selectors of content-heavy elements. Retry the fetch with `include_selectors` drawn from this list.

## Other responses

- `400` — Invalid request — missing `urls`, too many URLs (max 10), bad parameter value, or `if_none_match`/`if_modified_since` combined with a batch of URLs
- `401` — Unauthorized - Invalid or missing API key
- `429` — Rate limit exceeded
- `500` — Internal server error

---

[API](https://skmtc.dev/tinyfish/apis/tinyfish-web-agent-automation-api.md) · [All operations](https://skmtc.dev/tinyfish/apis/tinyfish-web-agent-automation-api/llms.txt) · [OpenAPI document](https://skmtc-service-production.skmtc.workers.dev/v1/apis/tinyfish/tinyfish-web-agent-automation-api/revisions/bea716bc627d/schema)
