---
title: "Scrape Markdown"
method: GET
path: "/web/scrape/markdown"
tags: ["Web Scraping"]
---

# Scrape Markdown

`GET /web/scrape/markdown`

Scrapes the given URL into LLM usable Markdown.

## Query parameters

- `url` string, uri, required
- `includeLinks` boolean
- `includeImages` boolean
- `shortenBase64Images` boolean
- `useMainContentOnly` boolean
- `pdf` object
  - `shouldParse` boolean — When true, PDF URLs are fetched and parsed. When false, PDF URLs are skipped and a 400 WEBSITE_ACCESS_ERROR is returned.
  - `start` integer — First 1-based PDF page to parse. When omitted, parsing starts at the first page.
  - `end` integer — Last 1-based PDF page to parse. When omitted, parsing ends at the last page. Must be greater than or equal to start when both are provided.
- `includeFrames` boolean
- `includeSelectors` string[]
- `excludeSelectors` string[]
- `maxAgeMs` integer
- `waitForMs` integer
- `settleAnimations` boolean
- `headers` HeaderMap
- `country` 'ad' | 'ae' | 'af' | 'ag' | 'ai' | 'al' | 'am' | 'ao' | 'ar' | 'at' | 'au' | 'aw' | 'az' | 'ba' | 'bb' | 'bd' | 'be' | 'bf' | 'bg' | 'bh' | 'bi' | 'bj' | 'bm' | 'bn' | 'bo' | 'bq' | 'br' | 'bs' | 'bw' | 'by' | 'bz' | 'ca' | 'cd' | 'cf' | 'cg' | 'ch' | 'ci' | 'cl' | 'cm' | 'cn' | 'co' | 'cr' | 'cv' | 'cw' | 'cy' | 'cz' | 'de' | 'dj' | 'dk' | 'dm' | 'do' | 'dz' | 'ec' | 'ee' | 'eg' | 'es' | 'et' | 'fi' | 'fj' | 'fr' | 'ga' | 'gb' | 'gd' | 'ge' | 'gf' | 'gg' | 'gh' | 'gm' | 'gn' | 'gp' | 'gq' | 'gr' | 'gt' | 'gu' | 'gw' | 'gy' | 'hk' | 'hn' | 'hr' | 'ht' | 'hu' | 'id' | 'ie' | 'il' | 'im' | 'in' | 'iq' | 'ir' | 'is' | 'it' | 'je' | 'jm' | 'jo' | 'jp' | 'ke' | 'kg' | 'kh' | 'kn' | 'kr' | 'kw' | 'ky' | 'kz' | 'la' | 'lb' | 'lc' | 'lk' | 'lr' | 'ls' | 'lt' | 'lu' | 'lv' | 'ly' | 'ma' | 'mc' | 'md' | 'me' | 'mf' | 'mg' | 'mk' | 'ml' | 'mm' | 'mn' | 'mo' | 'mq' | 'mr' | 'mt' | 'mu' | 'mv' | 'mw' | 'mx' | 'my' | 'mz' | 'na' | 'nc' | 'ne' | 'ng' | 'ni' | 'nl' | 'no' | 'np' | 'nz' | 'om' | 'pa' | 'pe' | 'pf' | 'pg' | 'ph' | 'pk' | 'pl' | 'pr' | 'ps' | 'pt' | 'py' | 'qa' | 're' | 'ro' | 'rs' | 'ru' | 'rw' | 'sa' | 'sc' | 'sd' | 'se' | 'sg' | 'si' | 'sk' | 'sl' | 'sm' | 'sn' | 'so' | 'sr' | 'ss' | 'st' | 'sv' | 'sx' | 'sy' | 'sz' | 'tc' | 'td' | 'tg' | 'th' | 'tj' | 'tl' | 'tm' | 'tn' | 'tr' | 'tt' | 'tw' | 'tz' | 'ua' | 'ug' | 'us' | 'uy' | 'uz' | 'vc' | 've' | 'vg' | 'vi' | 'vn' | 'ye' | 'yt' | 'za' | 'zm' | 'zw' — Two-letter ISO 3166-1 alpha-2 country code identifying a supported Context.dev residential proxy exit location. Must be one of Context.dev's supported countries. When provided, Context.dev fetches the target page from that country.
- `timeoutMS` integer — Optional timeout in milliseconds for the request. If the request takes longer than this value, it will be aborted with a 408 status code. Maximum allowed value is 300000ms (5 minutes).

## Response `200`

Successful response

- object
  - `success` true, required — Indicates success
  - `markdown` string, required — Page content converted to GitHub Flavored Markdown
  - `url` string, required — The URL that was scraped
  - `metadata` PageMetadata, required — Metadata extracted from the scraped page HTML.
    - `sourceUrl` string, required — Original URL requested by the caller.
    - `finalUrl` string, required — Final URL scraped after redirects or scraper fallback, when known. Falls back to sourceUrl when unavailable.
    - `title` string — Best title extracted from the page.
    - `description` string — Best description extracted from standard, Open Graph, or Twitter metadata.
    - `language` string — Language extracted from html lang or language meta tags.
    - `keywords` string[] — Keywords extracted from the page's keywords meta tag.
    - `canonicalUrl` string — Resolved canonical URL, when present.
    - `author` string — Author metadata, when present.
    - `siteName` string — Site or application name from page metadata.
    - `image` string — Primary resolved preview image from Open Graph, Twitter, or image metadata.
    - `favicon` string — Resolved favicon URL, when present.
    - `publishedTime` string — Published timestamp/date from page metadata, when present.
    - `modifiedTime` string — Modified timestamp/date from page metadata, when present.
    - `robots` string — Robots meta directive, when present.
    - `openGraph` object — Open Graph metadata with the og: prefix removed and keys camel-cased.
    - `twitter` object — Twitter card metadata with the twitter: prefix removed and keys camel-cased.
    - `alternates` PageMetadataAlternate[] — Resolved alternate links from link rel=alternate tags.
      - `href` string, required — Resolved alternate URL.
      - `hreflang` string — Language or locale for the alternate URL, when present.
      - `type` string — Alternate resource MIME type, when present.
      - `title` string — Alternate resource title, when present.
    - `jsonLd` object[] — JSON-LD structured data blocks parsed from the page.
    - `additionalMeta` object — Additional non-social meta tags not promoted to top-level metadata fields.
  - `key_metadata` KeyMetadata — Metadata about the API key used for the request. Included in every response whenever a valid API key is provided, even when the response status is not 200.
    - `credits_consumed` integer, required — The number of credits consumed by this request.
    - `credits_remaining` integer, required — The number of credits remaining for your organization after this request.

## Other responses

- `400` — Bad request - Invalid URL or failed to scrape
- `401` — Unauthorized - Invalid or missing API key
- `403` — Forbidden - Insufficient permissions or usage limit exceeded
- `404` — Target page returned a 404
- `408` — Request timeout
- `415` — Unsupported content type - the URL resolved to a content type that is not supported (e.g. an image, presentation, media, or archive). Supported types are HTML, XML, PDF, DOCX, DOC, XLSX, XLS, PPTX, PPT, and CSV.
- `429` — Rate limit exceeded
- `500` — Internal server error

## Changes

- **2026-07-09** `de91b92d5fb7` — 2 info
  - added the new optional `query` request parameter `country`
  - added the new optional `query` request parameter `settleAnimations`
- **2026-06-22** `4767df4f281f` — 3 info
  - added the non-success response with the status `404`
  - added the non-success response with the status `415`
  - added the required property `metadata` to the response with the `200` status

[Change history](https://skmtc.dev/context/apis/context-dev/changes/web/scrape/markdown/get.md)

---

[API](https://skmtc.dev/context/apis/context-dev.md) · [All operations](https://skmtc.dev/context/apis/context-dev/llms.txt) · [OpenAPI document](https://skmtc-service-production.skmtc.workers.dev/v1/apis/context/context-dev/revisions/de91b92d5fb7/schema)
