Parse

Parse Document

Parse a document using OCR.

This endpoint extracts text from uploaded documents (PDF, images, Office files, etc.) using advanced OCR technology.

Providers:

  • mistral: Mistral OCR (fast, accurate, supports markdown)
  • gemini: Google Gemini Vision (slower, very accurate)

Features:

  • Automatic provider fallback if primary fails
  • Page-level text extraction
  • Markdown-formatted output
  • Support for 30+ file types

Example:

curl -X POST https://api.chonkie.ai/v1/parse \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -F "file=@document.pdf" \
  -F "provider=mistral" \
  -F "return_pages=true"
post/v1/parse

Query parameters

provider'mistral' | 'gemini'

OCR provider to use

OCR provider to use

fallbackboolean

Whether to fallback to other providers on failure

Whether to fallback to other providers on failure

return_pagesboolean

Whether to return page-level information

Whether to return page-level information

Headers

authorizationstring nullable

Response

Successful Response

textstring required

Full extracted text

Changes

No recorded changes to this endpoint across all 1 revision of this API.