---
title: "Index URL"
method: POST
path: "/v2/collections/{collection_name}/index/url"
tags: ["indexing"]
---

# Index URL

`POST /v2/collections/{collection_name}/index/url`

Index documents from public URL(s) into a collection.

Accepts either a single `url` string or a `urls` array of strings.
Documents are downloaded from the provided URLs and processed through
the same pipeline as cloud storage indexing.

Supported formats: PDF, TXT, DOCX, CSV, XLSX, and other document types.

Headers:
- Authorization: Bearer {api_key} - Captain API key for authentication
- X-Organization-ID: Organization UUID
- Idempotency-Key: UUID for request deduplication (optional)

Args:
    collection_name: Name of the collection (path parameter)
    body: URL configuration with url or urls

Returns:
    { job_id, status: "pending" }

## Path parameters

- `collection_name` string, required

## Headers

- `authorization` string, nullable

## Request body

- IndexURLRequest
  - `custom_metadata` object, nullable — Custom metadata to attach to all indexed chunks. Keys must be strings. Values: str, int, float, bool, or List[str].
  - `parsing_script` string, nullable — Relative path to a JS parsing script for JSON files (e.g. 'research/paper-parser'). When provided, .json files are processed through a sandboxed V8 isolate. Without this, .json files are indexed as raw text.
  - `processing_type` 'advanced' | 'basic', required — Document processing type. 'advanced' uses agentic OCR with AI-enhanced extraction for complex layouts, tables, figures, charts, and documents containing images. 'basic' provides reliable OCR optimized for general document indexing and high-volume processing.
  - `transcription_language` string, nullable — AWS Transcribe language code for the spoken audio (e.g. 'es-US', 'pt-BR'). Omit to auto-detect per file. Video and audio files only. Supported codes: https://docs.aws.amazon.com/transcribe/latest/dg/supported-languages.html
  - `url` string, nullable — A single public URL to a hosted document. Supported types: PDF, DOCX, DOC, XLSX, XLS, CSV, TSV, TXT, MD, JSON, YAML, YML, PNG, JPG, JPEG, GIF, BMP, TIFF. Provide either 'url' or 'urls', not both.
  - `urls` string[], nullable — A list of public URLs to hosted documents. Provide either 'url' or 'urls', not both.
  - `mask_pii` boolean — When true, detected PII (emails, phone numbers, SSNs, credit cards, names, and locations) is masked in the parsed content before it is embedded and stored — replaced with entity tags like <PERSON> and <EMAIL_ADDRESS>. For images (including images embedded in PDFs), PII text visible in the image is also pixel-redacted. Opt-in; defaults to false, which leaves content unchanged.

## Response `200`

Successful Response

- IndexJobResponse
  - `job_id` string, required
  - `status` string
  - `custom_metadata` object, nullable — The custom_metadata Captain accepted for this job, echoed back as validated. Null when none was supplied.

---

[API](https://skmtc.dev/runcaptain/apis/api-reference.md) · [All operations](https://skmtc.dev/runcaptain/apis/api-reference/llms.txt) · [OpenAPI document](https://skmtc-service-production.skmtc.workers.dev/v1/apis/runcaptain/api-reference/revisions/fe4649e3f4aa/schema)
