---
title: "Index S3 Directory"
method: POST
path: "/v2/collections/{collection_name}/index/s3/directory"
tags: ["indexing"]
---

# Index S3 Directory

`POST /v2/collections/{collection_name}/index/s3/directory`

Index all files from a specific directory in an S3 bucket.

Headers:
- Authorization: Bearer {api_key} - Captain API key for authentication
- X-Organization-ID: Organization UUID
- Idempotency-Key: UUID for request deduplication (optional)

Args:
    collection_name: Name of the collection (path parameter)
    body: S3 directory configuration with directory_path

Returns:
    { job_id, status: "pending" }

## Path parameters

- `collection_name` string, required

## Headers

- `authorization` string, nullable

## Request body

- IndexS3DirectoryRequest
  - `auth` S3AssumeRoleAuth
    - `external_id` string, required — External ID Captain issued you when you enrolled the AWS integration. Prevents the confused-deputy problem — your role's trust policy must require this exact value via an sts:ExternalId condition.
    - `role_arn` string, required — ARN of the IAM role in your AWS account for Captain to assume (e.g. 'arn:aws:iam::123456789012:role/CaptainS3ReadRole'). The role's trust policy must allow Captain's principal and require the external_id below.
    - `type` 'assume_role', required — Auth type discriminator. Currently only 'assume_role' is supported.
  - `aws_access_key_id` string, nullable — AWS access key ID with read access to the bucket. Use this for long-lived IAM-user credentials. Omit when using the role-based 'auth' block.
  - `aws_secret_access_key` string, nullable — AWS secret access key. Use this for long-lived IAM-user credentials. Omit when using the role-based 'auth' block.
  - `bucket_name` string, required
  - `bucket_region` string
  - `custom_metadata` object, nullable — Custom metadata to attach to all indexed chunks. Keys must be strings. Values: str, int, float, bool, or List[str].
  - `directory_path` string, required — Path to the directory within the bucket. Accepts either a relative path (e.g., 'reports/2024/january') or a full S3 URI (e.g., 's3://my-bucket/reports/2024/january'). All files within this directory and its subdirectories will be indexed.
  - `max_files` integer, nullable
  - `overwrite_existing` boolean — When true, files that already exist in the collection are re-indexed and replaced with zero downtime: the new version is built alongside the live one and atomically swapped in when complete, so the previous version keeps serving search results throughout the rebuild. The document keeps the same document_id across overwrites, and its status reads 'updating' in the document listing while the rebuild runs. Requires skip_existing=false. Setting both to true returns a 400 error.
  - `parsing_script` string, nullable — Relative path to a JS parsing script for JSON files (e.g. 'research/paper-parser'). When provided, .json files are processed through a sandboxed V8 isolate. Without this, .json files are indexed as raw text.
  - `processing_type` 'advanced' | 'basic', required — Document processing type. 'advanced' uses agentic OCR with AI-enhanced extraction for complex layouts, tables, figures, charts, and documents containing images. 'basic' provides reliable OCR optimized for general document indexing and high-volume processing.
  - `skip_existing` boolean — When true, files already indexed in the collection are skipped and will not be re-indexed with incoming changes. When false, all incoming files are indexed regardless of whether they already exist.
  - `mask_pii` boolean — When true, detected PII (emails, phone numbers, SSNs, credit cards, names, and locations) is masked in the parsed content before it is embedded and stored — replaced with entity tags like <PERSON> and <EMAIL_ADDRESS>. For images (including images embedded in PDFs), PII text visible in the image is also pixel-redacted. Opt-in; defaults to false, which leaves content unchanged.
  - `transcription_language` string, nullable — AWS Transcribe language code for the spoken audio (e.g. 'es-US', 'pt-BR'). Omit to auto-detect per file. Video and audio files only. Supported codes: https://docs.aws.amazon.com/transcribe/latest/dg/supported-languages.html

## Response `200`

Successful Response

- IndexJobResponse
  - `job_id` string, required
  - `status` string
  - `custom_metadata` object, nullable — The custom_metadata Captain accepted for this job, echoed back as validated. Null when none was supplied.

## Other responses

- `400` — overwrite_existing and skip_existing cannot both be true

---

[API](https://skmtc.dev/runcaptain/apis/api-reference.md) · [All operations](https://skmtc.dev/runcaptain/apis/api-reference/llms.txt) · [OpenAPI document](https://skmtc-service-production.skmtc.workers.dev/v1/apis/runcaptain/api-reference/revisions/fe4649e3f4aa/schema)
