---
title: "Queue Evaluation"
method: POST
path: "/v3/collections/{collection_name}/evals"
tags: ["file-search"]
---

# Queue Evaluation

`POST /v3/collections/{collection_name}/evals`

Queue an evaluation of an uploaded case set against one to eight named query configurations. `Idempotency-Key` is required: within 24 hours the same key with the same body returns the stored `201` body with status `200`; the same key with a different body is `409` (`IDEMPOTENCY_KEY_REUSE`). Gold is resolved at accept time (document id, then unique basename); unresolved cases are `error` and never queried or billed. Each `configs[]` entry is the [v3 query](/reference/v3/query) body minus `query` and is validated the same way. Each unit (one question under one config) is billed as one query, so whether it costs credits depends on the plan; only units that ran (scored, or timed out at the 30 s limit) are counted. Scoring, limits, billing and the NDJSON format are on the [Evaluations](/guides/evaluations) guide.

## Path parameters

- `collection_name` string, required

## Headers

- `Idempotency-Key` string, nullable, required

## Request body

- CreateEvalRequestV3
  - `configs` EvalConfigV3[], required
    - `boost` BoostClauseV3[], nullable
      - `field` string, nullable — Required with `eq` or `in`; omit it for a `chunk_ids` rule. One of your own custom_metadata keys, the same keys `filter` accepts. Anything filterable is boostable, with no re-indexing. Captain's internal attributes (file_id, job_id, chunk_index and similar) return a 400.
      - `eq` union — Requires `field`. Boost chunks whose `field` equals this value. Strings, numbers, and booleans are supported. A list-valued field matches when any element equals it. Use `in` instead for several values.
        - string
        - integer
        - number, double
        - boolean
      - `in` BoostClauseV3InItems[], nullable — Requires `field`. Boost chunks whose `field` matches any of these values, up to 50. Use it for one key with several acceptable values, and for list-valued metadata.
        - union
          - string
          - integer
          - number, double
          - boolean
      - `chunk_ids` string[], nullable — Use instead of `field`/`eq`/`in` to boost specific chunks, up to 100. Accepts a `chunk_id` from a previous response, or `document_id:chunk_index`, which keeps resolving after the document is re-indexed. Use it to keep a follow-up question anchored to passages an earlier answer cited.
      - `weight` number, double, required — Required. Multiplier on the matching chunk's retrieval score, 0.2 to 5.0. The rule's retrieval pass is what makes a chunk appear; the weight decides where it lands. 1.5 to 2.0 is a good default and is usually enough to lift a tagged chunk to the top. Use 3.0 to 5.0 when boosted chunks compete with each other, or when a rule matches many chunks. Below 1.0 demotes, and only reorders chunks the query already found. When several rules match the same chunk their weights multiply, capped at 5.0.
      - `reserve` integer — With reranking on, guarantee this many of this rule's chunks a place on the page, chosen by reranker order. Reranking runs after boosting and can otherwise drop a boosted chunk whose text does not resemble the query. 0 (default) leaves the reranker's judgement in charge. The sum across all boost rules cannot exceed `limit`.
    - `exclude_chunk_types` string[], nullable
    - `filter` EvalConfigV3Filter
    - `include` QueryIncludeV3 — Optional expansions for v3 query results. Leave expensive context off unless the caller needs it.
      - `document` boolean — Include the parent document object for each returned chunk.
      - `metadata` boolean — Include Captain-generated retrieval metadata and application-supplied chunk metadata.
      - `regions` boolean — Include extracted layout regions, including bounding boxes when available.
      - `relations` boolean — Include graph relations connected to each returned chunk.
      - `related_chunks` boolean — Include linked chunks for returned relations.
      - `neighboring_chunks` boolean — Include the chunk immediately before and after each result in the source file.
      - `document_metadata` boolean
      - `archived` boolean — Include chunks archived by a sync `archive` deletion policy. Archived content is excluded from search by default; set true to surface it.
    - `limit` integer
    - `max_chunks_per_document` integer, nullable
    - `name` string, required
    - `relation_direction` 'outgoing' | 'incoming' | 'both'
    - `relation_types` string[], nullable
    - `rerank` union
      - boolean
      - RerankOptions — Object form of the `rerank` parameter. The boolean form stays valid: `true` is equivalent to sending this object with every field at its default.
        - `enabled` boolean — Whether to rerank. Sending the object without this field means enabled — the object form exists to tune reranking.
        - `candidate_limit` integer, nullable — How many fused retrieval candidates are fetched and reranked before the top `limit` results are returned. Defaults to `limit` x 3 (the measured configuration). Must be >= `limit`; capped at 200. Larger pools can lift recall on corpora with many near-duplicate documents, at the cost of rerank latency.
        - `model` string, nullable — Reranker model. Voyage cross-encoders: `voyage-rerank-2.5` (default) and `voyage-rerank-3` (aliases `rerank-2.5`, `rerank-3`). Gemini LLM rerankers: `gemini-2.5-flash` and `gemini-3.8-flash` — higher quality ceiling on complex queries, higher latency; best with small candidate pools. The family aliases `voyage` and `gemini` resolve to each family's default. Unknown values return a 400 listing the allowed set.
    - `semantic_ratio` number, double
  - `environment` string, nullable — Echo-only label for the scorecard. Does not switch collections.
  - `upload_id` string, required

## Response `200`

Idempotent replay: the stored original `201` body (including the original `eval_id`), not live status. Poll [Get Eval](/reference/evals/get) for progress.

- EvalResponseV3 — One shape for the 201 accept body, the 200 idempotent replay, and GET.
  - `billing` EvalBillingV3
    - `billable_units` integer, required
    - `failed_configs` string[]
    - `skipped_units` integer, required
  - `collection_name` string, required
  - `completed_at` string, nullable
  - `configs` ResolvedConfigV3[], required
    - `explicit_fields` string[], required
    - `name` string, required
    - `query` ResolvedQueryConfigV3, required — Every Query v3 field, defaults filled in, rerank always in object form.
      - `boost` ResolvedQueryConfigV3BoostItems[]
      - `exclude_chunk_types` string[]
      - `filter` ResolvedQueryConfigV3Filter
      - `include` object
      - `limit` integer, required
      - `max_chunks_per_document` integer, nullable
      - `relation_direction` string
      - `relation_types` string[], nullable
      - `rerank` ResolvedRerankV3, required
        - `candidate_limit` integer, nullable
        - `enabled` boolean, required
        - `model` string, nullable
      - `semantic_ratio` number, double, required
  - `created_at` string, required
  - `environment` string, nullable
  - `error_code` string, nullable
  - `error_message` string, nullable
  - `eval_id` string, required
  - `idempotency_key` string, required
  - `items` EvalItemV3[]
    - `error_code` string, nullable
    - `error_message` string, nullable
    - `expected_document_ids` string[], nullable
    - `expected_files` string[], required
    - `filters` EvalItemV3Filters
    - `id` string, required
    - `query` string, required
    - `results` object
    - `status` 'pending' | 'running' | 'scored' | 'error', required
  - `items_page` EvalItemsPageV3, required
    - `limit` integer, required
    - `next_cursor` string, nullable
    - `total` integer, required
  - `preview` EvalPreviewV3, required
    - `billable_units` integer, required
    - `cases` integer, required
    - `cases_error` integer, required
    - `configs` integer, required
    - `units` integer, required
  - `progress` EvalProgressV3, required
    - `cases_error` integer, required
    - `cases_total` integer, required
    - `configs_failed` integer, required
    - `configs_total` integer, required
    - `elapsed_seconds` integer, required
    - `eta_seconds` integer, nullable
    - `percent` integer, required
    - `qps` number, double, nullable
    - `units_completed` integer, required
    - `units_error` integer, required
    - `units_failed` integer, required
    - `units_total` integer, required
  - `request_id` string, required
  - `scorecards` object
  - `started_at` string, nullable
  - `status` 'pending' | 'running' | 'completed' | 'completed_with_errors' | 'failed', required
  - `updated_at` string, required
  - `upload_id` string, required

## Other responses

- `400` — `INVALID_ID_PREFIX`, missing or invalid `Idempotency-Key`, upload not yet PUT or expired, duplicate config names, a config failing v3 query validation, or a whole-file NDJSON problem (blank line, invalid JSON, duplicate case `id`, more than 10,000 lines, more than 80,000 units).
- `401` — Missing or invalid authentication.
- `403` — API key does not have the query permission for this collection.
- `404` — Collection not found (or an eval / upload that belongs to another organization: no existence oracle).
- `409` — `IDEMPOTENCY_KEY_REUSE` (same key, different body) or `UPLOAD_ALREADY_BOUND` (upload already used by another eval).
- `429` — `TOO_MANY_PENDING_EVALS`: the organization already has 20 pending evals (`Retry-After` is set). Evals are paced as jobs, so this is never an org-wide query 429.

## Changes

- **2026-09-24** `61a9364ad042` — 6 breaking, 24 info
  - the response's body type changed from no type to `object` for status `400`
  - the response's body type changed from no type to `object` for status `401`
  - the response's body type changed from no type to `object` for status `403`
  - the response's body type changed from no type to `object` for status `404`
  - …26 more
- **2026-09-16** `e82391fa5b0b` — 1 info
  - endpoint added

[Change history](https://skmtc.dev/runcaptain/apis/api-reference/changes/v3/collections/:collection_name/evals/post.md)

---

[API](https://skmtc.dev/runcaptain/apis/api-reference.md) · [All operations](https://skmtc.dev/runcaptain/apis/api-reference/llms.txt) · [OpenAPI document](https://skmtc.dev/runcaptain/apis/api-reference/revisions/ac61e472bb7d?raw)
