---
title: "Search document chunks"
method: POST
path: "/api/v3/search"
tags: ["Search"]
---

# Search document chunks

`POST /api/v3/search`

Embedding → hybrid vector search → optional reranking, returning ranked chunks
with provenance. No LLM generation is performed.

Billing: 1 retrieval credit per request.

**Relevance scoring (`relevance_scoring`):** controls the relevance scoring stage.
- `scoring_and_filtering` (default): Score candidates for relevance and only
  return those above the quality threshold.
- `scoring_only`: Score every candidate for relevance but return them all, even
  low-scoring ones. Useful for building your own filtering logic.
- `none`: Skip the relevance scoring step and return all candidates unfiltered.
  Fastest option, useful when you handle scoring yourself.

Omit `relevance_scoring` for the default; send `none` to skip scoring.
`skip_rerank` is deprecated — `true` maps to `none`, `false` to `scoring_and_filtering`.

**Result ordering:** results are returned in descending order of `score`.
With `scoring_and_filtering` or `scoring_only`, `score` equals the relevance
score (`scores.relevance`, 0–1). With `none`, `score` is the combined retrieval
score (higher is better, no fixed upper bound).

If the scoring model is temporarily unavailable, results are returned in
retrieval order and a `warnings` array is included. Each warning has a `code`
matching the degraded `scores` key (e.g. `relevance`) and a `reason` classifying the
failure: `model_not_found`, `timeout`, `service_error`, or `unknown`.
The `warnings` key is absent when all pipeline steps succeed.

**Scoping:** use `workspace_id` and/or `tag_id` to narrow results, or `file_id`
to target specific files. `file_id` cannot be combined with `workspace_id` or
`tag_id` (422). A 403 is returned when filters resolve to no authorized resources.
When no filters are provided, search runs across all documents authorized for the
API key.

**Facet filtering:** use `content_type` and `attribute` to narrow results by facet
metadata. Content type uses colon-separated paths (e.g. `legal:contract:nda`).
**Repeated `attribute` entries are ANDed; values inside one entry are ORed with
`|` (pipe, recommended).** Example: `attribute=fiscal_year:2024|2025&attribute=status:active`
→ (fiscal_year 2024 OR 2025) AND (status active). Supports operators (`>`, `>=`,
`<`, `<=`), prefix (`name:prefix*`), smart dates, and content-type scoping.

**Modes:**
- `text` (default): hybrid text search
- `vision`: VLM-embedded page image search

**Images:** set `include_image=true` to receive a base64-encoded page image with
each result. In text mode the image is fetched from the VisionChunk covering the
chunk's start page (empty string if no vision index exists for that page).

**Bounding boxes (PDF only):** set `include_bboxes=true` to append a `bboxes` array to each
result, giving the merged rectangles of the chunk's text on the source PDF (raw
PDF points, top-left origin with y extending downward) so you can overlay highlights without re-locating
the chunk. One rectangle per logical group; a chunk spanning two pages produces at
least one rectangle per page. Available for PDF documents in text mode only — returns an
empty list for non-PDF, vision-mode, or pre-v2.2.1 chunks. When `include_bboxes=false`
(default) the `bboxes` key is omitted.

## Request body

- SearchRequest — DRF serializer mixin providing ``content_type`` and ``attribute`` fields. Compose into any request serializer via multiple inheritance:: class SearchRequestSerializer(FacetFilterFieldsMixin, serializers.Serializer): query = serializers.CharField(...) # content_type and attribute inherited from the mixin
  - `content_type` string[] — Filter by content type path. Multiple values are OR. Exact-or-subtree matching by default (e.g. `legal` matches legal, legal:contract). Wildcards: `*contract*` (contains), `legal:contract*` (prefix).
  - `attribute` string[] — Filter by attribute value. **Repeated `attribute` entries are ANDed; values inside one entry are ORed with `|`** (pipe is the recommended OR delimiter — comma also works but can be ambiguous with multi-key values). Example: `attribute=fiscal_year:2024|2025&attribute=status:active` → (fiscal_year 2024 OR 2025) AND (status active). Formats: `name` (has any value), `name:value` (exact), `name:>value` / `name:>=value` (gt/gte), `name:<value` / `name:<=value` (lt/lte), `name:prefix*` (starts with, case-insensitive), `name:*text*` (contains, case-insensitive), `name:a|b` (OR). Smart dates: `filing_date:2023` (year), `filing_date:2023-06` (month). Type-aware: booleans (true/false), multi-select (membership check). Scoped: `content_type(legal:compliance).regulation:AML`.
  - `query` string, required — Natural-language search query. Maximum 1500 characters.
  - `max_results` integer — Maximum number of chunks to return after reranking. Range: 1–50.
  - `workspace_id` integer[] — Restrict search to these workspace IDs. Cannot combine with file_id.
  - `tag_id` integer[] — Restrict to documents carrying any of these tag IDs (OR). Cannot combine with file_id.
  - `file_id` integer[] — Restrict to specific file IDs. Cannot combine with workspace_id or tag_id.
  - `mode` 'text' | 'vision' — * `text` - text * `vision` - vision
  - `relevance_scoring` 'none' | 'scoring_only' | 'scoring_and_filtering' — * `none` - none * `scoring_only` - scoring_only * `scoring_and_filtering` - scoring_and_filtering
  - `skip_rerank` boolean — Deprecated — use relevance_scoring. true → relevance_scoring=none, false → relevance_scoring=scoring_and_filtering. Ignored when relevance_scoring is provided.
  - `include_image` boolean — Append a base64-encoded page image to each result.
  - `include_bboxes` boolean — Append merged bounding boxes (in PDF points, top-left origin) to each result so callers can overlay chunk highlights on PDF pages. PDF documents in text mode only — non-PDF and vision-mode results always return an empty list. Omitted from the response entirely when false.

## Response `200`

Ranked search results. Empty array if no documents match.

- SearchResponse
  - `results` SearchResultItem[], required — Retrieved chunks, ordered by score descending.
    - `chunk_id` string, uuid, required — Chunk UUID.
    - `content` string, nullable, required — Chunk text content. Null for vision-mode chunks.
    - `score` number, double, required — Effective relevance score — the sort key. Equals scores.relevance (0–1) when relevance scoring ran, otherwise the combined retrieval score (higher is better, no fixed upper bound). Results are ordered by this value descending.
    - `scores` SearchScores, required
      - `text` number, double, nullable, required — Semantic text similarity (0–1, higher is better). Null in vision mode.
      - `vision` number, double, nullable, required — Vision page similarity (0–1, higher is better). Null when the document has no vision index.
      - `keyword` number, double, nullable, required — Keyword match score (higher is better, no fixed upper bound). Null in vision mode.
      - `multivector` number, double, nullable, required — Token-level similarity score (higher is better, no fixed upper bound). Null when multi-vector scoring is disabled.
      - `relevance` number, double, nullable, required — Relevance score (0–1, higher is better). Populated when relevance_scoring is "scoring_only" or "scoring_and_filtering". Null when relevance_scoring is "none" or when the scoring model is unavailable.
    - `image` SearchImage
      - `b64_content` string, required — Base64-encoded page image. Empty string when no vision index exists for the page.
    - `source` SearchSource, required
      - `file_id` integer, required — File ID.
      - `filename` string, required — Original filename.
      - `title` string, nullable, required — Document title.
      - `mime_type` string, nullable, required — File type (e.g. pdf, docx).
      - `size_bytes` integer, nullable, required — File size in bytes.
      - `page_start` integer, nullable, required — Start page of the chunk (1-indexed).
      - `page_end` integer, nullable, required — End page of the chunk (1-indexed).
      - `total_pages` integer, required — Total pages in the document.
      - `tags` SearchTag[], required — Tags associated with the document.
        - `id` integer, required — Tag ID.
        - `name` string, required — Tag name.
      - `content_types` object[] — Facet content type classifications and attribute values.
      - `external_metadata` SearchExternalMetadata, required
        - `external_id` string, required — ID of the document in the external system.
        - `external_url` string, nullable, required — Deep-link back to the document in the source system. Null if not provided.
        - `additional_metadata` object, required — Freeform connector metadata. external_url is lifted to its own field and excluded here.
    - `workspace` SearchWorkspace, required
      - `id` integer, required — Workspace ID.
      - `name` string, required — Workspace name.
    - `bboxes` SearchBbox[] — Merged bounding boxes for the chunk's text on the source PDF. Present only when include_bboxes=true. Empty list for vision-mode, non-PDF, or pre-v2.2.1 chunks.
      - `page_number` integer, required — 1-indexed page the rectangle sits on.
      - `x` number, double, required — Left edge in PDF points, top-left origin.
      - `y` number, double, required — Top edge in PDF points, top-left origin (y extends downward).
      - `width` number, double, required — Width in PDF points.
      - `height` number, double, required — Height in PDF points.
      - `unit` string, required — Coordinate unit. Always "pdf_point" in v1.
      - `origin` string, required — Coordinate origin. Always "top_left" in v1.
  - `warnings` SearchWarning[] — Present only when a pipeline signal degrades. Absent in the happy path.
    - `code` string, required — Signal name from the scores object that degraded (e.g. 'relevance').
    - `reason` string — Machine-readable failure reason (model_not_found, timeout, service_error, unknown).
  - `explain` object — Scoring breakdown. Present only when explain=true and SEARCH_EXPLAIN_MODE is enabled.

## Other responses

- `400` — Request body is not parsable JSON.
- `401` — Missing or invalid API key.
- `403` — API key has no authorized resources matching the provided filters.
- `422` — Field validation failure.
- `429` — Rate limit exceeded.
- `500` — Unexpected server error.

---

[API](https://skmtc.dev/lighton/apis/lighton-api.md) · [All operations](https://skmtc.dev/lighton/apis/lighton-api/llms.txt) · [OpenAPI document](https://skmtc-service-production.skmtc.workers.dev/v1/apis/lighton/lighton-api/revisions/86c57e94ef15/schema)
