---
title: "Get Query Latency"
method: GET
path: "/v2/queries/latency"
tags: ["queries"]
---

# Get Query Latency

`GET /v2/queries/latency`

Latency percentiles and a histogram for a window of your stored queries.

Unlike GET /v2/queries, which pages through individual records, this
summarises the whole filtered population in one bounded response: counts,
exact p50/p95/p99, explicit histogram edges, and the same bins per
collection so their shapes can be compared.

Two numbers are reported separately on purpose. `counts.measured` are the
queries with a recorded processing time; `counts.missing` are matching
queries without one, which are left out of the percentiles rather than
counted as zero milliseconds. `coverage` says how much of the population
recorded each filterable setting, because older queries predate some of
them and answer `unknown`.

Windows are capped at 90 days. Use GET /v2/queries to drill into
individual queries.

## Query parameters

- `from` string, nullable
- `to` string, nullable
- `collection` string, nullable
- `api_version` string, nullable
- `rerank` string, nullable
- `search_mode` string, nullable
- `has_filter` string, nullable
- `bins` integer, nullable
- `refresh` boolean

## Response `200`

Successful Response

- QueryLatencyResponseV2 — Bounded latency metrics for a filtered population of stored queries.
  - `collections` LatencyCollectionSeriesV2[]
    - `collection_name` string, required — The collection, or `Other` for the folded remainder.
    - `count` integer, required — Measured queries against this collection.
    - `counts` integer[], required — Per-bin counts against the same edges as the overall histogram, so shapes are comparable.
    - `is_other` boolean — True for the folded remainder rather than a real collection.
    - `p50` integer, nullable — Median in ms; null for the `Other` group, whose percentiles cannot be derived exactly.
    - `p95` integer, nullable — 95th percentile in ms; null for `Other`.
    - `p99` integer, nullable — 99th percentile in ms; null for `Other`.
  - `completeness` LatencyCompletenessV2, required
    - `complete` boolean, required — True when the window has settled past the ingestion lag, so no further rows are expected for it. This is a settled-not-filling signal, not a guarantee of wholeness: usage rows are written after the response and a failed write is not retried, so compare `watermark` against the window when exactness matters.
    - `watermark` string, nullable — The most recent query time seen in the window. How far the data actually reaches.
  - `counts` LatencyCountsV2, required
    - `matching` integer, required — Queries in the window that matched every filter.
    - `measured` integer, required — Of those, the ones with a recorded processing time.
    - `missing` integer, required — Of those, the ones with no recorded processing time. Not zero-millisecond queries.
  - `coverage` object — Per filter, the share of matching queries that recorded that dimension (0.0-1.0). A low value means most of the population answers `unknown`.
  - `filters_applied` LatencyFiltersAppliedV2, required
    - `api_version` string, nullable — `v3` or `v2`.
    - `collection` string, nullable — Only queries against this collection.
    - `has_filter` 'true' | 'false' | 'unknown' — Whether the request carried a metadata filter.
    - `query_environment` string, nullable — The API key environment the queries ran in.
    - `rerank` 'on' | 'off' | 'unknown' — The recorded rerank request setting.
    - `search_mode` 'keyword' | 'semantic' | 'hybrid' | 'unknown' — Derived from the recorded semantic ratio.
  - `generated_at` string, required — When these numbers were computed. A cached response keeps its original value.
  - `histogram` LatencyHistogramV2, required
    - `counts` integer[], required — Measured queries per bin. Bin i spans edges[i] (inclusive) to edges[i+1] (exclusive); the last bin includes its upper edge.
    - `edges` integer[], required — Bin edges in ms, ascending. One more entry than `counts`.
  - `metric` 'query_processing_time' — What is being measured: server-side query processing time, not end-to-end client latency.
  - `metric_version` string
  - `percentile_method` 'exact' — Percentiles come from the measurements themselves, not an approximate sketch.
  - `percentiles` LatencyPercentilesV2, required
    - `p50` integer, nullable — Median processing time in ms; null when nothing was measured.
    - `p95` integer, nullable — 95th percentile in ms; null when nothing was measured.
    - `p99` integer, nullable — 99th percentile in ms; null when nothing was measured.
  - `population` string, required — Which queries the metric covers. Completed, non-evaluation queries: a failed query records no usage row, and evaluation units are excluded.
  - `unit` 'ms'
  - `window` LatencyWindowV2, required
    - `from` string, required — Start of the window, inclusive (UTC).
    - `to` string, required — End of the window, exclusive (UTC).

## Other responses

- `400` — Missing or invalid window, filter value, or bin count.
- `401` — Missing or invalid API key.
- `403` — The API key does not have query permission.
- `422` — Validation Error
- `503` — Analytics are unavailable; no metrics were computed.

## Changes

- **2026-09-24** `61a9364ad042` — 4 breaking, 1 warning, 19 info
  - the response's body type changed from no type to `object` for status `400`
  - the response's body type changed from no type to `object` for status `401`
  - the response's body type changed from no type to `object` for status `403`
  - the response's body type changed from no type to `object` for status `503`
  - …20 more
- **2026-09-17** `6ed36830de71` — 1 info
  - endpoint added

[Change history](https://skmtc.dev/runcaptain/apis/api-reference/changes/v2/queries/latency/get.md)

---

[API](https://skmtc.dev/runcaptain/apis/api-reference.md) · [All operations](https://skmtc.dev/runcaptain/apis/api-reference/llms.txt) · [OpenAPI document](https://skmtc.dev/runcaptain/apis/api-reference/revisions/ac61e472bb7d?raw)
