---
title: "Get a dedicated endpoint's monitoring series"
method: GET
path: "/v1/inference/endpoints/{id}/series"
tags: ["Dedicated Endpoints"]
---

# Get a dedicated endpoint's monitoring series

`GET /v1/inference/endpoints/{id}/series`

Requests, error rate, latency and time-to-first-token percentiles, and tokens for the endpoint over the range, on one time grid, plus a whole-range figure for every series. This is what the dashboard's Monitoring tab shows. Data appears within a minute of the endpoint serving its first request. The example cuts each `values` array to its first three points; a `1h` response has 241.

## Path parameters

- `id` string, required

## Query parameters

- `range` '1h' | '6h' | '24h' | '7d', required

## Response `200`

Series

- InferenceMetrics
  - `end` string, date-time, required — the last grid point
  - `metrics` InferenceMetricSet, required — Every chart, always all five. requests: success and error per second. errors: client_error, rate_limited and server_error, each as a share of all requests. latency (request start to last byte) and ttft (streaming requests only): p50, p90, p95 and p99 in seconds. tokens: per second by kind (input, cached_input, cache_write, output), only the kinds metered in the range.
    - `errors` InferenceMetric, required
      - `series` InferenceMetricSeries[], required
        - `key` string, required — the series within its metric: success, server_error, p95, output, ...
        - `values` number[], required — one value per grid point, oldest first: the value at index i is at start plus i times stepSeconds. A per-second series reads 0 while the endpoint is idle, and a ratio or percentile is null at a point with no requests. null in a per-second series means the store had no data for the endpoint at that point, for example before its first request; a gap is never filled with 0.
        - `window` number, double, nullable, required — the figure for the whole range, computed by the metrics store: the count for a per-second unit, the share for a ratio, the percentile over every request in the range for seconds (never an average of the per-point percentiles). null where it is undefined: a ratio or a percentile over a range with no requests.
      - `unit` 'requests_per_second' | 'ratio' | 'seconds' | 'tokens_per_second', required
    - `latency` InferenceMetric, required
      - `series` InferenceMetricSeries[], required
        - `key` string, required — the series within its metric: success, server_error, p95, output, ...
        - `values` number[], required — one value per grid point, oldest first: the value at index i is at start plus i times stepSeconds. A per-second series reads 0 while the endpoint is idle, and a ratio or percentile is null at a point with no requests. null in a per-second series means the store had no data for the endpoint at that point, for example before its first request; a gap is never filled with 0.
        - `window` number, double, nullable, required — the figure for the whole range, computed by the metrics store: the count for a per-second unit, the share for a ratio, the percentile over every request in the range for seconds (never an average of the per-point percentiles). null where it is undefined: a ratio or a percentile over a range with no requests.
      - `unit` 'requests_per_second' | 'ratio' | 'seconds' | 'tokens_per_second', required
    - `requests` InferenceMetric, required
      - `series` InferenceMetricSeries[], required
        - `key` string, required — the series within its metric: success, server_error, p95, output, ...
        - `values` number[], required — one value per grid point, oldest first: the value at index i is at start plus i times stepSeconds. A per-second series reads 0 while the endpoint is idle, and a ratio or percentile is null at a point with no requests. null in a per-second series means the store had no data for the endpoint at that point, for example before its first request; a gap is never filled with 0.
        - `window` number, double, nullable, required — the figure for the whole range, computed by the metrics store: the count for a per-second unit, the share for a ratio, the percentile over every request in the range for seconds (never an average of the per-point percentiles). null where it is undefined: a ratio or a percentile over a range with no requests.
      - `unit` 'requests_per_second' | 'ratio' | 'seconds' | 'tokens_per_second', required
    - `tokens` InferenceMetric, required
      - `series` InferenceMetricSeries[], required
        - `key` string, required — the series within its metric: success, server_error, p95, output, ...
        - `values` number[], required — one value per grid point, oldest first: the value at index i is at start plus i times stepSeconds. A per-second series reads 0 while the endpoint is idle, and a ratio or percentile is null at a point with no requests. null in a per-second series means the store had no data for the endpoint at that point, for example before its first request; a gap is never filled with 0.
        - `window` number, double, nullable, required — the figure for the whole range, computed by the metrics store: the count for a per-second unit, the share for a ratio, the percentile over every request in the range for seconds (never an average of the per-point percentiles). null where it is undefined: a ratio or a percentile over a range with no requests.
      - `unit` 'requests_per_second' | 'ratio' | 'seconds' | 'tokens_per_second', required
    - `ttft` InferenceMetric, required
      - `series` InferenceMetricSeries[], required
        - `key` string, required — the series within its metric: success, server_error, p95, output, ...
        - `values` number[], required — one value per grid point, oldest first: the value at index i is at start plus i times stepSeconds. A per-second series reads 0 while the endpoint is idle, and a ratio or percentile is null at a point with no requests. null in a per-second series means the store had no data for the endpoint at that point, for example before its first request; a gap is never filled with 0.
        - `window` number, double, nullable, required — the figure for the whole range, computed by the metrics store: the count for a per-second unit, the share for a ratio, the percentile over every request in the range for seconds (never an average of the per-point percentiles). null where it is undefined: a ratio or a percentile over a range with no requests.
      - `unit` 'requests_per_second' | 'ratio' | 'seconds' | 'tokens_per_second', required
  - `model` string, required — the public model id every series is filtered to
  - `pointCount` integer, required — the length of every series' values
  - `range` '1h' | '6h' | '24h' | '7d', required
  - `start` string, date-time, required — the first grid point
  - `stepSeconds` integer, required

## Other responses

- `400` — The request is invalid
- `403` — API key lacks the required scope
- `404` — Resource not found
- `502` — An upstream dependency answered badly; retry
- `503` — A required integration (e.g. payments) is not configured

## Changes

- **2026-09-27** `b684124c1a8d` — 1 info
  - endpoint added

[Change history](https://skmtc.dev/openrelay/apis/openrelay-api/changes/v1/inference/endpoints/:id/series/get.md)

---

[API](https://skmtc.dev/openrelay/apis/openrelay-api.md) · [All operations](https://skmtc.dev/openrelay/apis/openrelay-api/llms.txt) · [OpenAPI document](https://skmtc.dev/openrelay/apis/openrelay-api/revisions/b684124c1a8d?raw)
