Dedicated Endpoints

Get a dedicated endpoint's monitoring series

Requests, error rate, latency and time-to-first-token percentiles, and tokens for the endpoint over the range, on one time grid, plus a whole-range figure for every series. This is what the dashboard's Monitoring tab shows. Data appears within a minute of the endpoint serving its first request. The example cuts each values array to its first three points; a 1h response has 241.

get/v1/inference/endpoints/{id}/series

Path parameters

idstring required

Query parameters

range'1h' | '6h' | '24h' | '7d' required

Response

Series

endstring date-time required

the last grid point

modelstring required

the public model id every series is filtered to

pointCountinteger required

the length of every series' values

range'1h' | '6h' | '24h' | '7d' required
startstring date-time required

the first grid point

stepSecondsinteger required

Changes

Changed in 1 of the 36 revisions of this API.1