---
title: "GPU utilization, memory, power and temperature time series for a pod"
method: GET
path: "/v1/vms/{id}/metrics"
tags: ["VMs"]
---

# GPU utilization, memory, power and temperature time series for a pod

`GET /v1/vms/{id}/metrics`

Host-side GPU telemetry for a pod's GPUs (nvidia-smi on NVIDIA hosts, amd-smi on AMD), sampled on the node that runs it and returned as one continuous time grid. Every series and every GPU in a response shares that grid: the value at index i of any values array was measured at start plus i times stepSeconds, and a null means the node reported no sample at that point. The window ends now and covers only the periods this pod actually held the GPUs it is reported against, so a card that ran another customer's pod earlier never contributes a sample here.
A GPU attached to a VM is passed through to the guest, so the host cannot read it. That case answers 200 with an accelerator status of unavailable and an empty gpus list rather than an error, and so do a pod that has not started yet, a pod that was not running in the selected range, and a metrics store that cannot be reached: the accelerator object is the answer, and the client renders its detail sentence. A host that stopped reporting while the pod is still running answers degraded with the history up to the last sample still attached.
This operation is exposed as an agent tool by default, in the same class as a pod's burn rate: "how busy were my GPUs" is a read-only question about a resource the caller already owns.

## Path parameters

- `id` string, required

## Query parameters

- `range` '1h' | '6h' | '24h' | '7d'

## Response `200`

The pod's GPU time series

- VmMetrics
  - `accelerator` AcceleratorCapability, required — whether accelerator telemetry could be read at all, and why not when it could not. The status is the answer: an unreadable accelerator is still a 200.
    - `class` 'accelerator', required
    - `detail` string — one sentence explaining the status, safe to show a customer as-is
    - `reason` 'accelerator_passthrough' | 'guest_not_started' | 'no_session_in_range' | 'no_gpus_recorded' | 'source_unreachable' | 'source_stale' — machine-readable cause when status is not available. Absent when the cause has no code of its own, in which case detail carries it.
    - `source` 'nvidia-smi' | 'amd-smi' — how the numbers were measured: nvidia-smi on an NVIDIA host, amd-smi on an AMD host; present whenever the response carries series
    - `status` 'available' | 'degraded' | 'unavailable', required — available = the GPUs were readable; degraded = the metrics store could not be read, or the host stopped reporting while the pod is still running; unavailable = there is nothing to read for this pod.
  - `end` string, date-time, required — last point of the returned grid. It is the newest grid point at or before the moment the request was served, so it can be up to one step older than that moment.
  - `gpus` VmGpuMetrics[], required — one entry per GPU this pod holds or held during the returned window, in the order the pod claimed them. Empty whenever the response carries no series at all, which is every unavailable status and the source_unreachable one.
    - `bdf` string, required — the GPU's PCI address on its host, lowercase with a four-digit domain (for example 0000:04:00.0)
    - `busySeconds` number, double, required — estimated GPU-busy time over the returned window, integrated from the node's utilization samples. Counted from the first sample inside the window, so up to one step of busy time at the window start is not included. 0 when the node reported nothing.
    - `series` VmGpuSeries[], required
      - `key` 'utilization_ratio' | 'memory_used_bytes' | 'memory_total_bytes' | 'power_watts' | 'temperature_celsius', required
      - `unit` 'ratio' | 'bytes' | 'watts' | 'celsius', required
      - `values` number[], required — one value per grid point, oldest first: the value at index i was measured at start plus i times stepSeconds. The array is pointCount long in every series and every GPU of one response, so a client reads the grid index-parallel and never needs the timestamps. null where the node reported no sample; a gap is never filled with 0.
  - `pointCount` integer, required — length of the values array of every series in this response. 0 when nothing was queried, in which case start equals end.
  - `range` '1h' | '6h' | '24h' | '7d', required — the requested window length; the window always ends at the moment the request was served
  - `start` string, date-time, required — first point of the returned grid, after the requested range was clipped to the periods this pod actually held GPUs on its current node. Equal to end when nothing was queried.
  - `stepSeconds` integer, required — seconds between grid points: 60 for a 1h range, 300 for 6h, 900 for 24h, 3600 for 7d
  - `vmId` string, required

## Other responses

- `400` — The request is invalid
- `401` — Missing or invalid API key
- `403` — API key lacks the required scope
- `404` — Resource not found

## Changes

- **2026-09-24** `fb9c09eb89bc` — 1 warning
  - added the new `amd-smi` enum value to the `accelerator/source` response property for the response status `200`
- **2026-09-07** `cff62d8d636e` — 1 info
  - endpoint added

[Change history](https://skmtc.dev/openrelay/apis/openrelay-api/changes/v1/vms/:id/metrics/get.md)

---

[API](https://skmtc.dev/openrelay/apis/openrelay-api.md) · [All operations](https://skmtc.dev/openrelay/apis/openrelay-api/llms.txt) · [OpenAPI document](https://skmtc.dev/openrelay/apis/openrelay-api/revisions/b684124c1a8d?raw)
