Get an iteration's step results
One row per authored test step, in order, with status (ok/fail/skipped/pending), a reason, and any evidence (screenshots, replay-video offset, widget tool calls). The fastest way to see which step failed and why. Unlike /trace, a missing trace is not a 404 — step verdicts still return, just without evidence.
Path parameters
ID of the hosted project that contains the server.
Eval run ID, as returned by POST /eval-runs.
Iteration ID, as returned by the run's iterations list.
Headers
Which vocabulary this request and its response speak. Absent means 1, which is byte-for-byte today's contract: the same request fields, the same refusals, the same response projection. 2 is the canonical vocabulary. Any other value is a 400 with code: "VALIDATION_ERROR".
Today it decides one thing: the spelling of an evaluator's policy role. Vocabulary 1 accepts and returns gating; vocabulary 2 accepts both spellings and returns the canonical required. Sending required without the header is a 400, deliberately — vocabulary 1 is not widened to meet vocabulary 2 half way, because a boundary that accepts a spelling it does not announce is one two implementations can disagree about.
A response that varies by vocabulary sends Vary: x-mcpjam-eval-vocabulary.
Response
The ordered step results as a page envelope.