Get run status
Run status, result, and summary. Poll until status is terminal (completed, failed, or cancelled).
Path parameters
ID of the hosted project that contains the server.
Eval run ID, as returned by POST /eval-runs.
Response
The run.
Poll until terminal: completed, failed, or cancelled.
Verdict once terminal. inconclusive exists only under verdictPolicyVersion: 2 and is NOT a failure: the run did not measure the server well enough to say (too few gradeable trials, too many evaluator errors), so a gate that folds it into failed reports a defect the run never observed. Read verdictSummary.reasons for the check that withheld the verdict.
Run origin. API-created runs are api.
Epoch milliseconds.
Epoch milliseconds, null until terminal.
Whether the run's score evidence verified at ingest. TRI-STATE, and the third state matters: valid means the backend checked and definitions and results agree; invalid means they do not; null (or absent) means NO VERDICT was produced, on a deployment that predates integrity checking. A score gate must treat null exactly like invalid — absent evidence is not valid evidence.
The verdict policy this run was decided under, frozen at run start. ABSENT means legacy percent-threshold grading — result cannot then be inconclusive and there is no verdictSummary. A caller gating on fractions or on validity must read this FIRST rather than assume a missing summary means a clean run.
How a policy-2 verdict was reached: the resolved validity policy, the measured completion and evaluator-error rates with their denominators and exclusions, the per-case and per-execution-variant aggregates, and the exact reasons. Absent when the run is legacy, or when the stored summary failed contract validation at the boundary — a partially-valid decision is never published, because a gate cannot tell a missing field from a satisfied check.
Why a policy-2 run could not be decided from its own evidence (a missing or malformed policy snapshot, mixed evaluator configs). Accompanies an inconclusive result; it is never a task failure.
Shared by every per-target run from the same fan-out launch. Absent on a single-target launch and on rows created before run groups.
Model the run actually executed with. Absent on pre-attribution rows.
client_default inherited the host model; override used the environment's modelId.
Which engine executed the run: emulated (the platform's own turn loop) or harness:<id> (a real agent runtime such as Claude Code). ABSENT means the run recorded no engine — a run created before the platform attributed one. Treat that as UNKNOWN, never as emulated: those are different claims, and the runs whose engine was never recorded are exactly the ones a reader must not vouch for.