Get run status
Run status, result, and summary. Poll until status is terminal (completed, failed, cancelled or timed_out). grading is not terminal — every trial has finished and the run is held for its gating judge, with result still pending — so keep polling.
Path parameters
ID of the hosted project that contains the server.
Eval run ID, as returned by POST /eval-runs.
Response
The run.
Poll until TERMINAL. The four terminal statuses are completed, failed, cancelled and timed_out. grading is NOT terminal: every trial has finished and the run is being held for its gating judge — up to 30 minutes — with result still pending. A poller that stops at grading reports a run with no verdict as though it had one.
Verdict once terminal. inconclusive exists only under verdictPolicyVersion: 2 and is NOT a failure: the run did not measure the server well enough to say (too few gradeable trials, too many evaluator errors), so a gate that folds it into failed reports a defect the run never observed. Read verdictSummary.reasons for the check that withheld the verdict.
Run origin, STAMPED BY THE SERVER and not settable by a caller. Every run created through this API is api — including one launched by the CLI, by a GitHub Actions job, or by an MCP agent, because from the server's side all three are API calls. The caller's own claim about which of those it is rides on launcher. schedule and github_check are written by the platform's scheduled-eval and PR-check workers; benchmark marks a Connector Bench matrix cell and is hidden from project run lists.
Epoch milliseconds.
Epoch milliseconds, null until terminal.
Whether the run's score evidence verified at ingest. TRI-STATE, and the third state matters: valid means the backend checked and definitions and results agree; invalid means they do not; null (or absent) means NO VERDICT was produced, on a deployment that predates integrity checking. A score gate must treat null exactly like invalid — absent evidence is not valid evidence.
The verdict policy this run was decided under, frozen at run start. ABSENT means legacy percent-threshold grading — result cannot then be inconclusive and there is no verdictSummary. A caller gating on fractions or on validity must read this FIRST rather than assume a missing summary means a clean run.
How a policy-2 verdict was reached: the resolved validity policy, the measured completion and evaluator-error rates with their denominators and exclusions, the per-case and per-execution-variant aggregates, and the exact reasons. Absent when the run is legacy, or when the stored summary failed contract validation at the boundary — a partially-valid decision is never published, because a gate cannot tell a missing field from a satisfied check.
Why a policy-2 run could not be decided from its own evidence (a missing or malformed policy snapshot, mixed evaluator configs). Accompanies an inconclusive result; it is never a task failure.
Shared by every per-target run from the same fan-out launch. Absent on a single-target launch and on rows created before run groups.
Model the run actually executed with. Absent on pre-attribution rows.
client_default inherited the host model; override used the environment's modelId; case used the sole model in the case snapshot.
Which engine executed the run: emulated (the platform's own turn loop) or harness:<id> (a real agent runtime such as Claude Code). ABSENT means the run recorded no engine — a run created before the platform attributed one. Treat that as UNKNOWN, never as emulated: those are different claims, and the runs whose engine was never recorded are exactly the ones a reader must not vouch for.
Bounded caller metadata; preserved as descriptive data, never authorization.
Bounded CI attribution and source revision metadata.
Case-scoped run evaluator observations with source and definition identity.