Eval runs

Get a run's decision summary

The canonical, versioned run decision contract: the verdict and where it came from, the counts with the population they count, the run's own verdict decision when it has one, and one page of per-trial diagnostics carrying the user-value chain, the first failed stage, the failure category, evidence scoped to that stage, and one next action.

ADDITIVE and composed: it is the same two reads a caller could make by hand (GET …/eval-runs/{runId} and GET …/iterations), assembled once here so every client shares one reading of a run instead of inventing its own. cursor and limit page the DIAGNOSTICS using the same cursors the iterations endpoint issues, and diagnostics.complete says honestly whether the page you got is the run's whole non-passing set.

get/projects/{projectId}/eval-runs/{runId}/decision-summary

Path parameters

projectIdstring required

ID of the hosted project that contains the server.

runIdstring required

Eval run ID, as returned by POST /eval-runs.

Query parameters

limitinteger

Iterations examined per page, 1–200. Defaults to 50.

cursorstring

Opaque cursor from a previous response's diagnostics.nextCursor. A page fetched with one is never reported as complete.

Headers

x-mcpjam-eval-vocabulary'1' | '2'

Which vocabulary this request and its response speak. Absent means 1, which is byte-for-byte today's contract: the same request fields, the same refusals, the same response projection. 2 is the canonical vocabulary. Any other value is a 400 with code: "VALIDATION_ERROR".

Today it decides one thing: the spelling of an evaluator's policy role. Vocabulary 1 accepts and returns gating; vocabulary 2 accepts both spellings and returns the canonical required. Sending required without the header is a 400, deliberately — vocabulary 1 is not widened to meet vocabulary 2 half way, because a boundary that accepts a spelling it does not announce is one two implementations can disagree about.

A response that varies by vocabulary sends Vary: x-mcpjam-eval-vocabulary.

Response

The run's decision summary.

schemaVersion1 required
runIdstring required
runStatusstring required

The run's lifecycle status, verbatim. Not a verdict.

verdict'passed' | 'failed' | 'inconclusive' | 'notEstablished' required
verdictSource'policyV2' | 'legacy' | 'none' required

Which criterion decided the run, which is also what its counts are in. policyV2 — PER-CASE GRADING: the run's own decision, carried on decision; verdict is its verdict and the counts are case-execution variants. legacy — the SUITE-WIDE ACCURACY THRESHOLD: one percentage over the whole run, so there is no decision object and any counts are trials. The two are not one criterion in two units, so never convert a rate across this line. none — no verdict; verdict is notEstablished and undecided says why.

decisionobject

The run's authoritative v2 verdict decision, copied verbatim after validation — the same shape as EvalRun.verdictSummary, published at https://mcpjam.com/schemas/eval-verdict-policy/v2.json. Present exactly when verdictSource is policyV2.

Changes