Eval runs

Preview draft judge grading on one recorded iteration

Uses model budget and complete recorded evidence without changing the run verdict. Supply rubric.instructions, optional structured criteria, or null for objective-only grading. Continue with the unchanged rubric and returned cursor, sourceHash and reservationId. Completed page retries return cached results. Error cases have a null draft measurement.

post/projects/{projectId}/eval-runs/{runId}/judge/backtest

Path parameters

projectIdstring required

ID of the hosted project that contains the server.

runIdstring required

Eval run ID, as returned by POST /eval-runs.

Headers

x-mcpjam-eval-vocabulary'1' | '2'

Which vocabulary this request and its response speak. Absent means 1, which is byte-for-byte today's contract: the same request fields, the same refusals, the same response projection. 2 is the canonical vocabulary. Any other value is a 400 with code: "VALIDATION_ERROR".

Today it decides one thing: the spelling of an evaluator's policy role. Vocabulary 1 accepts and returns gating; vocabulary 2 accepts both spellings and returns the canonical required. Sending required without the header is a 400, deliberately — vocabulary 1 is not widened to meet vocabulary 2 half way, because a boundary that accepts a spelling it does not announce is one two implementations can disagree about.

A response that varies by vocabulary sends Vary: x-mcpjam-eval-vocabulary.

Request body

Response

One cached or newly graded page. Errors are unscored; isDone indicates the end of the population.

okboolean
casesobject[]
cursorinteger
sourceHashstring
reservationIdstring
isDoneboolean

Changes

Changed in 1 of the 122 revisions of this API.1