Eval runs

Preview deterministic draft evaluators against stored evidence

Requires suite configuration permission and a terminal run. Reserves a 60-second deterministic-preview cooldown, independent of judge previews. Supply the returned continuation object with the unchanged draft to read the next bounded population. Reads at most 100 iterations within a 30-second operation deadline; partial evidence is explicitly ungradable. Runs no model calls and never overwrites original verdicts. Draft assertions explicitly replace, extend, or inherit the frozen rules.

post/projects/{projectId}/eval-runs/{runId}/backtest

Path parameters

projectIdstring required

ID of the hosted project that contains the server.

runIdstring required

Eval run ID, as returned by POST /eval-runs.

Headers

x-mcpjam-eval-vocabulary'1' | '2'

Which vocabulary this request and its response speak. Absent means 1, which is byte-for-byte today's contract: the same request fields, the same refusals, the same response projection. 2 is the canonical vocabulary. Any other value is a 400 with code: "VALIDATION_ERROR".

Today it decides one thing: the spelling of an evaluator's policy role. Vocabulary 1 accepts and returns gating; vocabulary 2 accepts both spellings and returns the canonical required. Sending required without the header is a 400, deliberately — vocabulary 1 is not widened to meet vocabulary 2 half way, because a boundary that accepts a spelling it does not announce is one two implementations can disagree about.

A response that varies by vocabulary sends Vary: x-mcpjam-eval-vocabulary.

Request body

matchOptionsobject nullable

Optional deterministic tool matching options; null omits matching.

Response

Draft evidence comparison.

{"stackTrail":"components:schemas:EvalBacktestReport:properties:schemaVersion","oasType":"schema","type":"unknown"}
sourceRunIdstring required
sourceHashstring required
draftHashstring required
{"stackTrail":"components:schemas:EvalBacktestReport:properties:configRevision","oasType":"schema","type":"unknown"}
completeboolean required
continuationAvailableboolean required
differencesobject[] required
modelUse'none' required

Changes

Changed in 2 of the 122 revisions of this API.2