Grade an eval run with LLM as Judge
Spends. Runs the goal-completion judge over the finished run, scoring each case's final answer against its expected output.
202: scheduled, not done. Read the grades from the run detail's judges.goalCompletion rather than re-requesting — a second POST only spends again.
A run's grading config is pinned when the run is created, so turning the judge on for the suite does not reach an already-recorded run: enable: true is what grades one, and it changes nothing beyond that run. Omitting model and threshold clears any override a previous request left on the run.
Request body
Response
Scheduled.
Changes
Changed in 3 of the 122 revisions of this API.3
- ○
added the new optional
headerrequest parameterx-mcpjam-eval-vocabularyto all path's operationsnew-optional-request-default-parameter-to-existing-path
- ○
- ○
added the non-success response with the status
response-non-success-status-added
- ○
- ○
endpoint added
endpoint-added
- ○