evals

Compare Runs

Side-by-side compare of two suite runs. Returns aggregate scorecards and per-member breakdown so the UI can diff regressed scenarios at a glance — the killer view for the cheap-model-test workflow.

get/v1/admin/eval/runs/{run_id}/compare

Path parameters

run_idstring required

Query parameters

againststring required

Response

Successful Response

{"stackTrail":"paths:/v1/admin/eval/runs/{run_id}/compare:get:responses:200:content:application/json:schema","oasType":"schema","type":"unknown"}

Changes

No recorded changes to this endpoint across all 1 revision of this API.