Public Create Eval Run
Record a terminal eval run (with optional per-case results) for a task.
Path parameters
Headers
Request body
Model name snapshot, e.g. wilson-v3-step200
Training dataset snapshot
Eval dataset snapshot, e.g. apex-curated-eval-subset@2026-07-10
Mutation/question variant within the task, e.g. v2-polar-positive
Mark this run as a reference instead of a scored datapoint, e.g. "random", "SOTA", "gpt-5 baseline". Reference runs are excluded from every score rollup and the group chart draws them as dotted horizontal lines captioned with this label. Empty (default) = an ordinary scored run.
The {vars} question template that defines the variant
How the run was graded / answer logic
Name of the headline metric overall_score normalizes
Where the pipeline stored the raw outputs
Optional original record time, preserved by export→replay syncs. The timeline falls back to it when finished_at is unset; omitted = now.
Which pipeline pushed this run
Response
Successful Response
Non-empty for reference runs, which are excluded from the score rollups