---
title: "Get run status"
method: GET
path: "/projects/{projectId}/eval-runs/{runId}"
tags: ["Eval runs"]
---

# Get run status

`GET /projects/{projectId}/eval-runs/{runId}`

Run status, result, and summary. Poll until `status` is terminal (`completed`, `failed`, or `cancelled`).

## Path parameters

- `projectId` string, required
- `runId` string, required

## Response `200`

The run.

- EvalRun
  - `id` string, required
  - `suiteId` string, required
  - `runNumber` integer, nullable
  - `status` 'pending' | 'running' | 'completed' | 'failed' | 'cancelled', required — Poll until terminal: `completed`, `failed`, or `cancelled`.
  - `result` 'passed' | 'failed' | 'inconclusive' | 'null', nullable — Verdict once terminal. `inconclusive` exists only under `verdictPolicyVersion: 2` and is NOT a failure: the run did not measure the server well enough to say (too few gradeable trials, too many evaluator errors), so a gate that folds it into `failed` reports a defect the run never observed. Read `verdictSummary.reasons` for the check that withheld the verdict.
  - `summary` object, nullable
    - `total` integer
    - `passed` integer
    - `failed` integer
    - `passRate` number
  - `source` 'ui' | 'api' | 'sdk', required — Run origin. API-created runs are `api`.
  - `notes` string, nullable
  - `createdAt` number, required — Epoch milliseconds.
  - `completedAt` number, nullable — Epoch milliseconds, `null` until terminal.
  - `scoreIntegrity` 'valid' | 'invalid' | 'null', nullable — Whether the run's score evidence verified at ingest. TRI-STATE, and the third state matters: `valid` means the backend checked and definitions and results agree; `invalid` means they do not; `null` (or absent) means NO VERDICT was produced, on a deployment that predates integrity checking. A score gate must treat `null` exactly like `invalid` — absent evidence is not valid evidence.
  - `verdictPolicyVersion` 2 — The verdict policy this run was decided under, frozen at run start. ABSENT means legacy percent-threshold grading — `result` cannot then be `inconclusive` and there is no `verdictSummary`. A caller gating on fractions or on validity must read this FIRST rather than assume a missing summary means a clean run.
  - `verdictSummary` object — How a policy-2 verdict was reached: the resolved validity policy, the measured completion and evaluator-error rates with their denominators and exclusions, the per-case and per-execution-variant aggregates, and the exact reasons. Absent when the run is legacy, or when the stored summary failed contract validation at the boundary — a partially-valid decision is never published, because a gate cannot tell a missing field from a satisfied check.
  - `verdictPolicyIntegrityError` string — Why a policy-2 run could not be decided from its own evidence (a missing or malformed policy snapshot, mixed evaluator configs). Accompanies an `inconclusive` result; it is never a task failure.
  - `environment` EvalRunEnvironment, nullable — The project environment a run is pinned to, at the revision resolved when it launched. `null` for a legacy run that used the suite's saved server selection.
    - `id` string, required
    - `name` string, nullable
    - `revision` integer, nullable — The environment revision the run executed against.
  - `runGroupId` string — Shared by every per-target run from the same fan-out launch. Absent on a single-target launch and on rows created before run groups.
  - `effectiveModelId` string — Model the run actually executed with. Absent on pre-attribution rows.
  - `modelSource` 'client_default' | 'override' — `client_default` inherited the host model; `override` used the environment's `modelId`.
  - `executionEngine` string — Which engine executed the run: `emulated` (the platform's own turn loop) or `harness:<id>` (a real agent runtime such as Claude Code). ABSENT means the run recorded no engine — a run created before the platform attributed one. Treat that as UNKNOWN, never as `emulated`: those are different claims, and the runs whose engine was never recorded are exactly the ones a reader must not vouch for.
  - `insights` InsightsEnvelope — The common insights envelope, shared by eval runs, swarm waves and user-testing windows. One shape for three producers, so a caller writes the reading code once. An ABSENT envelope and `status: "not_available"` mean the same thing and both are normal: the field is an enrichment, and a caller who may not have it gets the resource without it rather than an error.
    - `schemaVersion` 1, required
    - `scope` InsightScope, required — What this envelope is about. The extra fields depend on `kind`.
      - `kind` 'eval_run' | 'swarm_wave' | 'user_testing_window', required
      - `id` string, required
      - `runId` string — `swarm_wave` only.
      - `scenarioId` string — `user_testing_window` only.
      - `windowStartAt` integer — `user_testing_window` only.
      - `windowEndAt` integer — `user_testing_window` only.
    - `status` 'not_available' | 'not_requested' | 'pending' | 'completed' | 'failed', required — `not_available` means this deployment cannot produce insights at all — treat an ABSENT envelope the same way. `not_requested` means nobody has asked. `pending` means one is running: poll, do not re-request.
    - `reasonCode` string, nullable, required
    - `retryable` boolean, required — Whether asking again could produce a different answer. False on a `failed` envelope means the input, not the attempt, was the problem.
    - `error` object, nullable, required
      - `code` string, required
      - `message` string, required
    - `generatedAt` integer, nullable, required
    - `updatedAt` integer, nullable, required
    - `summary` string, nullable, required
    - `coverage` object, required — READ THIS BEFORE QUOTING ANY FINDING. `truncated` and `lowConfidence` are the difference between "this happens" and "this happened in the part we looked at".
      - `unit` 'iterations' | 'sessions', required
      - `analyzed` integer, required
      - `total` integer, required
      - `gradedCount` integer
      - `feedbackCount` integer
      - `truncated` boolean, required — The analysis saw `analyzed` of `total`, not all of it.
      - `lowConfidence` boolean, required — Too little was analyzed to generalize. Findings still stand as observations of what WAS seen.
    - `findings` ActionableFinding[], required
      - `id` string, required — Stable remediation id (`rf_<16 hex>`). Survives dynamic error values, so the same problem keeps the same id across runs — dismiss it once and it stays dismissed.
      - `signalFingerprint` string, required — The registry signal this derives from. Several findings can share one.
      - `title` string, required
      - `category` 'unknown' | 'tool_contract' | 'tool_runtime' | 'capability_gap' | 'workflow' | 'agent_behavior' | 'test_design' | 'environment', required
      - `attribution` 'unknown' | 'server_contract' | 'server_runtime' | 'server_capability' | 'agent_or_prompt' | 'test_design' | 'environment', required — WHOSE problem this is. `server_*` points at the MCP server; `agent_or_prompt` and `test_design` point back at the caller.
      - `actionTarget` 'investigate' | 'mcp_server' | 'agent_configuration' | 'eval_case' | 'environment', required — What you would change to fix it.
      - `actionability` 'informational' | 'investigate' | 'ready', required — `ready` means the finding names a specific target and change. `investigate` means it does not yet. `informational` means there is nothing to do.
      - `severity` 'info' | 'low' | 'medium' | 'high', required
      - `confidence` 'low' | 'medium' | 'high', required
      - `observed` string, required — DETERMINISTIC observation — counts and identities, never model prose. This is the part you can verify yourself.
      - `rootCause` string
      - `recommendation` string, required
      - `acceptanceCriteria` string[], required — How you would know the fix worked.
      - `affected` object, required — How much of the analyzed population hit this. Read it as a ratio — `1/40` and `38/40` are different problems.
        - `count` integer, required
        - `total` integer, required
        - `unit` 'iterations' | 'sessions', required
      - `patternSlug` string
      - `target` object — Present only when a server (and, for tool surfaces, a tool) resolved against the pinned snapshot. Required for `mcp_server` / `ready`.
        - `serverId` string, required
        - `toolName` string
        - `surface` 'description' | 'input_schema' | 'output_schema' | 'handler' | 'server_instructions' | 'capability', required
        - `fieldPath` string
        - `snapshotHash` string, required — The pinned snapshot the target resolved against, so a finding cannot silently re-point at a definition that changed after it was written.
        - `currentDefinition` object
          - `description` string
          - `inputSchemaJson` string
          - `outputSchemaJson` string
          - `truncated` boolean, required
      - `evidence` ActionableFindingEvidence[], required
        - `sessionId` string
        - `iterationId` string
        - `kind` 'tool_error' | 'transcript' | 'feedback' | 'judge' | 'contrast', required
        - `excerpt` string, required — Scrubbed and clipped at the producer. Never a full transcript.
        - `toolName` string
        - `errorCode` string
    - `runHealth` object — Swarm only. Launch outcomes never appear as findings — a run that could not start is an operational fact, not something the server under test did.
      - `targets` object[], required
        - `subjectKind` 'environment' | 'host', required
        - `subjectId` string, required
        - `subjectLabel` string, required
        - `attempted` integer, required
        - `succeeded` integer, required
        - `failed` integer, required
        - `rateLimited` integer, required
    - `truncation` object, required — What this RESPONSE dropped to stay a sane size, as distinct from what the ANALYSIS did not look at (`coverage`).
      - `truncated` boolean, required
      - `omittedFindings` integer, required
      - `omittedEvidence` integer, required
      - `contractTruncated` boolean, required
  - `judges` EvalRunJudges — Advisory graders on a run, keyed by judge. An envelope rather than a bare `judge` field because goal completion is one of several; a future judge is a new key here, not a reshaped response.
    - `goalCompletion` EvalRunGoalCompletionJudge — Goal completion: grades each case's final answer against its expected output.
      - `status` 'pending' | 'completed' | 'failed' | 'null', nullable, required — `null` means this judge was NEVER requested for the run — a different answer from "requested and graded nothing". Poll rather than re-requesting while `pending`.
      - `errorCode` string, nullable, required — Machine-readable failure reason, set alongside `status: "failed"`.
      - `summary` string, nullable, required
      - `generatedAt` number, nullable, required — Epoch milliseconds.
      - `modelUsed` string, nullable, required
      - `threshold` number, nullable, required — The threshold these results were scored against (`passed = score >= threshold`).
      - `cases` EvalRunGoalCompletionCase[], required — Per-case grades. EMPTY unless `status` is `completed` — a pending or failed judge carries no cases, and `status` is what says which.
        - `caseKey` string, required — The stable AUTHORED-case identity, as persisted. NOT a case row id — do not join it against the ids the per-case routes take.
        - `iterationId` string — The iteration this case graded. The join key between a judge case and the iterations the run returns; `caseKey` is a storage key and joins to nothing. Absent on judge results written before it was persisted.
        - `score` number, nullable, required
        - `passed` boolean, required
        - `reason` string, nullable, required
        - `rubricHits` string[], required — Rubric criteria the answer satisfied.
    - `groundedness` EvalRunGroundednessJudge — Groundedness: grades whether each answer is SUPPORTED by its tool trajectory.
      - `status` 'pending' | 'completed' | 'failed' | 'null', nullable, required — `null` means this judge was NEVER requested for the run — a different answer from "requested and graded nothing". Poll rather than re-requesting while `pending`.
      - `errorCode` string, nullable, required — Machine-readable failure reason, set alongside `status: "failed"`.
      - `summary` string, nullable, required
      - `generatedAt` number, nullable, required — Epoch milliseconds.
      - `modelUsed` string, nullable, required
      - `threshold` number, nullable, required — The threshold these results were scored against (`passed = score >= threshold`).
      - `cases` EvalRunGroundednessCase[], required — Per-case grades. EMPTY unless `status` is `completed`.
        - `caseKey` string, required — The stable AUTHORED-case identity, as persisted. NOT a case row id — do not join it against the ids the per-case routes take.
        - `iterationId` string — The iteration this case graded. The join key between a judge case and the iterations the run returns; `caseKey` is a storage key and joins to nothing. Absent on judge results written before it was persisted.
        - `score` number, nullable, required
        - `passed` boolean, required
        - `reason` string, nullable, required
        - `unsupportedClaims` string[], required — Claims the tool trajectory does not support.
  - `gateWaiver` GateWaiver — An audited, time-boxed override of an eval run's release gate. A waiver never changes the run's own `result` — the run keeps its verdict and every surface that honors the waiver says so out loud, which is what makes "no silent waiver" checkable rather than promised.
    - `id` string, required
    - `suiteId` string, required
    - `runId` string, nullable, required — The run this waiver covers. Suite-wide waivers are not honored.
    - `reason` string, required — Why the gate was overridden, as the granter wrote it. Stored UNREDACTED and readable by anyone who can see the suite, for as long as the suite exists — never put secrets, tokens, or customer data in it.
    - `expiresAt` integer, required — Epoch ms. Always in the future when granted, and capped at 30 days out — there is no permanent waiver.
    - `createdAt` integer, required
    - `createdBy` string, required
    - `createdByEmail` string, nullable, required — `null`, never absent, when it cannot be resolved — a deleted user must not make a waiver look authorless.
    - `revokedAt` integer, nullable, required
    - `revokedBy` string, nullable, required
    - `active` boolean, required — Whether it is in force right now — neither revoked nor expired. Computed at read time; a client that must not honor a lapsed waiver should re-derive it from `expiresAt` rather than trust it.
    - `policySnapshot` object, nullable, required — WHAT was overridden, captured at waive time so a later edit to the suite cannot rewrite the record. `null` for a run decided by the v2 verdict policy, whose identity is recorded on the audit event instead — this shape cannot hold it, and filling it in would be a false record rather than an incomplete one.
      - `minimumPassRate` number, required
  - `importEligibility` ImportEligibility — Whether a run's imported cases carry evidence a gate may rely on. Computed by the platform from the run's OWN frozen snapshot, never from the suite's current cases — recomputing from those would let a later edit change what a finished run is allowed to prove. `legacy` means the run contains no imported cases at all (every native run, forever) and is gateable unchanged; `eligible` means every imported case carries a valid frozen decision; `incomplete` means the evidence cannot be trusted, which makes the run NOT GATEABLE and is not a test verdict. The whole field is ABSENT on deployments that predate import eligibility — a different fact from `legacy`, and one a gate must read as "no opinion, behave as before".
    - `status` 'legacy' | 'eligible' | 'incomplete', required
    - `gateable` boolean, required
    - `importedCaseCount` integer, required
    - `claimedExactCaseIds` string[], required
    - `approvedApproximationCaseIds` string[], required
    - `approvedApproximationReceipts` ImportApprovalReceipt[], required
      - `testCaseId` string, required
      - `caseKey` string
      - `sourceCaseKey` string
      - `approvedBy` string, required — Server-derived approver id.
      - `approvedAt` integer, required — Server-derived epoch milliseconds.
      - `reason` string, required
    - `issues` ImportEligibilityIssue[], required
      - `code` string, required — Stable machine-readable code.
      - `testCaseId` string
      - `caseKey` string
      - `toolName` string

## Other responses

- `401` — Missing, invalid, revoked, or orphaned key (`UNAUTHORIZED`) — or the **target MCP server** needs an OAuth grant (`OAUTH_REQUIRED`), which is a property of the server, not your key.
- `403` — Key is valid but not allowed to do this.
- `404` — Unknown project, server, or resource.
- `429` — Per-key rate limit exceeded (60 requests/minute sustained, bursts up to 10). Honor `Retry-After` and back off with jitter.
- `500` — Something failed on MCPJam's side.

## Changes

> 71 revisions in range; 28 could not be searched.

- **2026-08-19** `ece7d99ceaf3` — 1 info
  - added the optional property `executionEngine` to the response with the `200` status
- **2026-08-19** `fac4c14468c7` — 1 info
  - added the optional property `judges` to the response with the `200` status
- **2026-08-15** `d3adfe49fbbf` — 3 breaking, 2 warning, 1 info
  - added `#/components/schemas/EvalRunEnvironment, subschema #2` to the `environment` response property `oneOf` list for the response status `200`
  - the `environment` response's property type/format changed from `object, null`/`` to ``/`` for status `200`
  - removed the required property `environment/id` from the response with the `200` status
  - removed the optional property `environment/name` from the response with the `200` status
  - …2 more
- **2026-08-15** `2b4d84394b27` — 1 info
  - added the optional property `scoreIntegrity` to the response with the `200` status
- **2026-08-04** `898e20711854` — 1 info
  - added the optional property `environment` to the response with the `200` status

[Full history](https://skmtc.dev/mcpjam/apis/mcpjam-api/changes/projects/:projectId/eval-runs/:runId/get.md)

---

[API](https://skmtc.dev/mcpjam/apis/mcpjam-api.md) · [All operations](https://skmtc.dev/mcpjam/apis/mcpjam-api/llms.txt) · [OpenAPI document](https://skmtc-service-production.skmtc.workers.dev/v1/apis/mcpjam/mcpjam-api/revisions/c628a0e95b06/schema)
