---
title: "Create an eval run (async)"
method: POST
path: "/projects/{projectId}/eval-runs"
tags: ["Eval runs"]
---

# Create an eval run (async)

`POST /projects/{projectId}/eval-runs`

Creates a suite run from an existing `suiteId` (rerun) and/or inline `tests`, then **detaches execution and responds `202` immediately** with the `runId`. Validation and quota errors surface on this request; poll `GET /eval-runs/{runId}` for progress. The run appears live in the hosted UI Runs tab, tagged `source: "api"`.

A bare `suiteId` with no inline tests reruns the suite as configured. Per-organization concurrency is capped (default 2 concurrent runs); exceeding it returns `429` with `details.reason: "CONCURRENT_RUN_LIMIT"`.

## Path parameters

- `projectId` string, required

## Request body

- union — Two valid shapes: `suiteId` (rerun an existing suite, optionally upserting inline `tests` into it), or `suiteName` + `tests` + `serverIds` (create a new suite and run it). Inline `tests` alone — without a `suiteId` or a `suiteName` — are rejected with `VALIDATION_ERROR`.
  - object
    - `suiteId` string, required — Existing suite to rerun. A bare `suiteId` with no `tests` reruns the suite exactly as configured.
    - `suiteName` string — Name for a new suite. Required (non-empty) when no `suiteId` is given.
    - `suiteDescription` string
    - `tests` EvalTestCase[] — Inline test cases to upsert into the suite before running.
      - `title` string, required
      - `steps` EvalTestStep[], required — Ordered test steps. The first `prompt` step is the case query; `toolCalledWith` asserts are the expected tool calls; a single model-free `toolCall` step is a render-check.
        - `id` string, required
        - `kind` 'prompt' | 'toolCall' | 'interact' | 'assert', required
        - `prompt` string — User message (`kind: prompt`).
        - `serverName` string — Server that owns the tool (`kind: toolCall`).
        - `toolName` string — Tool name (`kind: toolCall` / `interact`).
        - `arguments` object — Tool-call arguments (`kind: toolCall`).
        - `action` object — Widget action (`kind: interact`).
        - `assertion` object — Predicate or widget assertion (`kind: assert`).
      - `runs` integer, required — Iterations to execute for this case.
      - `model` string, required — Model ID. Hosted-catalog ids use `provider/name` form (e.g. `anthropic/claude-haiku-4.5`) and run on org credits; provider-native ids require a matching `modelApiKeys` entry (BYOK). Unknown models are rejected with VALIDATION_ERROR at create time.
      - `provider` string, required — Model provider, e.g. `anthropic`, `openai`.
      - `isNegativeTest` boolean — When `true`, the case passes if NO tools are called.
      - `expectedOutput` string
      - `advancedConfig` object — Optional `system`, `temperature`, `toolChoice` overrides.
    - `serverIds` string[] — Servers (by ID) the run connects to. Required when creating a new suite; optional on reruns — when omitted, the run connects the suite's saved server selection (the set its snapshot references). A rerun of a suite with no saved selection is rejected with `VALIDATION_ERROR` (`details.reason: "NO_SAVED_SERVER_SELECTION"`).
    - `modelApiKeys` object — Optional per-provider model API keys (e.g. `{ "anthropic": "sk-ant-…" }`). Falls back to your organization's configured providers when omitted.
    - `notes` string
    - `passCriteria` object
      - `minimumPassRate` number
    - `iterationOverride` integer — Override the per-case `runs` count for this run only.
  - object
    - `tests` EvalTestCase[], required — Inline test cases to upsert into the suite before running.
      - `title` string, required
      - `steps` EvalTestStep[], required — Ordered test steps. The first `prompt` step is the case query; `toolCalledWith` asserts are the expected tool calls; a single model-free `toolCall` step is a render-check.
        - `id` string, required
        - `kind` 'prompt' | 'toolCall' | 'interact' | 'assert', required
        - `prompt` string — User message (`kind: prompt`).
        - `serverName` string — Server that owns the tool (`kind: toolCall`).
        - `toolName` string — Tool name (`kind: toolCall` / `interact`).
        - `arguments` object — Tool-call arguments (`kind: toolCall`).
        - `action` object — Widget action (`kind: interact`).
        - `assertion` object — Predicate or widget assertion (`kind: assert`).
      - `runs` integer, required — Iterations to execute for this case.
      - `model` string, required — Model ID. Hosted-catalog ids use `provider/name` form (e.g. `anthropic/claude-haiku-4.5`) and run on org credits; provider-native ids require a matching `modelApiKeys` entry (BYOK). Unknown models are rejected with VALIDATION_ERROR at create time.
      - `provider` string, required — Model provider, e.g. `anthropic`, `openai`.
      - `isNegativeTest` boolean — When `true`, the case passes if NO tools are called.
      - `expectedOutput` string
      - `advancedConfig` object — Optional `system`, `temperature`, `toolChoice` overrides.
    - `suiteId` string — Existing suite to rerun. A bare `suiteId` with no `tests` reruns the suite exactly as configured.
    - `suiteName` string, required — Name for a new suite. Required (non-empty) when no `suiteId` is given.
    - `suiteDescription` string
    - `serverIds` string[], required — Servers (by ID) the run connects to. Required when creating a new suite; optional on reruns — when omitted, the run connects the suite's saved server selection (the set its snapshot references). A rerun of a suite with no saved selection is rejected with `VALIDATION_ERROR` (`details.reason: "NO_SAVED_SERVER_SELECTION"`).
    - `modelApiKeys` object — Optional per-provider model API keys (e.g. `{ "anthropic": "sk-ant-…" }`). Falls back to your organization's configured providers when omitted.
    - `notes` string
    - `passCriteria` object
      - `minimumPassRate` number
    - `iterationOverride` integer — Override the per-case `runs` count for this run only.

## Response `202`

Run created; execution continues in the background.

- EvalRunCreated
  - `runId` string, required
  - `suiteId` string, required
  - `status` 'running', required
  - `servers` object[] — The servers the run connects to — explicit or derived from the suite's saved selection. `name` is present when known (always, on the derived path).
    - `id` string, required
    - `name` string
  - `caseUpsert` object, required — Per-case upsert outcomes for inline `tests`. Partial failures don't abort the run.
    - `committed` object[]
      - `id` string
      - `name` string
    - `failed` object[]
      - `id` string
      - `name` string
      - `error` string

## Other responses

- `400` — Malformed body or parameters.
- `401` — Missing, invalid, revoked, or orphaned key (`UNAUTHORIZED`) — or the **target MCP server** needs an OAuth grant (`OAUTH_REQUIRED`), which is a property of the server, not your key.
- `403` — Key is valid but not allowed to do this.
- `404` — Unknown project, server, or resource.
- `429` — Per-key rate limit exceeded (60 requests/minute sustained, bursts up to 10). Honor `Retry-After` and back off with jitter.
- `500` — Something failed on MCPJam's side.
- `502` — Could not connect to the target MCP server.
- `504` — The target MCP server connected but didn't respond in time.

## Changes

- **2026-06-26** `f311b5e01264` — 1 breaking, 2 warning
  - added the new required request property `tests/items/steps`
  - removed the request property `tests/items/expectedToolCalls`
  - removed the request property `tests/items/query`
- **2026-06-11** `08ae18c9c1ea` — 2 info
  - the request property `serverIds` became optional
  - added the optional property `servers` to the response with the `202` status
- **2026-06-11** `ad9bf0de7ed4` — 1 info
  - endpoint added

[Change history](https://skmtc.dev/mcpjam/apis/mcpjam-api/changes/projects/:projectId/eval-runs/post.md)

---

[API](https://skmtc.dev/mcpjam/apis/mcpjam-api.md) · [All operations](https://skmtc.dev/mcpjam/apis/mcpjam-api/llms.txt) · [OpenAPI document](https://skmtc-service-production.skmtc.workers.dev/v1/apis/mcpjam/mcpjam-api/revisions/302cc2d08c56/schema)
