---
title: "Create a batch inference job"
method: POST
path: "/v1/batches"
tags: ["Batches"]
---

# Create a batch inference job

`POST /v1/batches`

Submit a batch from an uploaded input file (input_file_id) or inline request objects (requests); provide exactly one of the two. The batch completes within the 24h window at a discounted rate; poll it with get, then download results via the output and error file ids.

## Request body

- object
  - `completion_window` '24h', required — Time budget for the batch; only "24h" is supported.
  - `endpoint` '/v1/chat/completions' | '/v1/completions' | '/v1/embeddings', required — The API family every record in the batch calls.
  - `input_file_id` string — Id of an uploaded JSONL input file (from the file upload).
  - `metadata` object — Up to 16 key-value pairs echoed back on the batch object.
  - `model` string — Optional display hint shown on the batch immediately; validation settles the authoritative value from the input file.
  - `priority` 'standard' | 'expedited' — Scheduling tier. standard (the default) runs at the discounted batch rate. expedited is granted a freeing pool slot ahead of any waiting standard batch, grows the fleet to a tighter drain target, and is billed at the online rate with no batch discount; it is available to every organization unless an operator has disabled it for yours, in which case the request is rejected with 400. There is no completion deadline on either tier.
  - `requests` BatchRequestRecord[] — Inline request objects (max 50,000), each with a custom_id and the request body for the chosen endpoint. Alternative to input_file_id for small batches.
    - `body` object — The request body you would send to that endpoint online, for example a {model, messages} object for /v1/chat/completions. Streaming is not supported, so "stream": true fails that record. A chat record may also carry OpenRelay request extensions under the reserved "openrelay" key. A bare top-level tool_config is the retired spelling of that extension and fails the record with code tool_config_moved.
      - `openrelay` BatchOpenRelayExtensions — The reserved "openrelay" object on a batch record body: the one top-level key OpenRelay request extensions live under, mirroring the "openrelay" response namespace that carries tool_rounds. Additive, so new keys may appear without a breaking change. The worker strips the keys it owns from the body before every model invocation.
        - `tool_config` BatchToolConfig — The OpenRelay tool-calling extension, carried at body.openrelay.tool_config on a /v1/chat/completions record: where to execute the model's tool calls between turns. When a turn ends with finish_reason "tool_calls", the batch worker POSTs the calls to this endpoint, appends the assistant and tool messages, and re-invokes the model, up to max_rounds times. The field is stripped from the body before every model invocation and never appears in result files. A malformed value fails that one record with code invalid_tool_config.
          - `authorization` string — Sent verbatim as the Authorization header on every executor call.
          - `context` object — Opaque JSON forwarded to the executor on every call.
          - `max_rounds` integer — Maximum tool round-trips for this record. 0 or omitted takes the default of 8; values above 16 are clamped to 16. A record can therefore cost up to max_rounds + 1 model invocations.
          - `timeout_ms` integer — Per-executor-call timeout in milliseconds. 0 or omitted takes the default of 30000; values above 120000 are clamped to 120000.
          - `type` 'http' | 'mcp' — Executor protocol. "http" is the plain POST contract ({custom_id, context, tool_calls} in, {results:[{tool_call_id, content}]} out); "mcp" is a Model Context Protocol server speaking Streamable HTTP. Defaults to "http".
          - `url` string, uri, required — Your tool-executor endpoint. HTTPS only, because the request carries your authorization value. Endpoints resolving to internal or private addresses are rejected.
        - `workflow` BatchWorkflowConfig — Callback-driven continuation for one batch record. OpenRelay posts an event to your callback, runs the inference requests the callback returns inside this record, posts their results back, and repeats until the callback answers completed or failed. The requests returned in one step run concurrently; steps run in order, and each results event carries every result of the previous step in request order. How many requests run at once is set per model by OpenRelay and shared by every record in the batch, not set per record. A record is bounded by work (max_steps, max_requests, max_tokens) and by the batch's 24 hour completion window, never by a per-record clock. Every request that returns a successful response is billed at the batch rate when it finishes, even if the record later fails, hits a limit, or expires. Budget failures use the codes workflow_step_limit, workflow_request_limit, and workflow_token_limit; a callback that fails all four delivery attempts of one event fails the record with workflow_callback_unreachable, and the rest of the batch keeps running. Dependent requests stay inside this batch and never create another batch. Available only to organizations an operator has enabled for workflow records; otherwise the record fails validation with code workflow_not_enabled and the rest of the batch runs unaffected. Records start in input order; the batch-level priority field selects the scheduling tier for all of them.
          - `authorization` string, required — Sent verbatim as the Authorization header on callback requests.
          - `callback_timeout_ms` integer — Timeout for one callback HTTP attempt, in milliseconds. Omitted takes the default of 120000; values outside 1000 to 600000 fail validation with invalid_workflow. Each event is delivered up to four times, and a timeout, connection error, 429, or 5xx is retried; when all four attempts fail the record fails with workflow_callback_unreachable. It bounds a single callback call, not how long the record runs.
          - `context` object — Opaque JSON echoed to the callback on every event.
          - `max_requests` integer — Maximum inference requests this record may run across all steps. Omitted takes the default of 512; values outside 1 to 4096 fail validation with invalid_workflow. Checked before a step starts: if the step's requests would take the record past the limit, none of them run and the record fails with workflow_request_limit. Every request counts, including one that fails validation or returns an error.
          - `max_steps` integer — Maximum callback transitions for this record: each continue or wait answer is one step. Omitted takes the default of 128; values outside 1 to 128 fail validation with invalid_workflow. Checked after each callback answer, so a callback may still answer completed or failed on its last step. Going past it fails the record with workflow_step_limit.
          - `max_tokens` integer — Maximum prompt plus output tokens across all of this record's requests. This is a record budget, separate from the max_tokens a request body sets for one completion. Omitted takes the default of 10000000; values outside 1 to 100000000 fail validation with invalid_workflow. Checked as each request finishes: once the total passes the limit, the step's requests still running are cancelled and not billed, and the record fails with workflow_token_limit. The total can therefore exceed the limit by the requests that finished in that step.
          - `timeout_ms` integer — Deprecated alias of callback_timeout_ms with the same meaning: the timeout for one callback HTTP attempt. It no longer limits how long a record runs. Accepts 1000 to 1800000 so existing records still validate, and applies min(timeout_ms, 600000). Sending both fields with different values fails validation with invalid_workflow. Use callback_timeout_ms.
          - `url` string, uri, required — HTTPS callback endpoint. Internal and private address targets are rejected by the worker's server-side request policy.
    - `custom_id` string — Your identifier for this request, echoed on the matching result line so you can join results back to inputs. Unique within the batch.
    - `method` 'POST' — HTTP method for the record; only POST is supported.
    - `url` string — The endpoint this record targets; must match the batch's endpoint.

## Response `200`

OK

- BatchObject — A batch inference job (OpenAI-compatible).
  - `cancelled_at` integer
  - `completed_at` integer
  - `completion_window` '24h', required
  - `created_at` integer, required — Unix timestamp (seconds).
  - `endpoint` string, required — The API family every record in the batch calls.
  - `error_file_id` string — Present once results are written; fetch its content for failed records.
  - `expired_at` integer
  - `expires_at` integer — When the completion window closes (created_at + 24h).
  - `failed_at` integer
  - `id` string, required — Batch id (batch_…).
  - `input_file_id` string, required
  - `metadata` object — Your key-value pairs, echoed back unchanged.
  - `model` string — Model id, settled from the input file during validation.
  - `object` 'batch', required
  - `output_file_id` string — Present once results are written; fetch its content for successful records.
  - `priority` 'standard' | 'expedited', required — The scheduling tier this batch runs and is billed at. Batches created before the tier existed read as standard.
  - `request_counts` BatchRequestCounts, required — Progress counters, settled as the batch validates and shards complete.
    - `completed` integer, required
    - `failed` integer, required
    - `total` integer, required
  - `status` 'validating' | 'in_progress' | 'finalizing' | 'completed' | 'failed' | 'expired' | 'cancelling' | 'cancelled', required — Lifecycle state. validating → in_progress → finalizing → completed | failed | expired; cancelling → cancelled.
  - `usage` BatchUsage — Rolled-up token and cost totals, present once any progress is recorded. cost_nano_usd is the price you pay (batch discount applied), in nano-USD so sub-cent batches stay exact.
    - `cost_nano_usd` integer, required
    - `input_tokens` integer, required
    - `output_tokens` integer, required

## Other responses

- `400` — The request is invalid
- `401` — Missing or invalid API key
- `404` — Not found (also returned when the organization does not have Batch API access)
- `413` — The upload exceeds the 200 MB limit
- `429` — Too many concurrent uploads; retry shortly
- `503` — The action could not be dispatched; retry

## Changes

- **2026-09-26** `24a11ebdab5e` — 5 info
  - added the new optional request property `requests/items/body/openrelay/workflow/callback_timeout_ms`
  - added the new optional request property `requests/items/body/openrelay/workflow/max_requests`
  - added the new optional request property `requests/items/body/openrelay/workflow/max_tokens`
  - the `timeout_ms` request property default value `120000` was removed
  - …1 more
- **2026-09-24** `fb9c09eb89bc` — 1 info
  - added the new optional request property `requests/items/body/openrelay/workflow`
- **2026-09-19** `702f7cdaf46f` — 2 info
  - added the new optional request property `priority`
  - added the required property `priority` to the response with the `200` status
- **2026-08-07** `4c42fd61ed41` — 4 info
  - added the new optional request property `requests/items/body`
  - added the new optional request property `requests/items/custom_id`
  - added the new optional request property `requests/items/method`
  - added the new optional request property `requests/items/url`

[Change history](https://skmtc.dev/openrelay/apis/openrelay-api/changes/v1/batches/post.md)

---

[API](https://skmtc.dev/openrelay/apis/openrelay-api.md) · [All operations](https://skmtc.dev/openrelay/apis/openrelay-api/llms.txt) · [OpenAPI document](https://skmtc.dev/openrelay/apis/openrelay-api/revisions/24a11ebdab5e?raw)
