---
title: "Async"
method: POST
path: "/v1/inference/async"
tags: ["Inference v2"]
---

# Async

`POST /v1/inference/async`

Asynchronous inference — returns job_id immediately for polling.

## Request body

- V1InferenceRequest — OpenAI-compatible inference request for /v1/inference.
  - `model` string, required — Model name or HuggingFace repo
  - `inputs` union, required — Text input(s) for inference
    - string[]
    - string
  - `max_tokens` integer
  - `temperature` number
  - `stream` boolean — Stream response via SSE

## Response `200`

Successful Response

- unknown

## Other responses

- `422` — Validation Error

## Changes

- **2026-04-01** `526c3dd3d9d7` — 1 info
  - endpoint added

[Change history](https://skmtc.dev/aabiro/apis/xcelsior/changes/v1/inference/async/post.md)

---

[API](https://skmtc.dev/aabiro/apis/xcelsior.md) · [All operations](https://skmtc.dev/aabiro/apis/xcelsior/llms.txt) · [OpenAPI document](https://skmtc-service-production.skmtc.workers.dev/v1/apis/aabiro/xcelsior/revisions/32211b6f9d65/schema)
