---
title: "Api Inference Create Endpoint"
method: POST
path: "/api/v2/inference/endpoints"
tags: ["Inference v2"]
---

# Api Inference Create Endpoint

`POST /api/v2/inference/endpoints`

Create a serverless inference endpoint.

If min_workers >= 1, a worker container is provisioned immediately.
Billing is real-time from credits (per-second compute + per-token inference).

## Request body

- InferenceEndpointCreate
  - `model_name` string, required
  - `gpu_type` string
  - `region` string
  - `docker_image` string
  - `min_workers` integer
  - `max_workers` integer
  - `max_batch_size` integer
  - `max_concurrent` integer
  - `scaledown_window_sec` integer
  - `mode` string
  - `health_endpoint` string
  - `api_format` string

## Response `200`

Successful Response

- unknown

## Other responses

- `422` — Validation Error

## Changes

- **2026-04-03** `74dc97653bc3` — 8 info
  - added the new optional request property `api_format`
  - added the new optional request property `docker_image`
  - added the new optional request property `gpu_type`
  - added the new optional request property `health_endpoint`
  - …4 more
- **2026-04-01** `526c3dd3d9d7` — 1 info
  - endpoint added

[Change history](https://skmtc.dev/aabiro/apis/xcelsior/changes/api/v2/inference/endpoints/post.md)

---

[API](https://skmtc.dev/aabiro/apis/xcelsior.md) · [All operations](https://skmtc.dev/aabiro/apis/xcelsior/llms.txt) · [OpenAPI document](https://skmtc-service-production.skmtc.workers.dev/v1/apis/aabiro/xcelsior/revisions/95960849edcd/schema)
