---
title: "Replace the endpoint's traffic allocation across ALL backends (platform admin)"
method: PUT
path: "/v1/admin/inference/catalog/{modelId}/allocation"
tags: ["Internal"]
---

# Replace the endpoint's traffic allocation across ALL backends (platform admin)

`PUT /v1/admin/inference/catalog/{modelId}/allocation`

Percentage-based traffic split over the model's backends. Every backend must be listed; an unlisted backend is rejected rather than silently zeroed. Targets are percentages (they must sum to 100); the server normalizes them to integer weights once, with largest-remainder rounding, so every consumer of this API shares one deterministic rule. A PUT that targets a self-hosted pool whose lane has a live rollout campaign answers 409: pause the rollout, then edit the allocation.

## Path parameters

- `modelId` string, required

## Request body

- PutInferenceCatalogAllocationRequest — Percentage traffic targets over EVERY backend of the model.
  - `allocations` object[], required
    - `endpoint` string, required
    - `weightPct` number, double, required — Percent of traffic; all entries must sum to 100

## Response `200`

Applied; echoes the persisted integer weights

- InferenceCatalogModel — The endpoint-level routing view of one catalog model.
  - `backends` InferenceCatalogBackend[], required
    - `accelerator` string — Hardware class tag, when known
    - `endpoint` string, required — pool://<poolId> for self-hosted capacity; an external URL for managed providers
    - `healthy` boolean, required — false = excluded from routing entirely (the kill switch)
    - `maxConcurrency` integer — Per-backend in-flight cap; 0 = model default
    - `maxPromptBytes` integer
    - `minPromptBytes` integer
    - `nativeFamilies` string[]
    - `pools` object[] — Self-hosted only: the deployments registering into this pool
      - `deploymentId` string
      - `nodePool` string
      - `status` string
      - `version` string
    - `provider` string, required — selfhosted|bedrock|openrouter|openai|anthropic (openrelay pools report selfhosted)
    - `ref` string, required — Provider-side model id / deployment id
    - `standby` boolean — true = held out of even the overflow tail (an unramped rollout green)
    - `transport` string — direct|tunnel|privatelink
    - `version` string — Low-cardinality rollout tag, when set
    - `weight` integer, required — Live routing weight (0 = overflow tail)
  - `enabled` boolean
  - `modelId` string, required
  - `routing` object
    - `families` object
    - `maxAttempts` integer
    - `strategy` string — first_healthy|weighted|sticky (empty = gateway default)
  - `status` string
  - `updatedAt` string

## Other responses

- `400` — The request is invalid
- `401` — Missing or invalid API key
- `403` — API key lacks the required scope
- `404` — Resource not found
- `409` — The request conflicts with existing state
- `503` — A required integration (e.g. payments) is not configured

---

[API](https://skmtc.dev/openrelay/apis/openrelay-api.md) · [All operations](https://skmtc.dev/openrelay/apis/openrelay-api/llms.txt) · [OpenAPI document](https://skmtc-service-production.skmtc.workers.dev/v1/apis/openrelay/openrelay-api/revisions/3dc47c9667e8/schema)
