---
title: "Create a chat completion"
method: POST
path: "/chat/completions"
tags: ["Chat"]
---

# Create a chat completion

`POST /chat/completions`

Sends a request for a model response for the given chat conversation. Supports both streaming and non-streaming modes.

## Headers

- `X-OpenRouter-Metadata` 'disabled' | 'enabled' — Opt-in level for surfacing routing metadata on the response under `openrouter_metadata`.

## Request body

- ChatRequest — Chat completion request parameters
  - `cache_control` AnthropicCacheControlDirective — Enable automatic prompt caching. When set at the top level, the system automatically applies cache breakpoints to the last cacheable block in the request. When set on an individual content block, it marks an explicit cache breakpoint; block-level markers also work on OpenAI models that support explicit prompt caching — OpenRouter converts them to the provider's native format.
    - `ttl` '5m' | '1h'
    - `type` 'ephemeral', required
  - `debug` ChatDebugOptions — Debug options for inspecting request transformations (streaming only)
    - `echo_upstream_body` boolean — If true, includes the transformed upstream request body in a debug chunk at the start of the stream. Only works with streaming mode.
  - `frequency_penalty` number, double, nullable — Frequency penalty (-2.0 to 2.0)
  - `image_config` ImageConfig — Provider-specific image configuration options. Keys and values vary by model/provider. See https://openrouter.ai/docs/guides/overview/multimodal/image-generation for more details.
  - `logit_bias` object, nullable — Token logit bias adjustments
  - `logprobs` boolean, nullable — Return log probabilities
  - `max_completion_tokens` integer, nullable — Maximum tokens in completion
  - `max_tokens` integer, nullable — Maximum tokens (deprecated, use max_completion_tokens). Note: some providers enforce a minimum of 16.
  - `messages` ChatMessages[], required — List of messages for the conversation
    - union — Chat completion message with role-based discrimination
      - ChatSystemMessage — System message for setting behavior
        - `content` union, required — System message content
          - string
          - ChatContentText[]
            - `cache_control` ChatContentCacheControl — Anthropic-style cache breakpoint for the content part. Interchangeable with the OpenAI-style `prompt_cache_breakpoint` marker: OpenRouter converts between the two based on the provider serving the request.
              - …
            - `prompt_cache_breakpoint` PromptCacheBreakpoint, nullable — Marks an explicit prompt-cache boundary on this content block (OpenAI-style). Everything through the block carrying this marker is part of the candidate cached prefix. Supported natively by OpenAI GPT-5.6 and newer; on providers that use Anthropic-style `cache_control`, OpenRouter converts the marker to that format automatically.
              - …
            - `text` string, required
            - `type` 'text', required
        - `name` string — Optional name for the system message
        - `role` 'system', required
      - ChatUserMessage — User message
        - `content` union, required — User message content
          - string
          - ChatContentItems[]
            - union — Content part for chat completion messages
              - …
        - `name` string — Optional name for the user
        - `role` 'user', required
      - ChatDeveloperMessage — Developer message
        - `content` union, required — Developer message content
          - string
          - ChatContentText[]
            - `cache_control` ChatContentCacheControl — Anthropic-style cache breakpoint for the content part. Interchangeable with the OpenAI-style `prompt_cache_breakpoint` marker: OpenRouter converts between the two based on the provider serving the request.
              - …
            - `prompt_cache_breakpoint` PromptCacheBreakpoint, nullable — Marks an explicit prompt-cache boundary on this content block (OpenAI-style). Everything through the block carrying this marker is part of the candidate cached prefix. Supported natively by OpenAI GPT-5.6 and newer; on providers that use Anthropic-style `cache_control`, OpenRouter converts the marker to that format automatically.
              - …
            - `text` string, required
            - `type` 'text', required
        - `name` string — Optional name for the developer message
        - `role` 'developer', required
      - ChatAssistantMessage — Assistant message for requests and responses
        - `audio` ChatAudioOutput — Audio output data or reference
          - `data` string — Base64 encoded audio data
          - `expires_at` integer — Audio expiration timestamp
          - `id` string — Audio output identifier
          - `transcript` string — Audio transcript
        - `content` union — Assistant message content
          - string
          - ChatContentItems[]
            - union — Content part for chat completion messages
              - …
          - unknown
        - `images` object[] — Generated images from image generation models
          - `image_url` object, required
            - `url` string, required — URL or base64-encoded data of the generated image
        - `name` string — Optional name for the assistant
        - `reasoning` string, nullable — Reasoning output
        - `reasoning_details` ReasoningDetailUnion[] — Reasoning details for extended thinking models
          - union — Reasoning detail union schema
            - ReasoningDetailSummary — Reasoning detail summary schema
              - …
            - ReasoningDetailEncrypted — Reasoning detail encrypted schema
              - …
            - ReasoningDetailText — Reasoning detail text schema
              - …
            - ReasoningDetailServerToolCall — Record of an OpenRouter server-tool invocation (e.g. openrouter:fusion), carried in reasoning_details so a prior tool call can be rehydrated into a later turn of the same conversation.
              - …
        - `refusal` string, nullable — Refusal message if content was refused
        - `role` 'assistant', required
        - `tool_calls` ChatToolCall[] — Tool calls made by the assistant
          - `function` object, required
            - `arguments` string, required — Function arguments as JSON string
            - `name` string, required — Function name to call
          - `id` string, required — Tool call identifier
          - `type` 'function', required
      - ChatToolMessage — Tool response message
        - `content` union, required — Tool response content
          - string
          - ChatContentItems[]
            - union — Content part for chat completion messages
              - …
        - `role` 'tool', required
        - `tool_call_id` string, required — ID of the assistant message tool call this message responds to
  - `metadata` object — Key-value pairs for additional object information (max 16 pairs, 64 char keys, 512 char values)
  - `min_p` number, double, nullable — Minimum probability threshold relative to the most likely token. Tokens with probability below min_p * (probability of top token) are filtered out. Not all providers support this parameter.
  - `modalities` string[] — Output modalities for the response. Supported values are "text", "image", and "audio".
  - `model` string — Model to use for completion
  - `models` string[] — Models to use for completion
  - `parallel_tool_calls` boolean, nullable — Whether to enable parallel function calling during tool use. When true, the model may generate multiple tool calls in a single response.
  - `plugins` union[] — Plugins you want to enable for this request, including their settings.
    - union
      - AutoRouterPlugin
        - `allowed_models` string[] — List of model patterns to filter which models the auto-router can route between. Supports wildcards (e.g., "anthropic/*" matches all Anthropic models). When not specified, uses the default supported models list.
        - `cost_quality_tradeoff` integer — Controls cost vs. quality routing tradeoff (0–10). 0 = pure quality (best model regardless of cost), 10 = maximize for cost (cheapest model wins). Intermediate values blend quality and cost signals continuously. Defaults to 7.
        - `enabled` boolean — Set to false to disable the auto-router plugin for this request. Defaults to true.
        - `id` 'auto-router', required
      - ModerationPlugin
        - `id` 'moderation', required
      - WebSearchPlugin
        - `enabled` boolean — Set to false to disable the web-search plugin for this request. Defaults to true.
        - `engine` 'native' | 'exa' | 'firecrawl' | 'parallel' | 'perplexity' — The search engine to use for web search.
        - `exclude_domains` string[] — A list of domains to exclude from web search results. Supports wildcards (e.g. "*.substack.com") and path filtering (e.g. "openai.com/blog").
        - `id` 'web', required
        - `include_domains` string[] — A list of domains to restrict web search results to. Supports wildcards (e.g. "*.substack.com") and path filtering (e.g. "openai.com/blog").
        - `max_results` integer
        - `max_uses` integer — Maximum number of times the model can invoke web search in a single turn. Passed through to native providers that support it (e.g. Anthropic).
        - `search_prompt` string
        - `user_location` object, nullable — Approximate user location for location-biased search results. Passed through to native providers that support it (e.g. Anthropic).
          - `city` string, nullable
          - `country` string, nullable
          - `region` string, nullable
          - `timezone` string, nullable
          - `type` 'approximate', required
      - WebFetchPlugin
        - `allowed_domains` string[] — Only fetch from these domains.
        - `blocked_domains` string[] — Never fetch from these domains.
        - `id` 'web-fetch', required
        - `max_content_tokens` integer — Maximum content length in approximate tokens. Content exceeding this limit is truncated.
        - `max_uses` integer — Maximum number of web fetches per request. Once exceeded, the tool returns an error.
      - FileParserPlugin
        - `enabled` boolean — Set to false to disable the file-parser plugin for this request. Defaults to true.
        - `id` 'file-parser', required
        - `pdf` PDFParserOptions — Options for PDF parsing.
          - `engine` union — The engine to use for parsing PDF files. "pdf-text" is deprecated and automatically redirected to "cloudflare-ai".
            - 'mistral-ocr' | 'native' | 'cloudflare-ai'
            - 'pdf-text'
      - ResponseHealingPlugin
        - `enabled` boolean — Set to false to disable the response-healing plugin for this request. Defaults to true.
        - `id` 'response-healing', required
      - ContextCompressionPlugin
        - `enabled` boolean — Set to false to disable the context-compression plugin for this request. Defaults to true.
        - `engine` 'middle-out' — The compression engine to use. Defaults to "middle-out".
        - `id` 'context-compression', required
      - ParetoRouterPlugin
        - `enabled` boolean — Set to false to disable the pareto-router plugin for this request. Defaults to true.
        - `id` 'pareto-router', required
        - `min_coding_score` number, double — Minimum coding quality score between 0 and 1. Maps to internal quality tiers: >= 0.66 → high (top coding models), >= 0.33 → medium (strong modern flagships), < 0.33 → low (capable coders above the median). Omit to default to the highest tier (equivalent to >= 0.66).
        - `price_source` 'prompt' | 'weighted_avg' — Price source for the Pareto frontier cost axis. "prompt" uses catalog list price (endpoint.pricing.prompt). "weighted_avg" uses traffic-weighted effective input price from ClickHouse, falling back to prompt price for models without traffic data. Defaults to "prompt".
      - FusionPlugin
        - `analysis_models` string[] — Slugs of models to run in parallel as the "expert panel" the judge analyzes. Each model receives the same user prompt with web_search + web_fetch enabled. Capped at 8 models to bound cost amplification. When omitted, defaults to the Quality preset from the /labs/fusion UI (~anthropic/claude-opus-latest, ~openai/gpt-latest, ~google/gemini-pro-latest).
        - `enabled` boolean — Set to false to disable the fusion plugin for this request. Defaults to true.
        - `id` 'fusion', required
        - `max_tool_calls` integer — Maximum number of tool-calling steps each panelist (analysis model) and the judge model may take during their agentic web-research loop. Models with web_search/web_fetch enabled iterate until they produce a text response or hit this ceiling. Defaults to 8. Capped at 16.
        - `model` string — Slug of the model that performs both the judge step (with web_search + web_fetch) and the final synthesis. When omitted, defaults to the first model in the Quality preset.
        - `preset` 'general-high' | 'general-budget' | 'general-fast' — A curated OpenRouter fusion preset (slugs follow `<task>-<tier>`, e.g. `general-high`). Expands server-side into the preset's analysis_models panel and judge model, so callers never name individual models. Explicitly provided `analysis_models` / `model` take precedence.
        - `tools` object[] — Server tools available to panelist and judge inner calls. Each entry uses the same `{ type, parameters? }` shorthand as the outer Chat Completions request. When omitted, defaults to `[{ type: "openrouter:web_search" }, { type: "openrouter:web_fetch" }]`. Pass an empty array to disable tools entirely (panelists answer from parametric knowledge only).
          - `parameters` object — Optional configuration forwarded as the tool's `parameters` object.
          - `type` string, required — Server tool type identifier (e.g. "openrouter:web_search", "openrouter:web_fetch").
  - `prediction` Prediction, nullable — Static predicted output content. Supported models can use this to reduce latency when much of the response is known in advance.
    - `content` union, required
      - string
      - PredictionContentText[]
        - `text` string, required
        - `type` 'text', required
    - `type` 'content', required
  - `presence_penalty` number, double, nullable — Presence penalty (-2.0 to 2.0)
  - `prompt_cache_key` string, nullable
  - `prompt_cache_options` PromptCacheOptions, nullable — Request-level prompt-cache controls. `mode: "explicit"` disables OpenAI-managed breakpoints so only blocks marked with `prompt_cache_breakpoint` are cached. Only supported by OpenAI GPT-5.6 and newer.
    - `mode` 'explicit', required
    - `ttl` string, nullable
  - `provider` ProviderPreferences, nullable — When multiple model providers are available, optionally indicate your routing preference.
    - `allow_fallbacks` boolean, nullable — Whether to allow backup providers to serve requests - true: (default) when the primary provider (or your custom providers in "order") is unavailable, use the next best provider. - false: use only the primary/custom provider, and return the upstream error if it's unavailable.
    - `data_collection` 'deny' | 'allow' | 'null', nullable — Data collection setting. If no available model provider meets the requirement, your request will return an error. - allow: (default) allow providers which store user data non-transiently and may train on it - deny: use only providers which do not collect user data.
    - `enforce_distillable_text` boolean, nullable — Whether to restrict routing to only models that allow text distillation. When true, only models where the author has allowed distillation will be used.
    - `ignore` union[], nullable — List of provider slugs to ignore. If provided, this list is merged with your account-wide ignored provider settings for this request.
      - union
        - 'Meta' | 'AkashML' | 'AI21' | 'AionLabs' | 'Alibaba' | 'Ambient' | 'Baidu' | 'Amazon Bedrock' | 'Amazon Nova' | 'Anthropic' | 'Arcee AI' | 'AtlasCloud' | 'Avian' | 'Azure' | 'BaseTen' | 'BytePlus' | 'Black Forest Labs' | 'Cerebras' | 'Chutes' | 'Cirrascale' | 'Clarifai' | 'Cloudflare' | 'Cohere' | 'Crucible' | 'Crusoe' | 'Darkbloom' | 'Decart' | 'Deepgram' | 'DeepInfra' | 'DeepSeek' | 'DekaLLM' | 'DigitalOcean' | 'Featherless' | 'Fireworks' | 'Friendli' | 'GMICloud' | 'Google' | 'Google AI Studio' | 'Groq' | 'HeyGen' | 'Inception' | 'Inceptron' | 'InferenceNet' | 'Ionstream' | 'Infermatic' | 'Io Net' | 'Inferact vLLM' | 'Inflection' | 'Liquid' | 'Mara' | 'Mancer 2' | 'Minimax' | 'ModelRun' | 'Mistral' | 'Modular' | 'Moonshot AI' | 'Morph' | 'NCompass' | 'Nebius' | 'Nex AGI' | 'NextBit' | 'Novita' | 'Nvidia' | 'OpenAI' | 'OpenInference' | 'Parasail' | 'Poolside' | 'Perceptron' | 'Perplexity' | 'Phala' | 'Recraft' | 'Reka' | 'Relace' | 'Sail Research' | 'Sakana AI' | 'SambaNova' | 'Seed' | 'SiliconFlow' | 'Sourceful' | 'StepFun' | 'Stealth' | 'StreamLake' | 'Switchpoint' | 'Tenstorrent' | 'Together' | 'Upstage' | 'Venice' | 'Wafer' | 'WandB' | 'Quiver' | 'Krea' | 'Xiaomi' | 'xAI' | 'Z.AI' | 'FakeProvider'
        - string
    - `max_price` object — The object specifying the maximum price you want to pay for this request. USD price per million tokens, for prompt and completion.
      - `audio` string — Maximum price in USD per audio unit
      - `completion` string — Maximum price in USD per million completion tokens
      - `image` string — Maximum price in USD per image
      - `prompt` string — Maximum price in USD per million prompt tokens
      - `request` string — Maximum price in USD per request
    - `only` union[], nullable — List of provider slugs to allow. If provided, this list is merged with your account-wide allowed provider settings for this request.
      - union
        - 'Meta' | 'AkashML' | 'AI21' | 'AionLabs' | 'Alibaba' | 'Ambient' | 'Baidu' | 'Amazon Bedrock' | 'Amazon Nova' | 'Anthropic' | 'Arcee AI' | 'AtlasCloud' | 'Avian' | 'Azure' | 'BaseTen' | 'BytePlus' | 'Black Forest Labs' | 'Cerebras' | 'Chutes' | 'Cirrascale' | 'Clarifai' | 'Cloudflare' | 'Cohere' | 'Crucible' | 'Crusoe' | 'Darkbloom' | 'Decart' | 'Deepgram' | 'DeepInfra' | 'DeepSeek' | 'DekaLLM' | 'DigitalOcean' | 'Featherless' | 'Fireworks' | 'Friendli' | 'GMICloud' | 'Google' | 'Google AI Studio' | 'Groq' | 'HeyGen' | 'Inception' | 'Inceptron' | 'InferenceNet' | 'Ionstream' | 'Infermatic' | 'Io Net' | 'Inferact vLLM' | 'Inflection' | 'Liquid' | 'Mara' | 'Mancer 2' | 'Minimax' | 'ModelRun' | 'Mistral' | 'Modular' | 'Moonshot AI' | 'Morph' | 'NCompass' | 'Nebius' | 'Nex AGI' | 'NextBit' | 'Novita' | 'Nvidia' | 'OpenAI' | 'OpenInference' | 'Parasail' | 'Poolside' | 'Perceptron' | 'Perplexity' | 'Phala' | 'Recraft' | 'Reka' | 'Relace' | 'Sail Research' | 'Sakana AI' | 'SambaNova' | 'Seed' | 'SiliconFlow' | 'Sourceful' | 'StepFun' | 'Stealth' | 'StreamLake' | 'Switchpoint' | 'Tenstorrent' | 'Together' | 'Upstage' | 'Venice' | 'Wafer' | 'WandB' | 'Quiver' | 'Krea' | 'Xiaomi' | 'xAI' | 'Z.AI' | 'FakeProvider'
        - string
    - `order` union[], nullable — An ordered list of provider slugs. The router will attempt to use the first provider in the subset of this list that supports your requested model, and fall back to the next if it is unavailable. If no providers are available, the request will fail with an error message.
      - union
        - 'Meta' | 'AkashML' | 'AI21' | 'AionLabs' | 'Alibaba' | 'Ambient' | 'Baidu' | 'Amazon Bedrock' | 'Amazon Nova' | 'Anthropic' | 'Arcee AI' | 'AtlasCloud' | 'Avian' | 'Azure' | 'BaseTen' | 'BytePlus' | 'Black Forest Labs' | 'Cerebras' | 'Chutes' | 'Cirrascale' | 'Clarifai' | 'Cloudflare' | 'Cohere' | 'Crucible' | 'Crusoe' | 'Darkbloom' | 'Decart' | 'Deepgram' | 'DeepInfra' | 'DeepSeek' | 'DekaLLM' | 'DigitalOcean' | 'Featherless' | 'Fireworks' | 'Friendli' | 'GMICloud' | 'Google' | 'Google AI Studio' | 'Groq' | 'HeyGen' | 'Inception' | 'Inceptron' | 'InferenceNet' | 'Ionstream' | 'Infermatic' | 'Io Net' | 'Inferact vLLM' | 'Inflection' | 'Liquid' | 'Mara' | 'Mancer 2' | 'Minimax' | 'ModelRun' | 'Mistral' | 'Modular' | 'Moonshot AI' | 'Morph' | 'NCompass' | 'Nebius' | 'Nex AGI' | 'NextBit' | 'Novita' | 'Nvidia' | 'OpenAI' | 'OpenInference' | 'Parasail' | 'Poolside' | 'Perceptron' | 'Perplexity' | 'Phala' | 'Recraft' | 'Reka' | 'Relace' | 'Sail Research' | 'Sakana AI' | 'SambaNova' | 'Seed' | 'SiliconFlow' | 'Sourceful' | 'StepFun' | 'Stealth' | 'StreamLake' | 'Switchpoint' | 'Tenstorrent' | 'Together' | 'Upstage' | 'Venice' | 'Wafer' | 'WandB' | 'Quiver' | 'Krea' | 'Xiaomi' | 'xAI' | 'Z.AI' | 'FakeProvider'
        - string
    - `preferred_max_latency` union — Preferred maximum latency (in seconds). Can be a number (applies to p50) or an object with percentile-specific cutoffs. Endpoints above the threshold(s) may still be used, but are deprioritized in routing. When using fallback models, this may cause a fallback model to be used instead of the primary model if it meets the threshold.
      - number, double
      - PercentileLatencyCutoffs — Percentile-based latency cutoffs. All specified cutoffs must be met for an endpoint to be preferred.
        - `p50` number, double, nullable — Maximum p50 latency (seconds)
        - `p75` number, double, nullable — Maximum p75 latency (seconds)
        - `p90` number, double, nullable — Maximum p90 latency (seconds)
        - `p99` number, double, nullable — Maximum p99 latency (seconds)
      - unknown
    - `preferred_min_throughput` union — Preferred minimum throughput (in tokens per second). Can be a number (applies to p50) or an object with percentile-specific cutoffs. Endpoints below the threshold(s) may still be used, but are deprioritized in routing. When using fallback models, this may cause a fallback model to be used instead of the primary model if it meets the threshold.
      - number, double
      - PercentileThroughputCutoffs — Percentile-based throughput cutoffs. All specified cutoffs must be met for an endpoint to be preferred.
        - `p50` number, double, nullable — Minimum p50 throughput (tokens/sec)
        - `p75` number, double, nullable — Minimum p75 throughput (tokens/sec)
        - `p90` number, double, nullable — Minimum p90 throughput (tokens/sec)
        - `p99` number, double, nullable — Minimum p99 throughput (tokens/sec)
      - unknown
    - `quantizations` Quantization[], nullable — A list of quantization levels to filter the provider by.
    - `require_parameters` boolean, nullable — Whether to filter providers to only those that support the parameters you've provided. If this setting is omitted or set to false, then providers will receive only the parameters they support, and ignore the rest.
    - `sort` union — The sorting strategy to use for this request, if "order" is not specified. When set, no load balancing is performed.
      - 'price' | 'throughput' | 'latency' | 'exacto' — The provider sorting strategy (price, throughput, latency)
      - ProviderSortConfig — The provider sorting strategy (price, throughput, latency)
        - `by` 'price' | 'throughput' | 'latency' | 'exacto' | 'null', nullable — The provider sorting strategy (price, throughput, latency)
        - `partition` 'model' | 'none' | 'null', nullable — Partitioning strategy for sorting: "model" (default) groups endpoints by model before sorting (fallback models remain fallbacks), "none" sorts all endpoints together regardless of model.
      - unknown
    - `zdr` boolean, nullable — Whether to restrict routing to only ZDR (Zero Data Retention) endpoints. When true, only endpoints that do not retain prompts will be used.
  - `reasoning` object — Configuration options for reasoning models
    - `effort` 'max' | 'xhigh' | 'high' | 'medium' | 'low' | 'minimal' | 'none' | 'null', nullable — Constrains effort on reasoning for reasoning models
    - `summary` 'auto' | 'concise' | 'detailed' | 'null', nullable
  - `reasoning_effort` 'max' | 'xhigh' | 'high' | 'medium' | 'low' | 'minimal' | 'none' | 'null', nullable — Shorthand for setting reasoning effort. Equivalent to setting reasoning.effort. Cannot be used simultaneously with reasoning.effort if they differ.
  - `repetition_penalty` number, double, nullable — Penalizes tokens based on how much they have already appeared in the text. A value of 1.0 means no penalty. Values above 1.0 penalize repeated tokens more strongly. Not all providers support this parameter.
  - `response_format` union — Response format configuration
    - ChatFormatTextConfig — Default text response format
      - `type` 'text', required
    - ChatFormatJsonObjectConfig — JSON object response format
      - `type` 'json_object', required
    - ChatFormatJsonSchemaConfig — JSON Schema response format for structured outputs
      - `json_schema` ChatJsonSchemaConfig, required — JSON Schema configuration object
        - `description` string — Schema description for the model
        - `name` string, required — Schema name (a-z, A-Z, 0-9, underscores, dashes, max 64 chars)
        - `schema` object — JSON Schema object
        - `strict` boolean, nullable — Enable strict schema adherence
      - `type` 'json_schema', required
    - ChatFormatGrammarConfig — Custom grammar response format
      - `grammar` string, required — Custom grammar for text generation
      - `type` 'grammar', required
    - ChatFormatPythonConfig — Python code response format
      - `type` 'python', required
  - `route` 'fallback' | 'sort' | 'null', nullable — **DEPRECATED** Use providers.sort.partition instead. Backwards-compatible alias for providers.sort.partition. Accepts legacy values: "fallback" (maps to "model"), "sort" (maps to "none").
  - `seed` integer, nullable — Random seed for deterministic outputs
  - `service_tier` 'auto' | 'default' | 'flex' | 'priority' | 'scale' | 'null', nullable — The service tier to use for processing this request.
  - `session_id` string — A unique identifier for grouping related requests (e.g., a conversation or agent workflow). When provided, OpenRouter uses it as the sticky routing key, routing all requests in the session to the same provider to maximize prompt cache hits. Also used for observability grouping. If provided in both the request body and the x-session-id header, the body value takes precedence. Maximum of 256 characters.
  - `stop` union — Stop sequences (up to 4)
    - string
    - string[]
    - unknown
  - `stop_server_tools_when` StopServerToolsWhenCondition[] — Stop conditions for the server-tool agent loop. Any condition firing halts the loop (OR logic). When set, this overrides `max_tool_calls`.
    - union — A single condition that, when met, halts the server-tool agent loop.
      - StopServerToolsWhenStepCountIs — Stop after the agent loop has executed this many steps.
        - `step_count` integer, required
        - `type` 'step_count_is', required
      - StopServerToolsWhenHasToolCall — Stop after a tool with this name has been called.
        - `tool_name` string, required
        - `type` 'has_tool_call', required
      - StopServerToolsWhenMaxTokensUsed — Stop once cumulative token usage across the loop exceeds this threshold.
        - `max_tokens` integer, required
        - `type` 'max_tokens_used', required
      - StopServerToolsWhenMaxCost — Stop once cumulative cost across the loop exceeds this dollar threshold.
        - `max_cost_in_dollars` number, double, required
        - `type` 'max_cost', required
      - StopServerToolsWhenFinishReasonIs — Stop when the upstream model emits this finish reason (e.g. `length`).
        - `reason` string, required
        - `type` 'finish_reason_is', required
  - `stream` boolean — Enable streaming response
  - `stream_options` ChatStreamOptions, nullable — Streaming configuration options
    - `include_usage` boolean — Deprecated: This field has no effect. Full usage details are always included.
  - `temperature` number, double, nullable — Sampling temperature (0-2)
  - `tool_choice` union — Tool choice configuration
    - 'none'
    - 'auto'
    - 'required'
    - ChatNamedToolChoice — Named tool choice for specific function
      - `function` object, required
        - `name` string, required — Function name to call
      - `type` 'function', required
    - ChatServerToolChoice — OpenRouter extension: force a specific server tool by naming it directly in `tool_choice.type` instead of wrapping it in `{ type: "function", function: { name } }`.
      - `type` string, required — OpenRouter server-tool type to force (e.g. `openrouter:web_search`, `web_search`, `web_search_preview`).
  - `tools` ChatFunctionTool[] — Available tools for function calling
    - union — Tool definition for function calling (regular function or OpenRouter built-in server tool)
      - object
        - `cache_control` ChatContentCacheControl — Anthropic-style cache breakpoint for the content part. Interchangeable with the OpenAI-style `prompt_cache_breakpoint` marker: OpenRouter converts between the two based on the provider serving the request.
          - `ttl` '5m' | '1h'
          - `type` 'ephemeral', required
        - `function` object, required — Function definition for tool calling
          - `description` string — Function description for the model
          - `name` string, required — Function name (a-z, A-Z, 0-9, underscores, dashes, max 64 chars)
          - `parameters` object — Function parameters as JSON Schema object
          - `strict` boolean, nullable — Enable strict schema adherence
        - `type` 'function', required
      - AdvisorServerToolOpenRouter — OpenRouter built-in server tool: consults a higher-intelligence advisor model (any OpenRouter model) for guidance mid-generation and returns its response. The advisor may run as a sub-agent with its own tools. Include multiple entries to offer several named advisors; at most one entry may omit `name` to act as the default advisor.
        - `parameters` AdvisorServerToolConfig — Configuration for one openrouter:advisor server tool entry.
          - `forward_transcript` boolean — When true, the full parent conversation is forwarded to the advisor so it sees the same context the executor does (and the tool-call `prompt`, if given, is appended as a final user turn). When false or omitted, the advisor receives only the `prompt` the executor passes in the tool call.
          - `instructions` string — System instructions for the advisor sub-agent. When omitted, the advisor responds with no system prompt of its own.
          - `max_completion_tokens` integer — Maximum number of output tokens (including reasoning) the advisor may produce. When omitted, the provider's default applies.
          - `max_tool_calls` integer — Maximum number of tool-calling steps the advisor sub-agent may take during its agentic loop. Capped at 25. Only relevant when the advisor is given tools.
          - `model` string — Slug of the advisor model to consult (any OpenRouter model). When omitted, the executor can choose it via the tool call's `model` argument; if neither is set, the model from the outer API request is used. The advisor tool itself cannot be the advisor model.
          - `name` string — Optional name for this advisor. The model sees one tool per named advisor (and one default for an unnamed entry). Names must be unique across advisor entries. Letters, digits, spaces, underscores, and dashes; trimmed; 1–64 chars.
          - `reasoning` AdvisorReasoning — Reasoning configuration forwarded to the advisor call. Use this to control reasoning effort and token budget for models that support extended thinking.
            - `effort` 'max' | 'xhigh' | 'high' | 'medium' | 'low' | 'minimal' | 'none' — Reasoning effort level for the advisor call.
            - `max_tokens` integer — Maximum number of reasoning tokens the advisor may use.
          - `stream` boolean — When true, the advisor's advice streams incrementally as it is produced. In the Responses API this emits `response.output_text.delta` events targeting the advisor output item; the final `advice` field is still set on the completed item. Has no effect on the Chat Completions API (where the advice arrives only as the final tool result). When false or omitted, the advice arrives only as the final result.
          - `temperature` number, double — Sampling temperature forwarded to the advisor call. When omitted, the provider's default applies.
          - `tools` AdvisorNestedTool[] — Tools the advisor sub-agent may use while forming its advice. The advisor runs as an agentic sub-agent over these tools, then returns its text. Only OpenRouter server tools are supported — function tools are rejected — and the list must not include the advisor tool itself.
            - `parameters` object
            - `type` string, required
        - `type` 'openrouter:advisor', required
      - BashServerTool — OpenRouter built-in server tool: runs shell commands server-side in a sandboxed container
        - `parameters` BashServerToolConfig — Configuration for the openrouter:bash server tool
          - `engine` 'auto' | 'native' | 'openrouter' — Which bash engine to use. "openrouter" runs commands server-side in the OpenRouter sandbox. "auto" (default) and "native" use native passthrough, returning the tool call to your application to run client-side; OpenRouter does not execute the commands.
          - `environment` union — Execution environment for the bash server tool.
            - ContainerAutoEnvironment — An OpenRouter-managed, auto-provisioned ephemeral container.
              - …
            - ContainerReferenceEnvironment — Reference to a previously created container to reuse.
              - …
          - `sleep_after_seconds` integer — How long (in seconds) the container stays warm after its last command before sleeping, freeing its capacity slot. Idle-based: each command renews the timer. Defaults to 900 (15 minutes); capped at 2592000 (30 days).
        - `type` 'openrouter:bash', required
      - DatetimeServerTool — OpenRouter built-in server tool: returns the current date and time
        - `parameters` DatetimeServerToolConfig — Configuration for the openrouter:datetime server tool
          - `timezone` string — IANA timezone name (e.g. "America/New_York"). Defaults to UTC.
        - `type` 'openrouter:datetime', required
      - FilesServerTool — OpenRouter built-in server tool: read, write, edit, and list workspace files via the Files API. Requires the `x-openrouter-file-ids: openrouter` request header.
        - `parameters` FilesServerToolConfig — Configuration for the openrouter:files server tool
        - `type` 'openrouter:files', required
      - FusionServerToolOpenRouter — OpenRouter built-in server tool: fans out the user prompt to a panel of analysis models, then asks a judge model to summarize their collective output as structured JSON the outer model can synthesize from.
        - `parameters` FusionServerToolConfig — Configuration for the openrouter:fusion server tool.
          - `analysis_models` string[] — Slugs of models to run in parallel as the analysis panel. Each model receives the user prompt with openrouter:web_search and openrouter:web_fetch enabled, then a judge model summarizes the collective output into structured analysis JSON. Capped at 8 models to bound cost amplification. Defaults to the Quality preset from /labs/fusion.
          - `cache_control` AnthropicCacheControlDirective — Enable automatic prompt caching. When set at the top level, the system automatically applies cache breakpoints to the last cacheable block in the request. When set on an individual content block, it marks an explicit cache breakpoint; block-level markers also work on OpenAI models that support explicit prompt caching — OpenRouter converts them to the provider's native format.
            - `ttl` '5m' | '1h'
            - `type` 'ephemeral', required
          - `max_completion_tokens` integer — Maximum number of output tokens (including reasoning tokens) each panelist and the judge model may produce per inner call. Controls the total output budget so reasoning-heavy models like GPT-5.5 do not exhaust their token allowance before producing visible text. When omitted, panelists default to 32000 and the judge to 50000.
          - `max_tool_calls` integer — Maximum number of tool-calling steps each panelist (analysis model) and the judge model may take during their agentic web-research loop. Models with web_search/web_fetch enabled iterate until they produce a text response or hit this ceiling. Defaults to 8. Capped at 16.
          - `model` string — Slug of the judge model that produces the structured analysis JSON. Defaults to the model used in the outer API request.
          - `reasoning` object — Reasoning configuration forwarded to panelist and judge inner calls. Use this to control reasoning effort and token budget for models that support extended thinking.
            - `effort` 'max' | 'xhigh' | 'high' | 'medium' | 'low' | 'minimal' | 'none' — Reasoning effort level for panelist and judge inner calls.
            - `max_tokens` integer — Maximum number of reasoning tokens each panelist and judge model may use. Helps bound cost when models allocate too much budget to chain-of-thought.
          - `temperature` number, double — Temperature forwarded to panelist inner calls. The judge always runs at temperature 0 regardless of this value. When omitted, the provider's default applies.
          - `tools` object[] — Server tools available to panelist and judge inner calls. Each entry uses the same `{ type, parameters? }` shorthand as the outer Chat Completions request. When omitted, defaults to `[{ type: "openrouter:web_search" }, { type: "openrouter:web_fetch" }]`. Pass an empty array to disable tools entirely (panelists answer from parametric knowledge only).
            - `parameters` object — Optional configuration forwarded as the tool's `parameters` object.
            - `type` string, required — Server tool type identifier (e.g. "openrouter:web_search", "openrouter:web_fetch").
        - `type` 'openrouter:fusion', required
      - ImageGenerationServerToolOpenRouter — OpenRouter built-in server tool: generates images from text prompts using an image generation model
        - `parameters` ImageGenerationServerToolConfig — Configuration for the openrouter:image_generation server tool. Accepts all image_config params (aspect_ratio, quality, size, background, output_format, output_compression, moderation, etc.) plus a model field.
          - `model` string — Which image generation model to use (e.g. "openai/gpt-5-image"). Defaults to "openai/gpt-5-image".
        - `type` 'openrouter:image_generation', required
      - ChatSearchModelsServerTool — OpenRouter built-in server tool: searches and filters AI models available on OpenRouter
        - `parameters` SearchModelsServerToolConfig — Configuration for the openrouter:experimental__search_models server tool
          - `max_results` integer — Maximum number of models to return. Defaults to 5, max 20.
        - `type` 'openrouter:experimental__search_models', required
      - SubagentServerToolOpenRouter — OpenRouter built-in server tool: delegates self-contained tasks to a smaller, cheaper, faster worker model (any OpenRouter model) mid-generation and returns its outcome. The worker may run as a sub-agent with its own tools.
        - `parameters` SubagentServerToolConfig — Configuration for the openrouter:subagent server tool.
          - `instructions` string — System instructions for the subagent. When omitted, the subagent responds with no system prompt of its own.
          - `max_completion_tokens` integer — Maximum number of output tokens (including reasoning) the subagent may produce. When omitted, the provider's default applies.
          - `max_tool_calls` integer — Maximum number of tool-calling steps the subagent may take during its agentic loop. Capped at 25. Only relevant when the subagent is given tools. Accepted and validated but not yet enforced on the subagent call.
          - `model` string — Slug of the model that executes delegated tasks (any OpenRouter model). Typically a smaller, cheaper, faster model than the one delegating. When omitted, the model from the outer API request is used. The subagent tool itself cannot be the subagent model.
          - `reasoning` SubagentReasoning — Reasoning configuration forwarded to the subagent call. Use this to control reasoning effort and token budget for models that support extended thinking.
            - `effort` 'max' | 'xhigh' | 'high' | 'medium' | 'low' | 'minimal' | 'none' — Reasoning effort level for the subagent call.
            - `max_tokens` integer — Maximum number of reasoning tokens the subagent may use. Accepted and validated but not yet forwarded to the subagent call.
          - `temperature` number, double — Sampling temperature forwarded to the subagent call. When omitted, the provider's default applies.
          - `tools` SubagentNestedTool[] — Tools the subagent may use while executing a delegated task. The subagent runs as an agentic sub-agent over these tools, then returns its outcome. Only OpenRouter server tools are supported — function tools are rejected — and the list must not include the subagent tool itself.
            - `parameters` object
            - `type` string, required
        - `type` 'openrouter:subagent', required
      - WebFetchServerTool — OpenRouter built-in server tool: fetches full content from a URL (web page or PDF)
        - `parameters` WebFetchServerToolConfig — Configuration for the openrouter:web_fetch server tool
          - `allowed_domains` string[] — Only fetch from these domains.
          - `blocked_domains` string[] — Never fetch from these domains.
          - `engine` 'auto' | 'native' | 'openrouter' | 'exa' | 'parallel' | 'firecrawl' — Which fetch engine to use. "auto" (default) uses native if the provider supports it, otherwise Exa. "native" forces the provider's built-in fetch. "exa" uses Exa Contents API. "openrouter" uses direct HTTP fetch. "firecrawl" uses Firecrawl scrape (requires BYOK). "parallel" uses the Parallel extract API.
          - `max_content_tokens` integer — Maximum content length in approximate tokens. Content exceeding this limit is truncated.
          - `max_uses` integer — Maximum number of web fetches per request. Once exceeded, the tool returns an error.
        - `type` 'openrouter:web_fetch', required
      - OpenRouterWebSearchServerTool — OpenRouter built-in server tool: searches the web for current information
        - `parameters` WebSearchConfig
          - `allowed_domains` string[] — Limit search results to these domains. Supported by Exa, Firecrawl, Parallel, Perplexity, and most native providers (Anthropic, OpenAI, xAI). Cannot be used with excluded_domains.
          - `engine` 'native' | 'exa' | 'parallel' | 'firecrawl' | 'perplexity' | 'auto' — Which search engine to use. "auto" (default) uses native if the provider supports it, otherwise Exa. "native" forces the provider's built-in search. "exa" forces the Exa search API. "firecrawl" uses Firecrawl (requires BYOK). "parallel" uses the Parallel search API. "perplexity" uses the Perplexity Search API (raw ranked results).
          - `excluded_domains` string[] — Exclude search results from these domains. Supported by Exa, Firecrawl, Parallel, Perplexity, Anthropic, and xAI. Not supported with OpenAI (silently ignored). Cannot be used with allowed_domains.
          - `max_characters` integer — Exact maximum number of characters of content per search result. Applies to the Exa, Parallel, and Perplexity engines; ignored with native provider search and Firecrawl. For Exa, caps highlight content per result. For Parallel, caps excerpt content per result (default 1,500 when omitted). For Perplexity, maps to the native `max_tokens_per_page` parameter (converted from characters to tokens) and trims the response to the exact character cap. When both `max_characters` and `search_context_size` are set, `max_characters` takes precedence. When omitted, falls back to `search_context_size` mapping (Exa) or engine defaults (Parallel, Perplexity).
          - `max_results` integer — Maximum number of search results to return per search call. Defaults to 5. Applies to Exa, Firecrawl, Parallel, and Perplexity engines; ignored with native provider search. Perplexity supports a maximum of 20; values above 20 are clamped.
          - `max_total_results` integer — Maximum total number of search results across all search calls in a single request. Once this limit is reached, the tool will stop returning new results. Useful for controlling cost and context size in agentic loops. Defaults to 50 when not specified.
          - `search_context_size` 'low' | 'medium' | 'high' — How much context to retrieve per result. Applies to Exa, Parallel, and Perplexity engines; ignored with native provider search and Firecrawl. For Exa, pins a fixed per-result character cap (low=5,000, medium=15,000, high=30,000); when omitted, Exa picks an adaptive size per query and document (typically ~2,000–4,000 characters per result). For Parallel, controls the total characters across all results; when omitted, Parallel uses its own default size. For Perplexity, maps directly to the Search API's native search_context_size parameter. Overridden by `max_characters` when both are set.
          - `user_location` WebSearchUserLocationServerTool — Approximate user location for location-biased results.
            - `city` string, nullable
            - `country` string, nullable
            - `region` string, nullable
            - `timezone` string, nullable
            - `type` 'approximate'
        - `type` 'openrouter:web_search', required
      - ChatWebSearchShorthand — Web search tool using OpenAI Responses API syntax. Automatically converted to openrouter:web_search.
        - `allowed_domains` string[] — Limit search results to these domains. Supported by Exa, Firecrawl, Parallel, Perplexity, and most native providers (Anthropic, OpenAI, xAI). Cannot be used with excluded_domains.
        - `engine` 'native' | 'exa' | 'parallel' | 'firecrawl' | 'perplexity' | 'auto' — Which search engine to use. "auto" (default) uses native if the provider supports it, otherwise Exa. "native" forces the provider's built-in search. "exa" forces the Exa search API. "firecrawl" uses Firecrawl (requires BYOK). "parallel" uses the Parallel search API. "perplexity" uses the Perplexity Search API (raw ranked results).
        - `excluded_domains` string[] — Exclude search results from these domains. Supported by Exa, Firecrawl, Parallel, Perplexity, Anthropic, and xAI. Not supported with OpenAI (silently ignored). Cannot be used with allowed_domains.
        - `max_characters` integer — Exact maximum number of characters of content per search result. Applies to the Exa, Parallel, and Perplexity engines; ignored with native provider search and Firecrawl. For Exa, caps highlight content per result. For Parallel, caps excerpt content per result (default 1,500 when omitted). For Perplexity, maps to the native `max_tokens_per_page` parameter (converted from characters to tokens) and trims the response to the exact character cap. When both `max_characters` and `search_context_size` are set, `max_characters` takes precedence. When omitted, falls back to `search_context_size` mapping (Exa) or engine defaults (Parallel, Perplexity).
        - `max_results` integer — Maximum number of search results to return per search call. Defaults to 5. Applies to Exa, Firecrawl, Parallel, and Perplexity engines; ignored with native provider search. Perplexity supports a maximum of 20; values above 20 are clamped.
        - `max_total_results` integer — Maximum total number of search results across all search calls in a single request. Once this limit is reached, the tool will stop returning new results. Useful for controlling cost and context size in agentic loops. Defaults to 50 when not specified.
        - `parameters` WebSearchConfig
          - `allowed_domains` string[] — Limit search results to these domains. Supported by Exa, Firecrawl, Parallel, Perplexity, and most native providers (Anthropic, OpenAI, xAI). Cannot be used with excluded_domains.
          - `engine` 'native' | 'exa' | 'parallel' | 'firecrawl' | 'perplexity' | 'auto' — Which search engine to use. "auto" (default) uses native if the provider supports it, otherwise Exa. "native" forces the provider's built-in search. "exa" forces the Exa search API. "firecrawl" uses Firecrawl (requires BYOK). "parallel" uses the Parallel search API. "perplexity" uses the Perplexity Search API (raw ranked results).
          - `excluded_domains` string[] — Exclude search results from these domains. Supported by Exa, Firecrawl, Parallel, Perplexity, Anthropic, and xAI. Not supported with OpenAI (silently ignored). Cannot be used with allowed_domains.
          - `max_characters` integer — Exact maximum number of characters of content per search result. Applies to the Exa, Parallel, and Perplexity engines; ignored with native provider search and Firecrawl. For Exa, caps highlight content per result. For Parallel, caps excerpt content per result (default 1,500 when omitted). For Perplexity, maps to the native `max_tokens_per_page` parameter (converted from characters to tokens) and trims the response to the exact character cap. When both `max_characters` and `search_context_size` are set, `max_characters` takes precedence. When omitted, falls back to `search_context_size` mapping (Exa) or engine defaults (Parallel, Perplexity).
          - `max_results` integer — Maximum number of search results to return per search call. Defaults to 5. Applies to Exa, Firecrawl, Parallel, and Perplexity engines; ignored with native provider search. Perplexity supports a maximum of 20; values above 20 are clamped.
          - `max_total_results` integer — Maximum total number of search results across all search calls in a single request. Once this limit is reached, the tool will stop returning new results. Useful for controlling cost and context size in agentic loops. Defaults to 50 when not specified.
          - `search_context_size` 'low' | 'medium' | 'high' — How much context to retrieve per result. Applies to Exa, Parallel, and Perplexity engines; ignored with native provider search and Firecrawl. For Exa, pins a fixed per-result character cap (low=5,000, medium=15,000, high=30,000); when omitted, Exa picks an adaptive size per query and document (typically ~2,000–4,000 characters per result). For Parallel, controls the total characters across all results; when omitted, Parallel uses its own default size. For Perplexity, maps directly to the Search API's native search_context_size parameter. Overridden by `max_characters` when both are set.
          - `user_location` WebSearchUserLocationServerTool — Approximate user location for location-biased results.
            - `city` string, nullable
            - `country` string, nullable
            - `region` string, nullable
            - `timezone` string, nullable
            - `type` 'approximate'
        - `search_context_size` 'low' | 'medium' | 'high' — How much context to retrieve per result. Applies to Exa, Parallel, and Perplexity engines; ignored with native provider search and Firecrawl. For Exa, pins a fixed per-result character cap (low=5,000, medium=15,000, high=30,000); when omitted, Exa picks an adaptive size per query and document (typically ~2,000–4,000 characters per result). For Parallel, controls the total characters across all results; when omitted, Parallel uses its own default size. For Perplexity, maps directly to the Search API's native search_context_size parameter. Overridden by `max_characters` when both are set.
        - `type` 'web_search' | 'web_search_preview' | 'web_search_preview_2025_03_11' | 'web_search_2025_08_26', required
        - `user_location` WebSearchUserLocationServerTool — Approximate user location for location-biased results.
          - `city` string, nullable
          - `country` string, nullable
          - `region` string, nullable
          - `timezone` string, nullable
          - `type` 'approximate'
  - `top_a` number, double, nullable — Consider only tokens with "sufficiently high" probabilities based on the probability of the most likely token. Not all providers support this parameter.
  - `top_k` integer, nullable — Limits the model to choose from the top K most likely tokens at each step. A value of 1 means the model will always pick the most likely next token. Not all providers support this parameter.
  - `top_logprobs` integer, nullable — Number of top log probabilities to return (0-20)
  - `top_p` number, double, nullable — Nucleus sampling parameter (0-1)
  - `trace` TraceConfig — Metadata for observability and tracing. Known keys (trace_id, trace_name, span_name, generation_name, parent_span_id) have special handling. Additional keys are passed through as custom metadata to configured broadcast destinations.
    - `generation_name` string
    - `parent_span_id` string
    - `span_name` string
    - `trace_id` string
    - `trace_name` string
  - `user` string — Unique user identifier

## Response `200`

Successful chat completion response

- ChatResult — Chat completion response
  - `choices` ChatChoice[], required — List of completion choices
    - `finish_reason` 'tool_calls' | 'stop' | 'length' | 'content_filter' | 'error' | 'null', nullable, required
    - `index` integer, required — Choice index
    - `logprobs` ChatTokenLogprobs, nullable — Log probabilities for the completion
      - `content` ChatTokenLogprob[], nullable, required — Log probabilities for content tokens
        - `bytes` integer[], nullable, required — UTF-8 bytes of the token
        - `logprob` number, double, required — Log probability of the token
        - `token` string, required — The token
        - `top_logprobs` object[], required — Top alternative tokens with probabilities
          - `bytes` integer[], nullable, required
          - `logprob` number, double, required
          - `token` string, required
      - `refusal` ChatTokenLogprob[], nullable — Log probabilities for refusal tokens
        - `bytes` integer[], nullable, required — UTF-8 bytes of the token
        - `logprob` number, double, required — Log probability of the token
        - `token` string, required — The token
        - `top_logprobs` object[], required — Top alternative tokens with probabilities
          - `bytes` integer[], nullable, required
          - `logprob` number, double, required
          - `token` string, required
    - `message` ChatAssistantMessage, required — Assistant message for requests and responses
      - `audio` ChatAudioOutput — Audio output data or reference
        - `data` string — Base64 encoded audio data
        - `expires_at` integer — Audio expiration timestamp
        - `id` string — Audio output identifier
        - `transcript` string — Audio transcript
      - `content` union — Assistant message content
        - string
        - ChatContentItems[]
          - union — Content part for chat completion messages
            - ChatContentText — Text content part
              - …
            - ChatContentImage — Image content part for vision models
              - …
            - ChatContentAudio — Audio input content part. Supported audio formats vary by provider.
              - …
            - LegacyChatContentVideo — Video input content part (legacy format - deprecated)
              - …
            - ChatContentVideo — Video input content part
              - …
            - ChatContentFile — File content part for document processing
              - …
        - unknown
      - `images` object[] — Generated images from image generation models
        - `image_url` object, required
          - `url` string, required — URL or base64-encoded data of the generated image
      - `name` string — Optional name for the assistant
      - `reasoning` string, nullable — Reasoning output
      - `reasoning_details` ReasoningDetailUnion[] — Reasoning details for extended thinking models
        - union — Reasoning detail union schema
          - ReasoningDetailSummary — Reasoning detail summary schema
            - `format` 'unknown' | 'openai-responses-v1' | 'azure-openai-responses-v1' | 'xai-responses-v1' | 'anthropic-claude-v1' | 'google-gemini-v1' | 'null', nullable
            - `id` string, nullable
            - `index` integer
            - `summary` string, required
            - `type` 'reasoning.summary', required
          - ReasoningDetailEncrypted — Reasoning detail encrypted schema
            - `data` string, required
            - `format` 'unknown' | 'openai-responses-v1' | 'azure-openai-responses-v1' | 'xai-responses-v1' | 'anthropic-claude-v1' | 'google-gemini-v1' | 'null', nullable
            - `id` string, nullable
            - `index` integer
            - `type` 'reasoning.encrypted', required
          - ReasoningDetailText — Reasoning detail text schema
            - `format` 'unknown' | 'openai-responses-v1' | 'azure-openai-responses-v1' | 'xai-responses-v1' | 'anthropic-claude-v1' | 'google-gemini-v1' | 'null', nullable
            - `id` string, nullable
            - `index` integer
            - `signature` string, nullable
            - `text` string, nullable
            - `type` 'reasoning.text', required
          - ReasoningDetailServerToolCall — Record of an OpenRouter server-tool invocation (e.g. openrouter:fusion), carried in reasoning_details so a prior tool call can be rehydrated into a later turn of the same conversation.
            - `arguments` string, required
            - `format` 'unknown' | 'openai-responses-v1' | 'azure-openai-responses-v1' | 'xai-responses-v1' | 'anthropic-claude-v1' | 'google-gemini-v1' | 'null', nullable
            - `id` string, nullable
            - `index` integer
            - `result` string, required
            - `tool_call_id` string, nullable
            - `tool_name` string, required
            - `type` 'reasoning.server_tool_call', required
      - `refusal` string, nullable — Refusal message if content was refused
      - `role` 'assistant', required
      - `tool_calls` ChatToolCall[] — Tool calls made by the assistant
        - `function` object, required
          - `arguments` string, required — Function arguments as JSON string
          - `name` string, required — Function name to call
        - `id` string, required — Tool call identifier
        - `type` 'function', required
  - `created` integer, required — Unix timestamp of creation
  - `id` string, required — Unique completion identifier
  - `model` string, required — Model used for completion
  - `object` 'chat.completion', required
  - `openrouter_metadata` OpenRouterMetadata
    - `attempt` integer, required
    - `attempts` RouterAttempt[]
      - `model` string, required
      - `provider` string, required
      - `status` integer, required
    - `endpoints` EndpointsMetadata, required
      - `available` EndpointInfo[], required
        - `model` string, required
        - `provider` string, required
        - `selected` boolean, required
      - `total` integer, required
    - `is_byok` boolean, required
    - `params` RouterParams
      - `quality_floor` number, double
      - `throughput_floor` number, double
      - `version_group` string
    - `pipeline` PipelineStage[]
      - `cost_usd` number, double, nullable
      - `data` object
      - `guardrail_id` string
      - `guardrail_scope` string
      - `name` string, required
      - `summary` string
      - `type` 'guardrail' | 'plugin' | 'server_tools' | 'response_healing' | 'context_compression', required — Categorical kind of a pipeline stage. Multiple plugins can share a type (e.g. all guardrail-level plugins emit `guardrail`); the `name` field disambiguates which plugin emitted it.
    - `region` string, nullable, required
    - `requested` string, required
    - `strategy` 'direct' | 'auto' | 'free' | 'latest' | 'alias' | 'fallback' | 'pareto' | 'bodybuilder' | 'fusion', required
    - `summary` string, required
  - `service_tier` string, nullable — The service tier used by the upstream provider for this request
  - `system_fingerprint` string, nullable, required — System fingerprint
  - `usage` ChatUsage — Token usage statistics
    - `completion_tokens` integer, required — Number of tokens in the completion
    - `completion_tokens_details` object, nullable — Detailed completion token usage
      - `accepted_prediction_tokens` integer, nullable — Accepted prediction tokens
      - `audio_tokens` integer, nullable — Tokens used for audio output
      - `reasoning_tokens` integer, nullable — Tokens used for reasoning
      - `rejected_prediction_tokens` integer, nullable — Rejected prediction tokens
    - `cost` number, double, nullable — Cost of the completion
    - `cost_details` CostDetails, nullable — Breakdown of upstream inference costs
      - `upstream_inference_completions_cost` number, double, required
      - `upstream_inference_cost` number, double, nullable
      - `upstream_inference_prompt_cost` number, double, required
    - `is_byok` boolean — Whether a request was made using a Bring Your Own Key configuration
    - `prompt_tokens` integer, required — Number of tokens in the prompt
    - `prompt_tokens_details` object, nullable — Detailed prompt token usage
      - `audio_tokens` integer — Audio input tokens
      - `cache_write_tokens` integer — Tokens written to cache. Only returned for models with explicit caching and cache write pricing.
      - `cached_tokens` integer — Cached prompt tokens
      - `video_tokens` integer — Video input tokens
    - `server_tool_use_details` ServerToolUseDetails, nullable — Usage for server-side tool execution (e.g., web search)
      - `tool_calls_executed` integer, nullable — Number of OpenRouter server tool calls that executed and produced a result.
      - `tool_calls_requested` integer, nullable — Total number of OpenRouter server-orchestrated tool calls the model requested, across all tool types. Provider-native tools (e.g. native web search) are not counted here.
      - `web_search_requests` integer, nullable — Number of web searches performed by server-side tools. For server-orchestrated tool calls a web search is also counted in tool_calls_requested; provider-native web search may report web_search_requests only. Do not sum the two.
    - `total_tokens` integer, required — Total number of tokens

## Other responses

- `400` — Bad Request - Invalid request parameters or malformed input
- `401` — Unauthorized - Authentication required or invalid credentials
- `402` — Payment Required - Insufficient credits or quota to complete request
- `403` — Forbidden - Authentication successful but insufficient permissions, or a guardrail blocked the request. When guardrails block and the `X-OpenRouter-Metadata: enabled` header is present, the response includes `openrouter_metadata` with full routing context and a `pipeline` array containing guardrail stage details.
- `404` — Not Found - Resource does not exist
- `408` — Request Timeout - Operation exceeded time limit
- `413` — Payload Too Large - Request payload exceeds size limits
- `422` — Unprocessable Entity - Semantic validation failure
- `429` — Too Many Requests - Rate limit exceeded
- `500` — Internal Server Error - Unexpected server error
- `502` — Bad Gateway - Provider/upstream API failure
- `503` — Service Unavailable - Service temporarily unavailable
- `524` — Infrastructure Timeout - Provider request timed out at edge network
- `529` — Provider Overloaded - Provider is temporarily overloaded

## Changes

- **2026-07-15** `ada7dbde69ed` — 3 info
  - added the new `Krea` enum value to the request property `provider/ignore/items/anyOf[#/components/schemas/ProviderName]/`
  - added the new `Krea` enum value to the request property `provider/only/items/anyOf[#/components/schemas/ProviderName]/`
  - added the new `Krea` enum value to the request property `provider/order/items/anyOf[#/components/schemas/ProviderName]/`
- **2026-07-15** `becc20273345` — 1 info
  - added the new optional request property `prediction`
- **2026-07-14** `8b656727d0a3` — 3 info
  - added the new `Sail Research` enum value to the request property `provider/ignore/items/anyOf[#/components/schemas/ProviderName]/`
  - added the new `Sail Research` enum value to the request property `provider/only/items/anyOf[#/components/schemas/ProviderName]/`
  - added the new `Sail Research` enum value to the request property `provider/order/items/anyOf[#/components/schemas/ProviderName]/`
- …earlier changes not shown

[Full history](https://skmtc.dev/openrouterteam/apis/openrouter-api/changes/chat/completions/post.md)

---

[API](https://skmtc.dev/openrouterteam/apis/openrouter-api.md) · [All operations](https://skmtc.dev/openrouterteam/apis/openrouter-api/llms.txt) · [OpenAPI document](https://skmtc.dev/openrouterteam/apis/openrouter-api/revisions/ada7dbde69ed?raw)
