Audio

Generate a sound effect from a text prompt

Generates a non-speech audio clip - a sound effect or ambience - from a text prompt. Gateway extension: OpenAI has no sound-effects endpoint, so this mirrors the shape of POST /audio/speech (JSON in, raw audio bytes out) rather than an upstream OpenAI operation. The response is the generated audio as raw binary bytes; the actual Content-Type header of the response reflects the requested response_format (e.g. audio/mpeg for mp3).

Not every provider implements sound-effect generation. Requests routed to a provider that does not support it return 400 Bad Request with an explanatory error message.

post/audio/sfx

Query parameters

provider'ollama' | 'ollama_cloud' | 'groq' | 'llamacpp' | 'openai' | 'cloudflare' | 'cohere' | 'anthropic' | 'deepseek' | 'elevenlabs' | 'google' | 'mistral' | 'minimax' | 'moonshot' | 'nvidia' | 'zai'

Specific provider to use (default determined by model)

Request body

modelstring required

Model ID to use for sound-effect generation (e.g. elevenlabs/eleven_text_to_sound_v2).

promptstring required

Description of the sound to generate (e.g. distant thunder rolling over a valley).

duration_secondsnumber

Length of the generated clip in seconds. Omit to let the provider pick a length that fits the prompt.

prompt_influencenumber

How closely the generation follows the prompt. Higher values stay closer to the prompt, lower values allow more variation. Omit to use the provider default.

loopboolean

Whether to generate a clip that loops seamlessly.

response_format'mp3' | 'opus' | 'aac' | 'flac' | 'pcm'

The audio format of the response.

Response

The generated audio as raw binary bytes. The actual Content-Type header of the response reflects the requested response_format.

Changes

Changed in 2 of the 77 revisions of this API.11