v1 inference (deprecated)

[Deprecated] Multimodal completion with optional audio/image output

DEPRECATED — Use the new dedicated endpoints instead: - Text → POST /v1/text/generations - Image → POST /v1/images/generations - Audio → POST /v1/audio/generations - All → POST /v1/chat/completions

This endpoint is kept for backward compatibility and internally delegates to
the appropriate new generation handler based on `output_type`.
post/v1/inference/multimodality-completion

Request body

user_idstring required

Unique identifier for the user (always required)

persona_idstring nullable

Optional persona ID. If provided, inference uses this persona's context instead of the user

project_idstring nullable

Optional project ID. If provided, inference uses project context (inherits from user)

questionstring nullable

User's question or prompt (optional if audio provided)

contextstring nullable

Additional context for the question

session_idstring nullable

Optional session identifier for conversation context

video_base64string nullable

Base64 encoded video content

image_base64string nullable

Base64 encoded image content

audio_base64string nullable

Base64 encoded audio content (supports webm, wav, mp3, mp4, and other formats)

audio_type'tts' | 'music' | 'sfx'

Type of audio output to generate

voicestring

Voice to use for TTS (alloy, echo, fable, onyx, nova, shimmer). Only used when audio_type='tts'

speednumber

Speed of the speech (0.25 to 4.0). Only used when audio_type='tts'

audio_durationnumber nullable

Duration in seconds for music/sfx generation. Default: 5s for sfx, 10s for music

modelstring nullable

LLM model to use for generating the response

output_type'text' | 'audio' | 'image'
disabled_learningboolean

Whether to disable learning/ingestion of the multimodal content

use_reasoningboolean

Use creative reasoning loop for constraint-satisfying generation (only for creative_design projects with output_type='image')

max_reasoning_iterationsinteger

Maximum repair iterations in reasoning loop

num_imagesinteger

Number of images to generate (each with a different seed for variation). Only used when output_type='image'

seedinteger nullable

Base seed for reproducible image generation. If not provided, a random seed is used. Only used when output_type='image'

Example request

{
  "audio_base64": "base64_encoded_audio...",
  "audio_type": "tts",
  "disabled_learning": false,
  "image_base64": "base64_encoded_image...",
  "model": "gemini-2.5-flash",
  "output_type": "text",
  "question": "What do you see?",
  "session_id": "session_123",
  "speed": 1,
  "user_id": "123e4567-e89b-12d3-a456-426614174000",
  "video_base64": "base64_encoded_video...",
  "voice": "alloy"
}

Response

Successful Response

text_responsestring required

Generated text response from LLM

audio_base64string nullable

Base64 encoded audio of the response (if output_type='audio')

audio_formatstring nullable

Format of the audio (e.g., mp3)

voice_usedstring nullable

Voice used for TTS

image_base64string nullable

Base64 encoded AI-generated image (if output_type='image'). First image when num_images > 1

generated_imagesstring[] nullable

List of all generated images as base64 (when num_images > 1). Each image is generated with a different seed for variation

memory_contextstring nullable

Formatted memory context used for generating the response

raw_resultsobject nullable

Raw results from memory retrieval

entity_imagesobject nullable

Reference images for matched entities used in generation (entity_name -> base64 image)

reasoning_traceobject nullable

Complete reasoning trace (if use_reasoning=True) with blueprint, grounding, constraints, verification, and repair steps

successboolean

Whether the request was successful

Example response

{
  "audio_base64": "base64_encoded_audio...",
  "audio_format": "mp3",
  "image_base64": "base64_encoded_image...",
  "memory_context": "Retrieved memories: ...",
  "raw_results": {
    "episodes": [],
    "preferences": []
  },
  "success": true,
  "text_response": "Based on your preferences, you prefer to wake up early...",
  "voice_used": "alloy"
}

Changes