---
title: "Get Squad"
method: GET
path: "/squad/{id}"
tags: ["Squads"]
---

# Get Squad

`GET /squad/{id}`

## Path parameters

- `id` string, uuid, required

## Response `200`

- Squad
  - `name` string — This is the name of the squad.
  - `members` SquadMemberDTO[], required — This is the list of assistants that make up the squad. The call will start with the first assistant in the list.
    - `assistantVersion` string, nullable — This is the assistant version (e.g. `v3`) to pin for this squad member. When set, the call uses the snapshot from `assistant_version` (by `(assistantId, version)`) instead of the latest. Valid only with `assistantId`; rejected with inline `assistant`. Omit to follow the latest version.
    - `assistantDestinations` union[]
      - union
        - TransferDestinationAssistant
          - `message` union — This is spoken to the customer before connecting them to the destination. Usage: - If this is not provided and transfer tool messages is not provided, default is "Transferring the call now". - If set to "", nothing is spoken. This is useful when you want to silently transfer. This is especially useful when transferring between assistants in a squad. In this scenario, you likely also want to set `assistant.firstMessageMode=assistant-speaks-first-with-model-generated-message` for the destination assistant. This accepts a string or a ToolMessageStart class. Latter is useful if you want to specify multiple messages for different languages through the `contents` field.
            - string
            - CustomMessage
              - …
          - `type` 'assistant', required
          - `transferMode` 'rolling-history' | 'swap-system-message-in-history' | 'swap-system-message-in-history-and-remove-transfer-tool-messages' | 'delete-history' — This is the mode to use for the transfer. Defaults to `rolling-history`. - `rolling-history`: This is the default mode. It keeps the entire conversation history and appends the new assistant's system message on transfer. Example: Pre-transfer: system: assistant1 system message assistant: assistant1 first message user: hey, good morning assistant: how can i help? user: i need help with my account assistant: (destination.message) Post-transfer: system: assistant1 system message assistant: assistant1 first message user: hey, good morning assistant: how can i help? user: i need help with my account assistant: (destination.message) system: assistant2 system message assistant: assistant2 first message (or model generated if firstMessageMode is set to `assistant-speaks-first-with-model-generated-message`) - `swap-system-message-in-history`: This replaces the original system message with the new assistant's system message on transfer. Example: Pre-transfer: system: assistant1 system message assistant: assistant1 first message user: hey, good morning assistant: how can i help? user: i need help with my account assistant: (destination.message) Post-transfer: system: assistant2 system message assistant: assistant1 first message user: hey, good morning assistant: how can i help? user: i need help with my account assistant: (destination.message) assistant: assistant2 first message (or model generated if firstMessageMode is set to `assistant-speaks-first-with-model-generated-message`) - `delete-history`: This deletes the entire conversation history on transfer. Example: Pre-transfer: system: assistant1 system message assistant: assistant1 first message user: hey, good morning assistant: how can i help? user: i need help with my account assistant: (destination.message) Post-transfer: system: assistant2 system message assistant: assistant2 first message user: Yes, please assistant: how can i help? user: i need help with my account - `swap-system-message-in-history-and-remove-transfer-tool-messages`: This replaces the original system message with the new assistant's system message on transfer and removes transfer tool messages from conversation history sent to the LLM. Example: Pre-transfer: system: assistant1 system message assistant: assistant1 first message user: hey, good morning assistant: how can i help? user: i need help with my account transfer-tool transfer-tool-result assistant: (destination.message) Post-transfer: system: assistant2 system message assistant: assistant1 first message user: hey, good morning assistant: how can i help? user: i need help with my account assistant: (destination.message) assistant: assistant2 first message (or model generated if firstMessageMode is set to `assistant-speaks-first-with-model-generated-message`) @default 'rolling-history'
          - `assistantName` string, required — This is the assistant to transfer the call to.
          - `name` string — This is the name of the transfer destination. This is just for your own reference. Usage: - Optional. Stored with the destination wherever it is supplied. For `number` and `sip` destinations it is also persisted on the transfer record in the call artifact after a transfer and displayed in the dashboard call log (on the transfer divider in the transcript view) alongside the destination. When omitted, everything behaves exactly as before. - Display-only. Unlike `description`, it is never included in prompts or tool descriptions and has no effect on model behavior or destination choice.
          - `description` string — This is the description of the destination, used by the AI to choose when and how to transfer the call.
        - HandoffDestinationAssistant
          - `type` 'assistant', required
          - `contextEngineeringPlan` union — This is the plan for manipulating the message context before handing off the call to the next assistant.
            - ContextEngineeringPlanLastNMessages
              - …
            - ContextEngineeringPlanNone
              - …
            - ContextEngineeringPlanAll
              - …
            - ContextEngineeringPlanUserAndAssistantMessages
              - …
            - ContextEngineeringPlanPreviousAssistantMessages
              - …
          - `assistantName` string — This is the assistant to transfer the call to. You must provide either assistantName or assistantId.
          - `assistantId` string — This is the assistant id to transfer the call to. You must provide either assistantName or assistantId.
          - `assistant` CreateAssistantDTO
            - `transcriber` union — These are the options for the assistant's transcriber.
              - …
            - `model` union — These are the options for the assistant's LLM.
              - …
            - `voice` union — These are the options for the assistant's voice.
              - …
            - `firstMessage` string — This is the first message that the assistant will say. This can also be a URL to a containerized audio file (mp3, wav, etc.). If unspecified, assistant will wait for user to speak and use the model to respond once they speak.
            - `firstMessageInterruptionsEnabled` boolean
            - `firstMessageMode` 'assistant-speaks-first' | 'assistant-speaks-first-with-model-generated-message' | 'assistant-waits-for-user' — This is the mode for the first message. Default is 'assistant-speaks-first'. Use: - 'assistant-speaks-first' to have the assistant speak first. - 'assistant-waits-for-user' to have the assistant wait for the user to speak first. - 'assistant-speaks-first-with-model-generated-message' to have the assistant speak first with a message generated by the model based on the conversation state. (`assistant.model.messages` at call start, `call.messages` at squad transfer points). @default 'assistant-speaks-first'
            - `voicemailDetection` union — These are the settings to configure or disable voicemail detection. Alternatively, voicemail detection can be configured using the model.tools=[VoicemailTool]. By default, voicemail detection is disabled.
              - …
            - `clientMessages` string[] — These are the messages that will be sent to your Client SDKs. Default is conversation-update,function-call,hang,model-output,speech-update,status-update,transfer-update,transcript,tool-calls,user-interrupted,voice-input,workflow.node.started,assistant.started. You can check the shape of the messages in ClientMessage schema.
            - `serverMessages` string[] — These are the messages that will be sent to your Server URL. Default is conversation-update,end-of-call-report,function-call,hang,speech-update,status-update,tool-calls,transfer-destination-request,handoff-destination-request,user-interrupted,assistant.started. You can check the shape of the messages in ServerMessage schema.
            - `maxDurationSeconds` number — This is the maximum number of seconds that the call will last. When the call reaches this duration, it will be ended. @default 600 (10 minutes)
            - `backgroundSound` union — This is the background sound in the call. Default for phone calls is 'office' and default for web calls is 'off'. You can also provide a custom sound by providing a URL to an audio file.
              - …
            - `modelOutputInMessagesEnabled` boolean — This determines whether the model's output is used in conversation history rather than the transcription of assistant's speech. @default false
            - `transportConfigurations` TransportConfigurationTwilio[] — These are the configurations to be passed to the transport providers of assistant's calls, like Twilio. You can store multiple configurations for different transport providers. For a call, only the configuration matching the call transport provider is used.
              - …
            - `observabilityPlan` LangfuseObservabilityPlan
              - …
            - `credentials` union[] — These are dynamic credentials that will be used for the assistant calls. By default, all the credentials are available for use in the call but you can supplement an additional credentials using this. Dynamic credentials override existing credentials.
              - …
            - `hooks` union[] — This is a set of actions that will be performed on certain events.
              - …
            - `name` string — This is the name of the assistant. This is required when you want to transfer between assistants in a call.
            - `voicemailMessage` string — This is the message that the assistant will say if the call is forwarded to voicemail. If unspecified, it will hang up.
            - `endCallMessage` string — This is the message that the assistant will say if it ends the call. If unspecified, it will hang up without saying anything.
            - `endCallPhrases` string[] — This list contains phrases that, if spoken by the assistant, will trigger the call to be hung up. Case insensitive.
            - `compliancePlan` CompliancePlan
              - …
            - `metadata` object — This is for metadata you want to store on the assistant.
            - `backgroundSpeechDenoisingPlan` BackgroundSpeechDenoisingPlan
              - …
            - `analysisPlan` AnalysisPlan
              - …
            - `artifactPlan` ArtifactPlan
              - …
            - `startSpeakingPlan` StartSpeakingPlan
              - …
            - `stopSpeakingPlan` StopSpeakingPlan
              - …
            - `monitorPlan` MonitorPlan
              - …
            - `credentialIds` string[] — These are the credentials that will be used for the assistant calls. By default, all the credentials are available for use in the call but you can provide a subset using this.
            - `server` Server
              - …
            - `keypadInputPlan` KeypadInputPlan
              - …
          - `variableExtractionPlan` VariableExtractionPlan
            - `schema` JsonSchema
              - …
            - `aliases` VariableExtractionAlias[] — These are additional variables to create. These will be accessible during the call as `{{key}}` and stored in `call.artifact.variableValues` after the call. Example: ```json { "aliases": [ { "key": "customerName", "value": "{{name}}" }, { "key": "fullName", "value": "{{firstName}} {{lastName}}" }, { "key": "greeting", "value": "Hello {{name}}, welcome to {{company}}!" }, { "key": "customerCity", "value": "{{addresses[0].city}}" }, { "key": "something", "value": "{{any liquid}}" } ] } ``` This will create variables `customerName`, `fullName`, `greeting`, `customerCity`, and `something`. To access these variables, you can reference them as `{{customerName}}`, `{{fullName}}`, `{{greeting}}`, `{{customerCity}}`, and `{{something}}`.
              - …
          - `assistantOverrides` AssistantOverrides
            - `transcriber` union — These are the options for the assistant's transcriber.
              - …
            - `model` union — These are the options for the assistant's LLM.
              - …
            - `voice` union — These are the options for the assistant's voice.
              - …
            - `firstMessage` string — This is the first message that the assistant will say. This can also be a URL to a containerized audio file (mp3, wav, etc.). If unspecified, assistant will wait for user to speak and use the model to respond once they speak.
            - `firstMessageInterruptionsEnabled` boolean
            - `firstMessageMode` 'assistant-speaks-first' | 'assistant-speaks-first-with-model-generated-message' | 'assistant-waits-for-user' — This is the mode for the first message. Default is 'assistant-speaks-first'. Use: - 'assistant-speaks-first' to have the assistant speak first. - 'assistant-waits-for-user' to have the assistant wait for the user to speak first. - 'assistant-speaks-first-with-model-generated-message' to have the assistant speak first with a message generated by the model based on the conversation state. (`assistant.model.messages` at call start, `call.messages` at squad transfer points). @default 'assistant-speaks-first'
            - `voicemailDetection` union — These are the settings to configure or disable voicemail detection. Alternatively, voicemail detection can be configured using the model.tools=[VoicemailTool]. By default, voicemail detection is disabled.
              - …
            - `clientMessages` string[] — These are the messages that will be sent to your Client SDKs. Default is conversation-update,function-call,hang,model-output,speech-update,status-update,transfer-update,transcript,tool-calls,user-interrupted,voice-input,workflow.node.started,assistant.started. You can check the shape of the messages in ClientMessage schema.
            - `serverMessages` string[] — These are the messages that will be sent to your Server URL. Default is conversation-update,end-of-call-report,function-call,hang,speech-update,status-update,tool-calls,transfer-destination-request,handoff-destination-request,user-interrupted,assistant.started. You can check the shape of the messages in ServerMessage schema.
            - `maxDurationSeconds` number — This is the maximum number of seconds that the call will last. When the call reaches this duration, it will be ended. @default 600 (10 minutes)
            - `backgroundSound` union — This is the background sound in the call. Default for phone calls is 'office' and default for web calls is 'off'. You can also provide a custom sound by providing a URL to an audio file.
              - …
            - `modelOutputInMessagesEnabled` boolean — This determines whether the model's output is used in conversation history rather than the transcription of assistant's speech. @default false
            - `transportConfigurations` TransportConfigurationTwilio[] — These are the configurations to be passed to the transport providers of assistant's calls, like Twilio. You can store multiple configurations for different transport providers. For a call, only the configuration matching the call transport provider is used.
              - …
            - `observabilityPlan` LangfuseObservabilityPlan
              - …
            - `credentials` union[] — These are dynamic credentials that will be used for the assistant calls. By default, all the credentials are available for use in the call but you can supplement an additional credentials using this. Dynamic credentials override existing credentials.
              - …
            - `hooks` union[] — This is a set of actions that will be performed on certain events.
              - …
            - `tools:append` union[]
              - …
            - `variableValues` object — These are values that will be used to replace the template variables in the assistant messages and other text-based fields. This uses LiquidJS syntax. https://liquidjs.com/tutorials/intro-to-liquid.html So for example, `{{ name }}` will be replaced with the value of `name` in `variableValues`. `{{"now" | date: "%b %d, %Y, %I:%M %p", "America/New_York"}}` will be replaced with the current date and time in New York. Some VAPI reserved defaults: - *customer* - the customer object
            - `name` string — This is the name of the assistant. This is required when you want to transfer between assistants in a call.
            - `voicemailMessage` string — This is the message that the assistant will say if the call is forwarded to voicemail. If unspecified, it will hang up.
            - `endCallMessage` string — This is the message that the assistant will say if it ends the call. If unspecified, it will hang up without saying anything.
            - `endCallPhrases` string[] — This list contains phrases that, if spoken by the assistant, will trigger the call to be hung up. Case insensitive.
            - `compliancePlan` CompliancePlan
              - …
            - `metadata` object — This is for metadata you want to store on the assistant.
            - `backgroundSpeechDenoisingPlan` BackgroundSpeechDenoisingPlan
              - …
            - `analysisPlan` AnalysisPlan
              - …
            - `artifactPlan` ArtifactPlan
              - …
            - `startSpeakingPlan` StartSpeakingPlan
              - …
            - `stopSpeakingPlan` StopSpeakingPlan
              - …
            - `monitorPlan` MonitorPlan
              - …
            - `credentialIds` string[] — These are the credentials that will be used for the assistant calls. By default, all the credentials are available for use in the call but you can provide a subset using this.
            - `server` Server
              - …
            - `keypadInputPlan` KeypadInputPlan
              - …
          - `description` string — This is the description of the destination, used by the AI to choose when and how to transfer the call.
    - `assistantId` string, nullable — This is the assistant that will be used for the call. To use a transient assistant, use `assistant` instead.
    - `assistant` CreateAssistantDTO
      - `transcriber` union — These are the options for the assistant's transcriber.
        - AssemblyAITranscriber
          - `provider` 'assembly-ai', required — This is the transcription provider that will be used.
          - `language` 'multi' | 'en' — This is the language that will be set for the transcription.
          - `confidenceThreshold` number — Transcripts below this confidence threshold will be discarded. @default 0.4
          - `formatTurns` boolean — This enables formatting of transcripts. @default true
          - `endOfTurnConfidenceThreshold` number — This is the end of turn confidence threshold. The minimum confidence that the end of turn is detected. Note: Only used if startSpeakingPlan.smartEndpointingPlan is not set. @min 0 @max 1 @default 0.7
          - `minEndOfTurnSilenceWhenConfident` number — This is the minimum end of turn silence when confident in milliseconds. Note: Only used if startSpeakingPlan.smartEndpointingPlan is not set. @default 160
          - `wordFinalizationMaxWaitTime` number
          - `maxTurnSilence` number — This is the maximum turn silence time in milliseconds. Note: Only used if startSpeakingPlan.smartEndpointingPlan is not set. @default 400
          - `vadAssistedEndpointingEnabled` boolean — Use VAD to assist with endpointing decisions from the transcriber. When enabled, transcriber endpointing will be buffered if VAD detects the user is still speaking, preventing premature turn-taking. When disabled, transcriber endpointing will be used immediately regardless of VAD state, allowing for quicker but more aggressive turn-taking. Note: Only used if startSpeakingPlan.smartEndpointingPlan is not set. @default true
          - `mode` 'max_accuracy' | 'min_latency' | 'balanced' — This is the transcription mode used by the `universal-3-5-pro` speech model. Only applies to the `universal-3-5-pro` speech model. @default 'balanced'
          - `prompt` string — This is a prompt that provides additional context to the transcription model. Only applies to the `universal-3-5-pro` speech model.
          - `agentContext` string — This is context about the voice agent that guides the transcription model. Only applies to the `universal-3-5-pro` speech model.
          - `languageCodes` string[] — These are language codes used to steer automatic language detection. Only applies to the `universal-3-5-pro` speech model.
          - `speechModel` 'universal-streaming-english' | 'universal-streaming-multilingual' | 'universal-3-5-pro' — This is the speech model used for the streaming session. Keyterms prompting is supported on universal-streaming-english and universal-3-5-pro. universal-3-5-pro is AssemblyAI's most accurate voice-agent model. @default 'universal-streaming-english'
          - `realtimeUrl` string — The WebSocket URL that the transcriber connects to.
          - `wordBoost` string[] — Add up to 2500 characters of custom vocabulary.
          - `keytermsPrompt` string[] — Keyterms prompting improves recognition accuracy for specific words and phrases. Can include up to 100 keyterms, each up to 50 characters. Costs an additional $0.04/hour on universal-streaming-english and is included at no extra cost on universal-3-5-pro.
          - `endUtteranceSilenceThreshold` number — The duration of the end utterance silence threshold in milliseconds.
          - `disablePartialTranscripts` boolean — Disable partial transcripts. Set to `true` to not receive partial transcripts. Defaults to `false`.
          - `fallbackPlan` FallbackTranscriberPlan
            - `transcribers` union[]
              - …
        - AzureSpeechTranscriber
          - `provider` 'azure', required — This is the transcription provider that will be used.
          - `language` 'af-ZA' | 'am-ET' | 'ar-AE' | 'ar-BH' | 'ar-DZ' | 'ar-EG' | 'ar-IL' | 'ar-IQ' | 'ar-JO' | 'ar-KW' | 'ar-LB' | 'ar-LY' | 'ar-MA' | 'ar-OM' | 'ar-PS' | 'ar-QA' | 'ar-SA' | 'ar-SY' | 'ar-TN' | 'ar-YE' | 'az-AZ' | 'bg-BG' | 'bn-IN' | 'bs-BA' | 'ca-ES' | 'cs-CZ' | 'cy-GB' | 'da-DK' | 'de-AT' | 'de-CH' | 'de-DE' | 'el-GR' | 'en-AU' | 'en-CA' | 'en-GB' | 'en-GH' | 'en-HK' | 'en-IE' | 'en-IN' | 'en-KE' | 'en-NG' | 'en-NZ' | 'en-PH' | 'en-SG' | 'en-TZ' | 'en-US' | 'en-ZA' | 'es-AR' | 'es-BO' | 'es-CL' | 'es-CO' | 'es-CR' | 'es-CU' | 'es-DO' | 'es-EC' | 'es-ES' | 'es-GQ' | 'es-GT' | 'es-HN' | 'es-MX' | 'es-NI' | 'es-PA' | 'es-PE' | 'es-PR' | 'es-PY' | 'es-SV' | 'es-US' | 'es-UY' | 'es-VE' | 'et-EE' | 'eu-ES' | 'fa-IR' | 'fi-FI' | 'fil-PH' | 'fr-BE' | 'fr-CA' | 'fr-CH' | 'fr-FR' | 'ga-IE' | 'gl-ES' | 'gu-IN' | 'he-IL' | 'hi-IN' | 'hr-HR' | 'hu-HU' | 'hy-AM' | 'id-ID' | 'is-IS' | 'it-CH' | 'it-IT' | 'ja-JP' | 'jv-ID' | 'ka-GE' | 'kk-KZ' | 'km-KH' | 'kn-IN' | 'ko-KR' | 'lo-LA' | 'lt-LT' | 'lv-LV' | 'mk-MK' | 'ml-IN' | 'mn-MN' | 'mr-IN' | 'ms-MY' | 'mt-MT' | 'my-MM' | 'nb-NO' | 'ne-NP' | 'nl-BE' | 'nl-NL' | 'pa-IN' | 'pl-PL' | 'ps-AF' | 'pt-BR' | 'pt-PT' | 'ro-RO' | 'ru-RU' | 'si-LK' | 'sk-SK' | 'sl-SI' | 'so-SO' | 'sq-AL' | 'sr-RS' | 'sv-SE' | 'sw-KE' | 'sw-TZ' | 'ta-IN' | 'te-IN' | 'th-TH' | 'tr-TR' | 'uk-UA' | 'ur-IN' | 'uz-UZ' | 'vi-VN' | 'wuu-CN' | 'yue-CN' | 'zh-CN' | 'zh-CN-shandong' | 'zh-CN-sichuan' | 'zh-HK' | 'zh-TW' | 'zu-ZA' — This is the language that will be set for the transcription. The list of languages Azure supports can be found here: https://learn.microsoft.com/en-us/azure/ai-services/speech-service/language-support?tabs=stt
          - `segmentationStrategy` 'Default' | 'Time' | 'Semantic' — Controls how phrase boundaries are detected, enabling either simple time/silence heuristics or more advanced semantic segmentation.
          - `segmentationSilenceTimeoutMs` number — Duration of detected silence after which the service finalizes a phrase. Configure to adjust sensitivity to pauses in speech.
          - `segmentationMaximumTimeMs` number — Maximum duration a segment can reach before being cut off when using time-based segmentation.
          - `fallbackPlan` FallbackTranscriberPlan
            - `transcribers` union[]
              - …
        - CustomTranscriber
          - `provider` 'custom-transcriber', required — This is the transcription provider that will be used. Use `custom-transcriber` for providers that are not natively supported.
          - `server` Server, required
            - `timeoutSeconds` number — This is the timeout in seconds for the request. Defaults to 20 seconds. @default 20
            - `credentialId` string — The credential ID for server authentication
            - `staticIpAddressesEnabled` boolean — If enabled, requests will originate from a static set of IPs owned and managed by Vapi. @default false
            - `encryptedPaths` string[] — This is the paths to encrypt in the request body if credentialId and encryptionPlan are defined.
            - `url` string — This is where the request will be sent.
            - `headers` object — These are the headers to include in the request. Each key-value pair represents a header name and its value. Note: Specifying an Authorization header here will override the authorization provided by the `credentialId` (if provided). This is an anti-pattern and should be avoided outside of edge case scenarios.
            - `backoffPlan` BackoffPlan
              - …
          - `fallbackPlan` FallbackTranscriberPlan
            - `transcribers` union[]
              - …
        - DeepgramTranscriber
          - `provider` 'deepgram', required — This is the transcription provider that will be used.
          - `model` union — This is the Deepgram model that will be used. A list of models can be found here: https://developers.deepgram.com/docs/models-languages-overview
            - 'nova-3' | 'nova-3-general' | 'nova-3-medical' | 'nova-2' | 'nova-2-general' | 'nova-2-meeting' | 'nova-2-phonecall' | 'nova-2-finance' | 'nova-2-conversationalai' | 'nova-2-voicemail' | 'nova-2-video' | 'nova-2-medical' | 'nova-2-drivethru' | 'nova-2-automotive' | 'nova' | 'nova-general' | 'nova-phonecall' | 'nova-medical' | 'enhanced' | 'enhanced-general' | 'enhanced-meeting' | 'enhanced-phonecall' | 'enhanced-finance' | 'base' | 'base-general' | 'base-meeting' | 'base-phonecall' | 'base-finance' | 'base-conversationalai' | 'base-voicemail' | 'base-video' | 'whisper' | 'flux-general-en' | 'flux-general-multi'
            - string
          - `language` 'ar' | 'az' | 'ba' | 'be' | 'bg' | 'bn' | 'br' | 'bs' | 'ca' | 'cs' | 'da' | 'da-DK' | 'de' | 'de-CH' | 'el' | 'en' | 'en-AU' | 'en-CA' | 'en-GB' | 'en-IE' | 'en-IN' | 'en-NZ' | 'en-US' | 'es' | 'es-419' | 'es-LATAM' | 'et' | 'eu' | 'fa' | 'fi' | 'fr' | 'fr-CA' | 'ha' | 'haw' | 'he' | 'hi' | 'hi-Latn' | 'hr' | 'hu' | 'id' | 'is' | 'it' | 'ja' | 'jw' | 'kn' | 'ko' | 'ko-KR' | 'ln' | 'lt' | 'lv' | 'mk' | 'mr' | 'ms' | 'multi' | 'nl' | 'nl-BE' | 'no' | 'pl' | 'pt' | 'pt-BR' | 'pt-PT' | 'ro' | 'ru' | 'sk' | 'sl' | 'sn' | 'so' | 'sr' | 'su' | 'sv' | 'sv-SE' | 'ta' | 'taq' | 'te' | 'th' | 'th-TH' | 'tl' | 'tr' | 'tt' | 'uk' | 'ur' | 'vi' | 'yo' | 'zh' | 'zh-CN' | 'zh-HK' | 'zh-Hans' | 'zh-Hant' | 'zh-TW' — This is the language that will be set for the transcription. The list of languages Deepgram supports can be found here: https://developers.deepgram.com/docs/models-languages-overview
          - `smartFormat` boolean — This will be use smart format option provided by Deepgram. It's default disabled because it can sometimes format numbers as times but it's getting better.
          - `mipOptOut` boolean — If set to true, this will add mip_opt_out=true as a query parameter of all API requests. See https://developers.deepgram.com/docs/the-deepgram-model-improvement-partnership-program#want-to-opt-out This will only be used if you are using your own Deepgram API key. @default false
          - `numerals` boolean — If set to true, this will cause deepgram to convert spoken numbers to literal numerals. For example, "my phone number is nine-seven-two..." would become "my phone number is 972..." @default false
          - `profanityFilter` boolean — If set to true, Deepgram will replace profanity in transcripts with surrounding asterisks, e.g. "f***". @default false
          - `redaction` string[] — Enables redaction of sensitive information from transcripts. Options include: - "pci": Redacts credit card numbers, expiration dates, and CVV. - "pii": Redacts personally identifiable information (names, locations, identifying numbers, etc.). - "phi": Redacts protected health information (medical conditions, drugs, injuries, etc.). - "numbers": Redacts numerical and identifying entities (dates, account numbers, SSNs, etc.). Multiple values can be provided to redact different categories simultaneously. Redacted content is replaced with entity labels like [CREDIT_CARD_1], [SSN_1], etc. See https://developers.deepgram.com/docs/redaction for details.
          - `confidenceThreshold` number — Transcripts below this confidence threshold will be discarded. @default 0.4
          - `eotThreshold` number — End-of-turn confidence required to finish a turn. Only used with Flux models. @default 0.7
          - `eotTimeoutMs` number — A turn will be finished when this much time has passed after speech, regardless of EOT confidence. Only used with Flux models. @default 5000
          - `languages` string[] — Language hints to bias Flux Multilingual (`flux-general-multi`) toward specific languages. Provide BCP-47 language codes (e.g. "en", "es", "fr"). Multiple hints can be given for multilingual or code-switching scenarios. Omit for auto-detection. Only used with `flux-general-multi`.
          - `keywords` string[] — These keywords are passed to the transcription model to help it pick up use-case specific words. Anything that may not be a common word, like your company name, should be added here.
          - `keyterm` string[] — Keyterm Prompting allows you improve Keyword Recall Rate (KRR) for important keyterms or phrases up to 90%.
          - `endpointing` number — This is the timeout after which Deepgram will send transcription on user silence. You can read in-depth documentation here: https://developers.deepgram.com/docs/endpointing. Here are the most important bits: - Defaults to 10. This is recommended for most use cases to optimize for latency. - 10 can cause some missing transcriptions since because of the shorter context. This mostly happens for one-word utterances. For those uses cases, it's recommended to try 300. It will add a bit of latency but the quality and reliability of the experience will be better. - If neither 10 nor 300 work, contact support@vapi.ai and we'll find another solution. @default 10
          - `fallbackPlan` FallbackTranscriberPlan
            - `transcribers` union[]
              - …
        - ElevenLabsTranscriber
          - `provider` '11labs', required — This is the transcription provider that will be used.
          - `model` 'scribe_v1' | 'scribe_v2' | 'scribe_v2_realtime' — This is the model that will be used for the transcription.
          - `language` 'aa' | 'ab' | 'ae' | 'af' | 'ak' | 'am' | 'an' | 'ar' | 'as' | 'av' | 'ay' | 'az' | 'ba' | 'be' | 'bg' | 'bh' | 'bi' | 'bm' | 'bn' | 'bo' | 'br' | 'bs' | 'ca' | 'ce' | 'ch' | 'co' | 'cr' | 'cs' | 'cu' | 'cv' | 'cy' | 'da' | 'de' | 'dv' | 'dz' | 'ee' | 'el' | 'en' | 'eo' | 'es' | 'et' | 'eu' | 'fa' | 'ff' | 'fi' | 'fj' | 'fo' | 'fr' | 'fy' | 'ga' | 'gd' | 'gl' | 'gn' | 'gu' | 'gv' | 'ha' | 'he' | 'hi' | 'ho' | 'hr' | 'ht' | 'hu' | 'hy' | 'hz' | 'ia' | 'id' | 'ie' | 'ig' | 'ii' | 'ik' | 'io' | 'is' | 'it' | 'iu' | 'ja' | 'jv' | 'ka' | 'kg' | 'ki' | 'kj' | 'kk' | 'kl' | 'km' | 'kn' | 'ko' | 'kr' | 'ks' | 'ku' | 'kv' | 'kw' | 'ky' | 'la' | 'lb' | 'lg' | 'li' | 'ln' | 'lo' | 'lt' | 'lu' | 'lv' | 'mg' | 'mh' | 'mi' | 'mk' | 'ml' | 'mn' | 'mr' | 'ms' | 'mt' | 'my' | 'na' | 'nb' | 'nd' | 'ne' | 'ng' | 'nl' | 'nn' | 'no' | 'nr' | 'nv' | 'ny' | 'oc' | 'oj' | 'om' | 'or' | 'os' | 'pa' | 'pi' | 'pl' | 'ps' | 'pt' | 'qu' | 'rm' | 'rn' | 'ro' | 'ru' | 'rw' | 'sa' | 'sc' | 'sd' | 'se' | 'sg' | 'si' | 'sk' | 'sl' | 'sm' | 'sn' | 'so' | 'sq' | 'sr' | 'ss' | 'st' | 'su' | 'sv' | 'sw' | 'ta' | 'te' | 'tg' | 'th' | 'ti' | 'tk' | 'tl' | 'tn' | 'to' | 'tr' | 'ts' | 'tt' | 'tw' | 'ty' | 'ug' | 'uk' | 'ur' | 'uz' | 've' | 'vi' | 'vo' | 'wa' | 'wo' | 'xh' | 'yi' | 'yue' | 'yo' | 'za' | 'zh' | 'zu' — This is the language that will be used for the transcription.
          - `silenceThresholdSeconds` number — This is the number of seconds of silence before VAD commits (0.3-3.0).
          - `confidenceThreshold` number — This is the VAD sensitivity (0.1-0.9, lower indicates more sensitive).
          - `minSpeechDurationMs` number — This is the minimum speech duration for VAD (50-2000ms).
          - `minSilenceDurationMs` number — This is the minimum silence duration for VAD (50-2000ms).
          - `fallbackPlan` FallbackTranscriberPlan
            - `transcribers` union[]
              - …
        - GladiaTranscriber
          - `provider` 'gladia', required — This is the transcription provider that will be used.
          - `model` 'fast' | 'accurate' | 'solaria-1' — This is the Gladia model that will be used. Default is 'fast'
          - `languageBehaviour` 'manual' | 'automatic single language' | 'automatic multiple languages' — Defines how the transcription model detects the audio language. Default value is 'automatic single language'.
          - `language` 'af' | 'sq' | 'am' | 'ar' | 'hy' | 'as' | 'az' | 'ba' | 'eu' | 'be' | 'bn' | 'bs' | 'br' | 'bg' | 'ca' | 'zh' | 'hr' | 'cs' | 'da' | 'nl' | 'en' | 'et' | 'fo' | 'fi' | 'fr' | 'gl' | 'ka' | 'de' | 'el' | 'gu' | 'ht' | 'ha' | 'haw' | 'he' | 'hi' | 'hu' | 'is' | 'id' | 'it' | 'ja' | 'jv' | 'kn' | 'kk' | 'km' | 'ko' | 'lo' | 'la' | 'lv' | 'ln' | 'lt' | 'lb' | 'mk' | 'mg' | 'ms' | 'ml' | 'mt' | 'mi' | 'mr' | 'mn' | 'my' | 'ne' | 'no' | 'nn' | 'oc' | 'ps' | 'fa' | 'pl' | 'pt' | 'pa' | 'ro' | 'ru' | 'sa' | 'sr' | 'sn' | 'sd' | 'si' | 'sk' | 'sl' | 'so' | 'es' | 'su' | 'sw' | 'sv' | 'tl' | 'tg' | 'ta' | 'tt' | 'te' | 'th' | 'bo' | 'tr' | 'tk' | 'uk' | 'ur' | 'uz' | 'vi' | 'cy' | 'yi' | 'yo' — Defines the language to use for the transcription. Required when languageBehaviour is 'manual'.
          - `languages` string[] — Defines the languages to use for the transcription. Required when languageBehaviour is 'manual'.
          - `transcriptionHint` string — Provides a custom vocabulary to the model to improve accuracy of transcribing context specific words, technical terms, names, etc. If empty, this argument is ignored. ⚠️ Warning ⚠️: Please be aware that the transcription_hint field has a character limit of 600. If you provide a transcription_hint longer than 600 characters, it will be automatically truncated to meet this limit.
          - `prosody` boolean — If prosody is true, you will get a transcription that can contain prosodies i.e. (laugh) (giggles) (malefic laugh) (toss) (music)… Default value is false.
          - `audioEnhancer` boolean — If true, audio will be pre-processed to improve accuracy but latency will increase. Default value is false.
          - `confidenceThreshold` number — Transcripts below this confidence threshold will be discarded. @default 0.4
          - `endpointing` number — Endpointing time in seconds - time to wait before considering speech ended
          - `speechThreshold` number — Speech threshold - sensitivity configuration for speech detection (0.0 to 1.0)
          - `customVocabularyEnabled` boolean — Enable custom vocabulary for improved accuracy
          - `customVocabularyConfig` GladiaCustomVocabularyConfigDTO
            - `vocabulary` union[], required — Array of vocabulary items (strings or objects with value, pronunciations, intensity, language)
              - …
            - `defaultIntensity` number — Default intensity for vocabulary items (0.0 to 1.0)
          - `region` 'us-west' | 'eu-west' — Region for processing audio (us-west or eu-west)
          - `receivePartialTranscripts` boolean — Enable partial transcripts for low-latency streaming transcription
          - `fallbackPlan` FallbackTranscriberPlan
            - `transcribers` union[]
              - …
        - GoogleTranscriber
          - `provider` 'google', required — This is the transcription provider that will be used.
          - `model` 'gemini-3.5-flash' | 'gemini-3.1-flash-lite' | 'gemini-3-flash-preview' | 'gemini-2.5-pro' | 'gemini-2.5-flash' | 'gemini-2.5-flash-lite' | 'gemini-2.0-flash-thinking-exp' | 'gemini-2.0-pro-exp-02-05' | 'gemini-2.0-flash' | 'gemini-2.0-flash-lite' | 'gemini-2.0-flash-exp' | 'gemini-2.0-flash-realtime-exp' | 'gemini-1.5-flash' | 'gemini-1.5-flash-002' | 'gemini-1.5-pro' | 'gemini-1.5-pro-002' | 'gemini-1.0-pro' — This is the model that will be used for the transcription.
          - `language` 'Multilingual' | 'Arabic' | 'Bengali' | 'Bulgarian' | 'Chinese' | 'Croatian' | 'Czech' | 'Danish' | 'Dutch' | 'English' | 'Estonian' | 'Finnish' | 'French' | 'German' | 'Greek' | 'Hebrew' | 'Hindi' | 'Hungarian' | 'Indonesian' | 'Italian' | 'Japanese' | 'Korean' | 'Latvian' | 'Lithuanian' | 'Norwegian' | 'Polish' | 'Portuguese' | 'Romanian' | 'Russian' | 'Serbian' | 'Slovak' | 'Slovenian' | 'Spanish' | 'Swahili' | 'Swedish' | 'Thai' | 'Turkish' | 'Ukrainian' | 'Vietnamese' — This is the language that will be set for the transcription.
          - `fallbackPlan` FallbackTranscriberPlan
            - `transcribers` union[]
              - …
        - SpeechmaticsTranscriber
          - `provider` 'speechmatics', required — This is the transcription provider that will be used.
          - `model` 'default' — This is the model that will be used for the transcription.
          - `language` 'auto' | 'ar' | 'ar_en' | 'ba' | 'eu' | 'be' | 'bn' | 'bg' | 'yue' | 'ca' | 'hr' | 'cs' | 'da' | 'nl' | 'en' | 'eo' | 'et' | 'fi' | 'fr' | 'gl' | 'de' | 'el' | 'he' | 'hi' | 'hu' | 'id' | 'ia' | 'ga' | 'it' | 'ja' | 'ko' | 'lv' | 'lt' | 'ms' | 'en_ms' | 'mt' | 'cmn' | 'cmn_en' | 'mr' | 'mn' | 'no' | 'fa' | 'pl' | 'pt' | 'ro' | 'ru' | 'sk' | 'sl' | 'es' | 'en_es' | 'sw' | 'sv' | 'tl' | 'ta' | 'en_ta' | 'th' | 'tr' | 'uk' | 'ur' | 'ug' | 'vi' | 'cy'
          - `operatingPoint` 'standard' | 'enhanced' — This is the operating point for the transcription. Choose between `standard` for faster turnaround with strong accuracy or `enhanced` for highest accuracy when precision is critical. @default 'enhanced'
          - `region` 'eu' | 'us' — This is the region for the Speechmatics API. Choose between EU (Europe) and US (United States) regions for lower latency and data sovereignty compliance. @default 'eu'
          - `enableDiarization` boolean — This enables speaker diarization, which identifies and separates speakers in the transcription. Essential for multi-speaker conversations and conference calls. @default false
          - `maxDelay` number — This sets the maximum delay in milliseconds for partial transcripts. Balances latency and accuracy. @default 3000
          - `customVocabulary` SpeechmaticsCustomVocabularyItem[], required
            - `content` string, required — The word or phrase to add to the custom vocabulary.
            - `soundsLike` string[] — Alternative phonetic representations of how the word might sound. This helps recognition when the word might be pronounced differently.
          - `numeralStyle` 'written' | 'spoken' — This controls how numbers, dates, currencies, and other entities are formatted in the transcription output. @default 'written'
          - `endOfTurnSensitivity` number — This is the sensitivity level for end-of-turn detection, which determines when a speaker has finished talking. Higher values are more sensitive. @default 0.5
          - `removeDisfluencies` boolean — This enables removal of disfluencies (um, uh) from the transcript to create cleaner, more professional output. This is only supported for the English language transcriber. @default false
          - `minimumSpeechDuration` number — This is the minimum duration in seconds for speech segments. Shorter segments will be filtered out. Helps remove noise and improve accuracy. @default 0.0
          - `fallbackPlan` FallbackTranscriberPlan
            - `transcribers` union[]
              - …
        - TalkscriberTranscriber
          - `provider` 'talkscriber', required — This is the transcription provider that will be used.
          - `model` 'whisper' — This is the model that will be used for the transcription.
          - `language` 'en' | 'zh' | 'de' | 'es' | 'ru' | 'ko' | 'fr' | 'ja' | 'pt' | 'tr' | 'pl' | 'ca' | 'nl' | 'ar' | 'sv' | 'it' | 'id' | 'hi' | 'fi' | 'vi' | 'he' | 'uk' | 'el' | 'ms' | 'cs' | 'ro' | 'da' | 'hu' | 'ta' | 'no' | 'th' | 'ur' | 'hr' | 'bg' | 'lt' | 'la' | 'mi' | 'ml' | 'cy' | 'sk' | 'te' | 'fa' | 'lv' | 'bn' | 'sr' | 'az' | 'sl' | 'kn' | 'et' | 'mk' | 'br' | 'eu' | 'is' | 'hy' | 'ne' | 'mn' | 'bs' | 'kk' | 'sq' | 'sw' | 'gl' | 'mr' | 'pa' | 'si' | 'km' | 'sn' | 'yo' | 'so' | 'af' | 'oc' | 'ka' | 'be' | 'tg' | 'sd' | 'gu' | 'am' | 'yi' | 'lo' | 'uz' | 'fo' | 'ht' | 'ps' | 'tk' | 'nn' | 'mt' | 'sa' | 'lb' | 'my' | 'bo' | 'tl' | 'mg' | 'as' | 'tt' | 'haw' | 'ln' | 'ha' | 'ba' | 'jw' | 'su' | 'yue' — This is the language that will be set for the transcription. The list of languages Whisper supports can be found here: https://github.com/openai/whisper/blob/main/whisper/tokenizer.py
          - `fallbackPlan` FallbackTranscriberPlan
            - `transcribers` union[]
              - …
        - OpenAITranscriber
          - `provider` 'openai', required — This is the transcription provider that will be used.
          - `model` 'gpt-4o-transcribe' | 'gpt-4o-mini-transcribe', required — This is the model that will be used for the transcription.
          - `language` 'af' | 'ar' | 'hy' | 'az' | 'be' | 'bs' | 'bg' | 'ca' | 'zh' | 'hr' | 'cs' | 'da' | 'nl' | 'en' | 'et' | 'fi' | 'fr' | 'gl' | 'de' | 'el' | 'he' | 'hi' | 'hu' | 'is' | 'id' | 'it' | 'ja' | 'kn' | 'kk' | 'ko' | 'lv' | 'lt' | 'mk' | 'ms' | 'mr' | 'mi' | 'ne' | 'no' | 'fa' | 'pl' | 'pt' | 'ro' | 'ru' | 'sr' | 'sk' | 'sl' | 'es' | 'sw' | 'sv' | 'tl' | 'ta' | 'th' | 'tr' | 'uk' | 'ur' | 'vi' | 'cy' — This is the language that will be set for the transcription.
          - `fallbackPlan` FallbackTranscriberPlan
            - `transcribers` union[]
              - …
        - CartesiaTranscriber
          - `provider` 'cartesia', required
          - `model` 'ink-whisper' | 'ink-2'
          - `language` 'aa' | 'ab' | 'ae' | 'af' | 'ak' | 'am' | 'an' | 'ar' | 'as' | 'av' | 'ay' | 'az' | 'ba' | 'be' | 'bg' | 'bh' | 'bi' | 'bm' | 'bn' | 'bo' | 'br' | 'bs' | 'ca' | 'ce' | 'ch' | 'co' | 'cr' | 'cs' | 'cu' | 'cv' | 'cy' | 'da' | 'de' | 'dv' | 'dz' | 'ee' | 'el' | 'en' | 'eo' | 'es' | 'et' | 'eu' | 'fa' | 'ff' | 'fi' | 'fj' | 'fo' | 'fr' | 'fy' | 'ga' | 'gd' | 'gl' | 'gn' | 'gu' | 'gv' | 'ha' | 'he' | 'hi' | 'ho' | 'hr' | 'ht' | 'hu' | 'hy' | 'hz' | 'ia' | 'id' | 'ie' | 'ig' | 'ii' | 'ik' | 'io' | 'is' | 'it' | 'iu' | 'ja' | 'jv' | 'ka' | 'kg' | 'ki' | 'kj' | 'kk' | 'kl' | 'km' | 'kn' | 'ko' | 'kr' | 'ks' | 'ku' | 'kv' | 'kw' | 'ky' | 'la' | 'lb' | 'lg' | 'li' | 'ln' | 'lo' | 'lt' | 'lu' | 'lv' | 'mg' | 'mh' | 'mi' | 'mk' | 'ml' | 'mn' | 'mr' | 'ms' | 'mt' | 'my' | 'na' | 'nb' | 'nd' | 'ne' | 'ng' | 'nl' | 'nn' | 'no' | 'nr' | 'nv' | 'ny' | 'oc' | 'oj' | 'om' | 'or' | 'os' | 'pa' | 'pi' | 'pl' | 'ps' | 'pt' | 'qu' | 'rm' | 'rn' | 'ro' | 'ru' | 'rw' | 'sa' | 'sc' | 'sd' | 'se' | 'sg' | 'si' | 'sk' | 'sl' | 'sm' | 'sn' | 'so' | 'sq' | 'sr' | 'ss' | 'st' | 'su' | 'sv' | 'sw' | 'ta' | 'te' | 'tg' | 'th' | 'ti' | 'tk' | 'tl' | 'tn' | 'to' | 'tr' | 'ts' | 'tt' | 'tw' | 'ty' | 'ug' | 'uk' | 'ur' | 'uz' | 've' | 'vi' | 'vo' | 'wa' | 'wo' | 'xh' | 'yi' | 'yue' | 'yo' | 'za' | 'zh' | 'zu'
          - `fallbackPlan` FallbackTranscriberPlan
            - `transcribers` union[]
              - …
        - SonioxTranscriber
          - `provider` 'soniox', required
          - `model` 'stt-rt-v4' | 'stt-rt-v5' — The Soniox model to use for transcription.
          - `language` 'aa' | 'ab' | 'ae' | 'af' | 'ak' | 'am' | 'an' | 'ar' | 'as' | 'av' | 'ay' | 'az' | 'ba' | 'be' | 'bg' | 'bh' | 'bi' | 'bm' | 'bn' | 'bo' | 'br' | 'bs' | 'ca' | 'ce' | 'ch' | 'co' | 'cr' | 'cs' | 'cu' | 'cv' | 'cy' | 'da' | 'de' | 'dv' | 'dz' | 'ee' | 'el' | 'en' | 'eo' | 'es' | 'et' | 'eu' | 'fa' | 'ff' | 'fi' | 'fj' | 'fo' | 'fr' | 'fy' | 'ga' | 'gd' | 'gl' | 'gn' | 'gu' | 'gv' | 'ha' | 'he' | 'hi' | 'ho' | 'hr' | 'ht' | 'hu' | 'hy' | 'hz' | 'ia' | 'id' | 'ie' | 'ig' | 'ii' | 'ik' | 'io' | 'is' | 'it' | 'iu' | 'ja' | 'jv' | 'ka' | 'kg' | 'ki' | 'kj' | 'kk' | 'kl' | 'km' | 'kn' | 'ko' | 'kr' | 'ks' | 'ku' | 'kv' | 'kw' | 'ky' | 'la' | 'lb' | 'lg' | 'li' | 'ln' | 'lo' | 'lt' | 'lu' | 'lv' | 'mg' | 'mh' | 'mi' | 'mk' | 'ml' | 'mn' | 'mr' | 'ms' | 'mt' | 'my' | 'na' | 'nb' | 'nd' | 'ne' | 'ng' | 'nl' | 'nn' | 'no' | 'nr' | 'nv' | 'ny' | 'oc' | 'oj' | 'om' | 'or' | 'os' | 'pa' | 'pi' | 'pl' | 'ps' | 'pt' | 'qu' | 'rm' | 'rn' | 'ro' | 'ru' | 'rw' | 'sa' | 'sc' | 'sd' | 'se' | 'sg' | 'si' | 'sk' | 'sl' | 'sm' | 'sn' | 'so' | 'sq' | 'sr' | 'ss' | 'st' | 'su' | 'sv' | 'sw' | 'ta' | 'te' | 'tg' | 'th' | 'ti' | 'tk' | 'tl' | 'tn' | 'to' | 'tr' | 'ts' | 'tt' | 'tw' | 'ty' | 'ug' | 'uk' | 'ur' | 'uz' | 've' | 'vi' | 'vo' | 'wa' | 'wo' | 'xh' | 'yi' | 'yue' | 'yo' | 'za' | 'zh' | 'zu' — Single language for transcription as an ISO 639-1 code (e.g., `en`, `es`). For multi-language hints or to enable Soniox auto-detect, use `languages` instead — when `languages` is set (including to an empty array), this field is ignored when building the Soniox request. Defaults to `en` if neither this nor `languages` is set.
          - `languages` string[] — Language hints sent to Soniox as `language_hints`. Provide `[lang1, lang2, ...]` (ISO 639-1 codes) to bias recognition toward specific languages, or provide an explicit empty array `[]` to enable Soniox auto-detect across all 60+ supported languages. When set (including the empty array), this field takes precedence over the singular `language` field. When omitted, falls back to the singular `language` (which defaults to `en` if also unset). Best accuracy is achieved with a single language.
          - `languageHintsStrict` boolean — When `true`, Soniox strictly restricts transcription to the languages in `languages` (or the singular `language` if `languages` is unset). When `false`, Soniox biases toward those languages but still allows transcription in other languages. Has no effect when no language hints are sent (e.g., `languages: []` for auto-detect). Defaults to `true` (strict mode).
          - `maxEndpointDelayMs` number — Maximum delay in milliseconds between when the speaker stops and when the endpoint is detected. Lower values mean faster turn-taking but more false endpoints. Range: 500-3000. Default: 500.
          - `endpointSensitivity` number — How likely Soniox is to emit an endpoint (end the caller turn). Higher values make endpoints more likely for faster turn-taking; negative values make them less likely, which helps when callers pause mid-sentence (e.g. reading numbers group by group). Range: -1.0 to 1.0. Default: 0.3 (the platform low-latency voice profile; Soniox's own default is 0.0). Supported by stt-rt-v5; omitted from the Soniox request on explicit stt-rt-v4. Soniox recommends tuning endpointLatencyAdjustmentLevel first, and advises against negative sensitivity while the level is above 0 (the settings work against each other).
          - `endpointLatencyAdjustmentLevel` number — How aggressively Soniox reduces endpoint latency. 0 is Soniox's default semantic endpointing; 3 is the most aggressive. Higher levels return endpoints sooner but may split speech into more segments and slightly reduce accuracy. Integer. Range: 0-3. Default: 2 (the platform low-latency voice profile; Soniox's own default is 0). Supported by stt-rt-v5; omitted from the Soniox request on explicit stt-rt-v4.
          - `customVocabulary` string[] — Custom vocabulary terms to boost recognition accuracy. Useful for brand names, product names, and domain-specific terminology. Maps to Soniox context.terms.
          - `contextGeneral` SonioxContextGeneralItem[] — General context key-value pairs that guide the AI model during transcription. Helps adapt vocabulary to the correct domain, improving accuracy. Recommended: 10 or fewer pairs. Maps to Soniox context.general.
            - `key` string, required — The key describing the type of context (e.g., "domain", "topic", "doctor", "organization").
            - `value` string, required — The value for the context key (e.g., "Healthcare", "Diabetes management consultation").
          - `fallbackPlan` FallbackTranscriberPlan
            - `transcribers` union[]
              - …
        - XaiTranscriber
          - `provider` 'xai', required
          - `model` 'default' — The xAI speech-to-text model to use. xAI currently exposes a single STT model — placeholder for future model selection.
          - `language` 'ar' | 'cs' | 'da' | 'nl' | 'en' | 'fil' | 'fr' | 'de' | 'hi' | 'id' | 'it' | 'ja' | 'ko' | 'mk' | 'ms' | 'fa' | 'pl' | 'pt' | 'ro' | 'ru' | 'es' | 'sv' | 'th' | 'tr' | 'vi' — Single language for transcription as an ISO 639-1 code (e.g., `en`, `es`). Defaults to `en` if not set. xAI auto-detects when omitted via the API but Vapi defaults to English for deterministic behavior.
          - `fallbackPlan` FallbackTranscriberPlan
            - `transcribers` union[]
              - …
        - VapiTranscriber
          - `provider` 'vapi', required
          - `version` 'latest' | '1' — This is the version of the Vapi transcriber. Vapi manages the underlying model and routing. When omitted, the latest version is used. Managed version params are additive-only and `'latest'` is an auto-update channel — see the param-evolution INVARIANT in `vapiManaged/types.ts`.
          - `language` 'aa' | 'ab' | 'ae' | 'af' | 'ak' | 'am' | 'an' | 'ar' | 'as' | 'av' | 'ay' | 'az' | 'ba' | 'be' | 'bg' | 'bh' | 'bi' | 'bm' | 'bn' | 'bo' | 'br' | 'bs' | 'ca' | 'ce' | 'ch' | 'co' | 'cr' | 'cs' | 'cu' | 'cv' | 'cy' | 'da' | 'de' | 'dv' | 'dz' | 'ee' | 'el' | 'en' | 'eo' | 'es' | 'et' | 'eu' | 'fa' | 'ff' | 'fi' | 'fj' | 'fo' | 'fr' | 'fy' | 'ga' | 'gd' | 'gl' | 'gn' | 'gu' | 'gv' | 'ha' | 'he' | 'hi' | 'ho' | 'hr' | 'ht' | 'hu' | 'hy' | 'hz' | 'ia' | 'id' | 'ie' | 'ig' | 'ii' | 'ik' | 'io' | 'is' | 'it' | 'iu' | 'ja' | 'jv' | 'ka' | 'kg' | 'ki' | 'kj' | 'kk' | 'kl' | 'km' | 'kn' | 'ko' | 'kr' | 'ks' | 'ku' | 'kv' | 'kw' | 'ky' | 'la' | 'lb' | 'lg' | 'li' | 'ln' | 'lo' | 'lt' | 'lu' | 'lv' | 'mg' | 'mh' | 'mi' | 'mk' | 'ml' | 'mn' | 'mr' | 'ms' | 'mt' | 'my' | 'na' | 'nb' | 'nd' | 'ne' | 'ng' | 'nl' | 'nn' | 'no' | 'nr' | 'nv' | 'ny' | 'oc' | 'oj' | 'om' | 'or' | 'os' | 'pa' | 'pi' | 'pl' | 'ps' | 'pt' | 'qu' | 'rm' | 'rn' | 'ro' | 'ru' | 'rw' | 'sa' | 'sc' | 'sd' | 'se' | 'sg' | 'si' | 'sk' | 'sl' | 'sm' | 'sn' | 'so' | 'sq' | 'sr' | 'ss' | 'st' | 'su' | 'sv' | 'sw' | 'ta' | 'te' | 'tg' | 'th' | 'ti' | 'tk' | 'tl' | 'tn' | 'to' | 'tr' | 'ts' | 'tt' | 'tw' | 'ty' | 'ug' | 'uk' | 'ur' | 'uz' | 've' | 'vi' | 'vo' | 'wa' | 'wo' | 'xh' | 'yi' | 'yue' | 'yo' | 'za' | 'zh' | 'zu' — This is the language for transcription as an ISO 639-1 code (e.g. `en`). Selecting a language locks transcription to it. For multiple languages, use `languages` instead. When neither `language` nor `languages` is set, the transcriber auto-detects the spoken language.
          - `languages` string[] — These are the languages for transcription as ISO 639-1 codes. Set one or more codes to restrict and bias recognition to those languages. An empty array `[]` (or omitting both this and `language`) enables auto-detection of the spoken language.
          - `keywords` string[] — These are custom keywords/vocabulary to boost recognition of use-case specific words (company names, product names, jargon).
          - `turnTaking` 'intelligent' | 'manual' — This is the turn-taking mode. `intelligent` uses the underlying model's native end-of-turn detection; `manual` ignores it and waits a fixed end-of-turn delay. Defaults to `intelligent`.
      - `model` union — These are the options for the assistant's LLM.
        - AnthropicModel
          - `messages` OpenAIMessage[] — This is the starting state for the conversation.
            - `content` string, nullable, required
            - `role` 'assistant' | 'function' | 'user' | 'system' | 'tool', required
          - `tools` union[] — These are the tools that the assistant can use during the call. To use existing tools, use `toolIds`. Both `tools` and `toolIds` can be used together.
            - union
              - …
          - `toolIds` string[] — These are the tools that the assistant can use during the call. To use transient tools, use `tools`. Both `tools` and `toolIds` can be used together.
          - `toolRefs` ToolRef[] — These are version-pinned references to tools. Each entry pins a specific version of a tool by `(toolId, version)`. When the same `toolId` appears in both `toolIds` and `toolRefs[]`, the `toolRefs` pin wins (the `toolIds` entry is dropped at write time).
            - `toolId` string, uuid, required — This is the unique identifier of the tool whose version is being pinned.
            - `version` string, required — Public version label of the tool, e.g. "v3"
          - `knowledgeBase` CreateCustomKnowledgeBaseDTO
            - `provider` 'custom-knowledge-base', required — This knowledge base is bring your own knowledge base implementation.
            - `server` Server, required
              - …
          - `model` 'claude-3-opus-20240229' | 'claude-3-sonnet-20240229' | 'claude-3-haiku-20240307' | 'claude-3-5-sonnet-20240620' | 'claude-3-5-sonnet-20241022' | 'claude-3-5-haiku-20241022' | 'claude-3-7-sonnet-20250219' | 'claude-opus-4-20250514' | 'claude-opus-4-5-20251101' | 'claude-opus-4-6' | 'claude-sonnet-4-20250514' | 'claude-sonnet-4-5-20250929' | 'claude-sonnet-4-6' | 'claude-sonnet-5' | 'claude-haiku-4-5-20251001', required — The specific Anthropic/Claude model that will be used.
          - `provider` 'anthropic', required — The provider identifier for Anthropic.
          - `thinking` AnthropicThinkingConfig
            - `type` 'enabled', required
            - `budgetTokens` number, required — The maximum number of tokens to allocate for thinking. Must be between 1024 and 100000 tokens.
          - `temperature` number — This is the temperature that will be used for calls. Default is 0.5.
          - `maxTokens` number — This is the max number of tokens that the assistant will be allowed to generate in each turn of the conversation. Default is 250.
          - `emotionRecognitionEnabled` boolean — This determines whether we detect user's emotion while they speak and send it as an additional info to model. Default `false` because the model is usually are good at understanding the user's emotion from text. @default false
          - `numFastTurns` number — This sets how many turns at the start of the conversation to use a smaller, faster model from the same provider before switching to the primary model. Example, gpt-3.5-turbo if provider is openai. Default is 0. @default 0
        - AnthropicBedrockModel
          - `messages` OpenAIMessage[] — This is the starting state for the conversation.
            - `content` string, nullable, required
            - `role` 'assistant' | 'function' | 'user' | 'system' | 'tool', required
          - `tools` union[] — These are the tools that the assistant can use during the call. To use existing tools, use `toolIds`. Both `tools` and `toolIds` can be used together.
            - union
              - …
          - `toolIds` string[] — These are the tools that the assistant can use during the call. To use transient tools, use `tools`. Both `tools` and `toolIds` can be used together.
          - `toolRefs` ToolRef[] — These are version-pinned references to tools. Each entry pins a specific version of a tool by `(toolId, version)`. When the same `toolId` appears in both `toolIds` and `toolRefs[]`, the `toolRefs` pin wins (the `toolIds` entry is dropped at write time).
            - `toolId` string, uuid, required — This is the unique identifier of the tool whose version is being pinned.
            - `version` string, required — Public version label of the tool, e.g. "v3"
          - `knowledgeBase` CreateCustomKnowledgeBaseDTO
            - `provider` 'custom-knowledge-base', required — This knowledge base is bring your own knowledge base implementation.
            - `server` Server, required
              - …
          - `provider` 'anthropic-bedrock', required — The provider identifier for Anthropic via AWS Bedrock.
          - `model` 'claude-3-opus-20240229' | 'claude-3-sonnet-20240229' | 'claude-3-haiku-20240307' | 'claude-3-5-sonnet-20240620' | 'claude-3-5-sonnet-20241022' | 'claude-3-5-haiku-20241022' | 'claude-3-7-sonnet-20250219' | 'claude-opus-4-20250514' | 'claude-opus-4-5-20251101' | 'claude-opus-4-6' | 'claude-sonnet-4-20250514' | 'claude-sonnet-4-5-20250929' | 'claude-sonnet-4-6' | 'claude-haiku-4-5-20251001' | 'global.anthropic.claude-haiku-4-5-20251001-v1:0', required — The specific Anthropic/Claude model that will be used via Bedrock.
          - `thinking` AnthropicThinkingConfig
            - `type` 'enabled', required
            - `budgetTokens` number, required — The maximum number of tokens to allocate for thinking. Must be between 1024 and 100000 tokens.
          - `temperature` number — This is the temperature that will be used for calls. Default is 0.5.
          - `maxTokens` number — This is the max number of tokens that the assistant will be allowed to generate in each turn of the conversation. Default is 250.
          - `emotionRecognitionEnabled` boolean — This determines whether we detect user's emotion while they speak and send it as an additional info to model. Default `false` because the model is usually are good at understanding the user's emotion from text. @default false
          - `numFastTurns` number — This sets how many turns at the start of the conversation to use a smaller, faster model from the same provider before switching to the primary model. Example, gpt-3.5-turbo if provider is openai. Default is 0. @default 0
        - AnyscaleModel
          - `messages` OpenAIMessage[] — This is the starting state for the conversation.
            - `content` string, nullable, required
            - `role` 'assistant' | 'function' | 'user' | 'system' | 'tool', required
          - `tools` union[] — These are the tools that the assistant can use during the call. To use existing tools, use `toolIds`. Both `tools` and `toolIds` can be used together.
            - union
              - …
          - `toolIds` string[] — These are the tools that the assistant can use during the call. To use transient tools, use `tools`. Both `tools` and `toolIds` can be used together.
          - `toolRefs` ToolRef[] — These are version-pinned references to tools. Each entry pins a specific version of a tool by `(toolId, version)`. When the same `toolId` appears in both `toolIds` and `toolRefs[]`, the `toolRefs` pin wins (the `toolIds` entry is dropped at write time).
            - `toolId` string, uuid, required — This is the unique identifier of the tool whose version is being pinned.
            - `version` string, required — Public version label of the tool, e.g. "v3"
          - `knowledgeBase` CreateCustomKnowledgeBaseDTO
            - `provider` 'custom-knowledge-base', required — This knowledge base is bring your own knowledge base implementation.
            - `server` Server, required
              - …
          - `provider` 'anyscale', required
          - `model` string, required — This is the name of the model. Ex. cognitivecomputations/dolphin-mixtral-8x7b
          - `temperature` number — This is the temperature that will be used for calls. Default is 0.5.
          - `maxTokens` number — This is the max number of tokens that the assistant will be allowed to generate in each turn of the conversation. Default is 250.
          - `emotionRecognitionEnabled` boolean — This determines whether we detect user's emotion while they speak and send it as an additional info to model. Default `false` because the model is usually are good at understanding the user's emotion from text. @default false
          - `numFastTurns` number — This sets how many turns at the start of the conversation to use a smaller, faster model from the same provider before switching to the primary model. Example, gpt-3.5-turbo if provider is openai. Default is 0. @default 0
        - CerebrasModel
          - `messages` OpenAIMessage[] — This is the starting state for the conversation.
            - `content` string, nullable, required
            - `role` 'assistant' | 'function' | 'user' | 'system' | 'tool', required
          - `tools` union[] — These are the tools that the assistant can use during the call. To use existing tools, use `toolIds`. Both `tools` and `toolIds` can be used together.
            - union
              - …
          - `toolIds` string[] — These are the tools that the assistant can use during the call. To use transient tools, use `tools`. Both `tools` and `toolIds` can be used together.
          - `toolRefs` ToolRef[] — These are version-pinned references to tools. Each entry pins a specific version of a tool by `(toolId, version)`. When the same `toolId` appears in both `toolIds` and `toolRefs[]`, the `toolRefs` pin wins (the `toolIds` entry is dropped at write time).
            - `toolId` string, uuid, required — This is the unique identifier of the tool whose version is being pinned.
            - `version` string, required — Public version label of the tool, e.g. "v3"
          - `knowledgeBase` CreateCustomKnowledgeBaseDTO
            - `provider` 'custom-knowledge-base', required — This knowledge base is bring your own knowledge base implementation.
            - `server` Server, required
              - …
          - `model` 'llama3.1-8b' | 'llama-3.3-70b', required — This is the name of the model. Ex. cognitivecomputations/dolphin-mixtral-8x7b
          - `provider` 'cerebras', required
          - `temperature` number — This is the temperature that will be used for calls. Default is 0.5.
          - `maxTokens` number — This is the max number of tokens that the assistant will be allowed to generate in each turn of the conversation. Default is 250.
          - `emotionRecognitionEnabled` boolean — This determines whether we detect user's emotion while they speak and send it as an additional info to model. Default `false` because the model is usually are good at understanding the user's emotion from text. @default false
          - `numFastTurns` number — This sets how many turns at the start of the conversation to use a smaller, faster model from the same provider before switching to the primary model. Example, gpt-3.5-turbo if provider is openai. Default is 0. @default 0
        - CustomLLMModel
          - `messages` OpenAIMessage[] — This is the starting state for the conversation.
            - `content` string, nullable, required
            - `role` 'assistant' | 'function' | 'user' | 'system' | 'tool', required
          - `tools` union[] — These are the tools that the assistant can use during the call. To use existing tools, use `toolIds`. Both `tools` and `toolIds` can be used together.
            - union
              - …
          - `toolIds` string[] — These are the tools that the assistant can use during the call. To use transient tools, use `tools`. Both `tools` and `toolIds` can be used together.
          - `toolRefs` ToolRef[] — These are version-pinned references to tools. Each entry pins a specific version of a tool by `(toolId, version)`. When the same `toolId` appears in both `toolIds` and `toolRefs[]`, the `toolRefs` pin wins (the `toolIds` entry is dropped at write time).
            - `toolId` string, uuid, required — This is the unique identifier of the tool whose version is being pinned.
            - `version` string, required — Public version label of the tool, e.g. "v3"
          - `knowledgeBase` CreateCustomKnowledgeBaseDTO
            - `provider` 'custom-knowledge-base', required — This knowledge base is bring your own knowledge base implementation.
            - `server` Server, required
              - …
          - `provider` 'custom-llm', required — This is the provider that will be used for the model. Any service, including your own server, that is compatible with the OpenAI API can be used.
          - `metadataSendMode` 'off' | 'variable' | 'destructured' — This determines whether metadata is sent in requests to the custom provider. - `off` will not send any metadata. payload will look like `{ messages }` - `variable` will send `assistant.metadata` as a variable on the payload. payload will look like `{ messages, metadata }` - `destructured` will send `assistant.metadata` fields directly on the payload. payload will look like `{ messages, ...metadata }` Further, `variable` and `destructured` will send `call`, `phoneNumber`, and `customer` objects in the payload. Default is `variable`.
          - `headers` object — Custom headers to send with requests. These headers can override default OpenAI headers except for Authorization (which should be specified using a custom-llm credential).
          - `url` string, required — These is the URL we'll use for the OpenAI client's `baseURL`. Ex. https://openrouter.ai/api/v1
          - `wordLevelConfidenceEnabled` boolean — This determines whether the transcriber's word level confidence is sent in requests to the custom provider. Default is false. This only works for Deepgram transcribers.
          - `timeoutSeconds` number — This sets the timeout for the connection to the custom provider without needing to stream any tokens back. Default is 20 seconds.
          - `model` string, required — This is the name of the model. Ex. cognitivecomputations/dolphin-mixtral-8x7b
          - `temperature` number — This is the temperature that will be used for calls. Default is 0.5.
          - `maxTokens` number — This is the max number of tokens that the assistant will be allowed to generate in each turn of the conversation. Default is 250.
          - `emotionRecognitionEnabled` boolean — This determines whether we detect user's emotion while they speak and send it as an additional info to model. Default `false` because the model is usually are good at understanding the user's emotion from text. @default false
          - `numFastTurns` number — This sets how many turns at the start of the conversation to use a smaller, faster model from the same provider before switching to the primary model. Example, gpt-3.5-turbo if provider is openai. Default is 0. @default 0
        - DeepInfraModel
          - `messages` OpenAIMessage[] — This is the starting state for the conversation.
            - `content` string, nullable, required
            - `role` 'assistant' | 'function' | 'user' | 'system' | 'tool', required
          - `tools` union[] — These are the tools that the assistant can use during the call. To use existing tools, use `toolIds`. Both `tools` and `toolIds` can be used together.
            - union
              - …
          - `toolIds` string[] — These are the tools that the assistant can use during the call. To use transient tools, use `tools`. Both `tools` and `toolIds` can be used together.
          - `toolRefs` ToolRef[] — These are version-pinned references to tools. Each entry pins a specific version of a tool by `(toolId, version)`. When the same `toolId` appears in both `toolIds` and `toolRefs[]`, the `toolRefs` pin wins (the `toolIds` entry is dropped at write time).
            - `toolId` string, uuid, required — This is the unique identifier of the tool whose version is being pinned.
            - `version` string, required — Public version label of the tool, e.g. "v3"
          - `knowledgeBase` CreateCustomKnowledgeBaseDTO
            - `provider` 'custom-knowledge-base', required — This knowledge base is bring your own knowledge base implementation.
            - `server` Server, required
              - …
          - `provider` 'deepinfra', required
          - `model` string, required — This is the name of the model. Ex. cognitivecomputations/dolphin-mixtral-8x7b
          - `temperature` number — This is the temperature that will be used for calls. Default is 0.5.
          - `maxTokens` number — This is the max number of tokens that the assistant will be allowed to generate in each turn of the conversation. Default is 250.
          - `emotionRecognitionEnabled` boolean — This determines whether we detect user's emotion while they speak and send it as an additional info to model. Default `false` because the model is usually are good at understanding the user's emotion from text. @default false
          - `numFastTurns` number — This sets how many turns at the start of the conversation to use a smaller, faster model from the same provider before switching to the primary model. Example, gpt-3.5-turbo if provider is openai. Default is 0. @default 0
        - DeepSeekModel
          - `messages` OpenAIMessage[] — This is the starting state for the conversation.
            - `content` string, nullable, required
            - `role` 'assistant' | 'function' | 'user' | 'system' | 'tool', required
          - `tools` union[] — These are the tools that the assistant can use during the call. To use existing tools, use `toolIds`. Both `tools` and `toolIds` can be used together.
            - union
              - …
          - `toolIds` string[] — These are the tools that the assistant can use during the call. To use transient tools, use `tools`. Both `tools` and `toolIds` can be used together.
          - `toolRefs` ToolRef[] — These are version-pinned references to tools. Each entry pins a specific version of a tool by `(toolId, version)`. When the same `toolId` appears in both `toolIds` and `toolRefs[]`, the `toolRefs` pin wins (the `toolIds` entry is dropped at write time).
            - `toolId` string, uuid, required — This is the unique identifier of the tool whose version is being pinned.
            - `version` string, required — Public version label of the tool, e.g. "v3"
          - `knowledgeBase` CreateCustomKnowledgeBaseDTO
            - `provider` 'custom-knowledge-base', required — This knowledge base is bring your own knowledge base implementation.
            - `server` Server, required
              - …
          - `model` 'deepseek-chat' | 'deepseek-reasoner', required — This is the name of the model. Ex. cognitivecomputations/dolphin-mixtral-8x7b
          - `provider` 'deep-seek', required
          - `temperature` number — This is the temperature that will be used for calls. Default is 0.5.
          - `maxTokens` number — This is the max number of tokens that the assistant will be allowed to generate in each turn of the conversation. Default is 250.
          - `emotionRecognitionEnabled` boolean — This determines whether we detect user's emotion while they speak and send it as an additional info to model. Default `false` because the model is usually are good at understanding the user's emotion from text. @default false
          - `numFastTurns` number — This sets how many turns at the start of the conversation to use a smaller, faster model from the same provider before switching to the primary model. Example, gpt-3.5-turbo if provider is openai. Default is 0. @default 0
        - GoogleModel
          - `messages` OpenAIMessage[] — This is the starting state for the conversation.
            - `content` string, nullable, required
            - `role` 'assistant' | 'function' | 'user' | 'system' | 'tool', required
          - `tools` union[] — These are the tools that the assistant can use during the call. To use existing tools, use `toolIds`. Both `tools` and `toolIds` can be used together.
            - union
              - …
          - `toolIds` string[] — These are the tools that the assistant can use during the call. To use transient tools, use `tools`. Both `tools` and `toolIds` can be used together.
          - `toolRefs` ToolRef[] — These are version-pinned references to tools. Each entry pins a specific version of a tool by `(toolId, version)`. When the same `toolId` appears in both `toolIds` and `toolRefs[]`, the `toolRefs` pin wins (the `toolIds` entry is dropped at write time).
            - `toolId` string, uuid, required — This is the unique identifier of the tool whose version is being pinned.
            - `version` string, required — Public version label of the tool, e.g. "v3"
          - `knowledgeBase` CreateCustomKnowledgeBaseDTO
            - `provider` 'custom-knowledge-base', required — This knowledge base is bring your own knowledge base implementation.
            - `server` Server, required
              - …
          - `model` 'gemini-3.5-flash' | 'gemini-3.1-flash-lite' | 'gemini-3-flash-preview' | 'gemini-2.5-pro' | 'gemini-2.5-flash' | 'gemini-2.5-flash-lite' | 'gemini-2.0-flash-thinking-exp' | 'gemini-2.0-pro-exp-02-05' | 'gemini-2.0-flash' | 'gemini-2.0-flash-lite' | 'gemini-2.0-flash-exp' | 'gemini-2.0-flash-realtime-exp' | 'gemini-1.5-flash' | 'gemini-1.5-flash-002' | 'gemini-1.5-pro' | 'gemini-1.5-pro-002' | 'gemini-1.0-pro', required — This is the Google model that will be used.
          - `provider` 'google', required
          - `realtimeConfig` GoogleRealtimeConfig
            - `topP` number — This is the nucleus sampling parameter that controls the cumulative probability of tokens considered during text generation. Only applicable with the Gemini Flash 2.0 Multimodal Live API.
            - `topK` number — This is the top-k sampling parameter that limits the number of highest probability tokens considered during text generation. Only applicable with the Gemini Flash 2.0 Multimodal Live API.
            - `presencePenalty` number — This is the presence penalty parameter that influences the model's likelihood to repeat information by penalizing tokens based on their presence in the text. Only applicable with the Gemini Flash 2.0 Multimodal Live API.
            - `frequencyPenalty` number — This is the frequency penalty parameter that influences the model's likelihood to repeat tokens by penalizing them based on their frequency in the text. Only applicable with the Gemini Flash 2.0 Multimodal Live API.
            - `speechConfig` GeminiMultimodalLiveSpeechConfig
              - …
          - `temperature` number — This is the temperature that will be used for calls. Default is 0.5.
          - `maxTokens` number — This is the max number of tokens that the assistant will be allowed to generate in each turn of the conversation. Default is 250.
          - `emotionRecognitionEnabled` boolean — This determines whether we detect user's emotion while they speak and send it as an additional info to model. Default `false` because the model is usually are good at understanding the user's emotion from text. @default false
          - `numFastTurns` number — This sets how many turns at the start of the conversation to use a smaller, faster model from the same provider before switching to the primary model. Example, gpt-3.5-turbo if provider is openai. Default is 0. @default 0
        - GroqModel
          - `messages` OpenAIMessage[] — This is the starting state for the conversation.
            - `content` string, nullable, required
            - `role` 'assistant' | 'function' | 'user' | 'system' | 'tool', required
          - `tools` union[] — These are the tools that the assistant can use during the call. To use existing tools, use `toolIds`. Both `tools` and `toolIds` can be used together.
            - union
              - …
          - `toolIds` string[] — These are the tools that the assistant can use during the call. To use transient tools, use `tools`. Both `tools` and `toolIds` can be used together.
          - `toolRefs` ToolRef[] — These are version-pinned references to tools. Each entry pins a specific version of a tool by `(toolId, version)`. When the same `toolId` appears in both `toolIds` and `toolRefs[]`, the `toolRefs` pin wins (the `toolIds` entry is dropped at write time).
            - `toolId` string, uuid, required — This is the unique identifier of the tool whose version is being pinned.
            - `version` string, required — Public version label of the tool, e.g. "v3"
          - `knowledgeBase` CreateCustomKnowledgeBaseDTO
            - `provider` 'custom-knowledge-base', required — This knowledge base is bring your own knowledge base implementation.
            - `server` Server, required
              - …
          - `model` 'openai/gpt-oss-20b' | 'openai/gpt-oss-120b' | 'deepseek-r1-distill-llama-70b' | 'llama-3.3-70b-versatile' | 'llama-3.1-405b-reasoning' | 'llama-3.1-8b-instant' | 'llama3-8b-8192' | 'llama3-70b-8192' | 'gemma2-9b-it' | 'moonshotai/kimi-k2-instruct-0905' | 'meta-llama/llama-4-scout-17b-16e-instruct' | 'mistral-saba-24b' | 'compound-beta' | 'compound-beta-mini', required — This is the name of the model. Ex. cognitivecomputations/dolphin-mixtral-8x7b
          - `provider` 'groq', required
          - `temperature` number — This is the temperature that will be used for calls. Default is 0.5.
          - `maxTokens` number — This is the max number of tokens that the assistant will be allowed to generate in each turn of the conversation. Default is 250.
          - `emotionRecognitionEnabled` boolean — This determines whether we detect user's emotion while they speak and send it as an additional info to model. Default `false` because the model is usually are good at understanding the user's emotion from text. @default false
          - `numFastTurns` number — This sets how many turns at the start of the conversation to use a smaller, faster model from the same provider before switching to the primary model. Example, gpt-3.5-turbo if provider is openai. Default is 0. @default 0
        - InflectionAIModel
          - `messages` OpenAIMessage[] — This is the starting state for the conversation.
            - `content` string, nullable, required
            - `role` 'assistant' | 'function' | 'user' | 'system' | 'tool', required
          - `tools` union[] — These are the tools that the assistant can use during the call. To use existing tools, use `toolIds`. Both `tools` and `toolIds` can be used together.
            - union
              - …
          - `toolIds` string[] — These are the tools that the assistant can use during the call. To use transient tools, use `tools`. Both `tools` and `toolIds` can be used together.
          - `toolRefs` ToolRef[] — These are version-pinned references to tools. Each entry pins a specific version of a tool by `(toolId, version)`. When the same `toolId` appears in both `toolIds` and `toolRefs[]`, the `toolRefs` pin wins (the `toolIds` entry is dropped at write time).
            - `toolId` string, uuid, required — This is the unique identifier of the tool whose version is being pinned.
            - `version` string, required — Public version label of the tool, e.g. "v3"
          - `knowledgeBase` CreateCustomKnowledgeBaseDTO
            - `provider` 'custom-knowledge-base', required — This knowledge base is bring your own knowledge base implementation.
            - `server` Server, required
              - …
          - `model` 'inflection_3_pi', required — This is the name of the model. Ex. cognitivecomputations/dolphin-mixtral-8x7b
          - `provider` 'inflection-ai', required
          - `temperature` number — This is the temperature that will be used for calls. Default is 0.5.
          - `maxTokens` number — This is the max number of tokens that the assistant will be allowed to generate in each turn of the conversation. Default is 250.
          - `emotionRecognitionEnabled` boolean — This determines whether we detect user's emotion while they speak and send it as an additional info to model. Default `false` because the model is usually are good at understanding the user's emotion from text. @default false
          - `numFastTurns` number — This sets how many turns at the start of the conversation to use a smaller, faster model from the same provider before switching to the primary model. Example, gpt-3.5-turbo if provider is openai. Default is 0. @default 0
        - MinimaxLLMModel
          - `messages` OpenAIMessage[] — This is the starting state for the conversation.
            - `content` string, nullable, required
            - `role` 'assistant' | 'function' | 'user' | 'system' | 'tool', required
          - `tools` union[] — These are the tools that the assistant can use during the call. To use existing tools, use `toolIds`. Both `tools` and `toolIds` can be used together.
            - union
              - …
          - `toolIds` string[] — These are the tools that the assistant can use during the call. To use transient tools, use `tools`. Both `tools` and `toolIds` can be used together.
          - `toolRefs` ToolRef[] — These are version-pinned references to tools. Each entry pins a specific version of a tool by `(toolId, version)`. When the same `toolId` appears in both `toolIds` and `toolRefs[]`, the `toolRefs` pin wins (the `toolIds` entry is dropped at write time).
            - `toolId` string, uuid, required — This is the unique identifier of the tool whose version is being pinned.
            - `version` string, required — Public version label of the tool, e.g. "v3"
          - `knowledgeBase` CreateCustomKnowledgeBaseDTO
            - `provider` 'custom-knowledge-base', required — This knowledge base is bring your own knowledge base implementation.
            - `server` Server, required
              - …
          - `provider` 'minimax', required
          - `model` 'MiniMax-M2.7', required — This is the name of the model. Ex. cognitivecomputations/dolphin-mixtral-8x7b
          - `temperature` number — This is the temperature that will be used for calls. Default is 0.5.
          - `maxTokens` number — This is the max number of tokens that the assistant will be allowed to generate in each turn of the conversation. Default is 250.
          - `emotionRecognitionEnabled` boolean — This determines whether we detect user's emotion while they speak and send it as an additional info to model. Default `false` because the model is usually are good at understanding the user's emotion from text. @default false
          - `numFastTurns` number — This sets how many turns at the start of the conversation to use a smaller, faster model from the same provider before switching to the primary model. Example, gpt-3.5-turbo if provider is openai. Default is 0. @default 0
        - OpenAIModel
          - `messages` OpenAIMessage[] — This is the starting state for the conversation.
            - `content` string, nullable, required
            - `role` 'assistant' | 'function' | 'user' | 'system' | 'tool', required
          - `tools` union[] — These are the tools that the assistant can use during the call. To use existing tools, use `toolIds`. Both `tools` and `toolIds` can be used together.
            - union
              - …
          - `toolIds` string[] — These are the tools that the assistant can use during the call. To use transient tools, use `tools`. Both `tools` and `toolIds` can be used together.
          - `toolRefs` ToolRef[] — These are version-pinned references to tools. Each entry pins a specific version of a tool by `(toolId, version)`. When the same `toolId` appears in both `toolIds` and `toolRefs[]`, the `toolRefs` pin wins (the `toolIds` entry is dropped at write time).
            - `toolId` string, uuid, required — This is the unique identifier of the tool whose version is being pinned.
            - `version` string, required — Public version label of the tool, e.g. "v3"
          - `knowledgeBase` CreateCustomKnowledgeBaseDTO
            - `provider` 'custom-knowledge-base', required — This knowledge base is bring your own knowledge base implementation.
            - `server` Server, required
              - …
          - `provider` 'openai', required — This is the provider that will be used for the model.
          - `model` 'gpt-5.6-sol' | 'gpt-5.6-terra' | 'gpt-5.6-luna' | 'gpt-5.5' | 'chat-latest' | 'gpt-5.4' | 'gpt-5.4-mini' | 'gpt-5.4-nano' | 'gpt-5.2' | 'gpt-5.2-chat-latest' | 'gpt-5.1' | 'gpt-5.1-chat-latest' | 'gpt-5' | 'gpt-5-chat-latest' | 'gpt-5-mini' | 'gpt-5-nano' | 'gpt-4.1-2025-04-14' | 'gpt-4.1-mini-2025-04-14' | 'gpt-4.1-nano-2025-04-14' | 'gpt-4.1' | 'gpt-4.1-mini' | 'gpt-4.1-nano' | 'chatgpt-4o-latest' | 'o3' | 'o3-mini' | 'o4-mini' | 'o1-mini' | 'o1-mini-2024-09-12' | 'gpt-4o-realtime-preview-2024-10-01' | 'gpt-4o-realtime-preview-2024-12-17' | 'gpt-4o-mini-realtime-preview-2024-12-17' | 'gpt-realtime-2025-08-28' | 'gpt-realtime-mini-2025-12-15' | 'gpt-realtime-2' | 'gpt-4o-mini-2024-07-18' | 'gpt-4o-mini' | 'gpt-4o' | 'gpt-4o-2024-05-13' | 'gpt-4o-2024-08-06' | 'gpt-4o-2024-11-20' | 'gpt-4-turbo' | 'gpt-4-turbo-2024-04-09' | 'gpt-4-turbo-preview' | 'gpt-4-0125-preview' | 'gpt-4-1106-preview' | 'gpt-4' | 'gpt-4-0613' | 'gpt-3.5-turbo' | 'gpt-3.5-turbo-0125' | 'gpt-3.5-turbo-1106' | 'gpt-3.5-turbo-16k' | 'gpt-3.5-turbo-0613' | 'gpt-5.6-luna:westus3' | 'gpt-5.6-terra:westus3' | 'gpt-5.6-sol:westus3' | 'gpt-5.4:eastus2' | 'gpt-5.4:swedencentral' | 'gpt-5.4-mini:eastus2' | 'gpt-5.4-mini:swedencentral' | 'gpt-5.4-nano:eastus2' | 'gpt-5.4-nano:swedencentral' | 'gpt-5.2:eastus2' | 'gpt-5.2:swedencentral' | 'gpt-5.1:eastus2' | 'gpt-5.1:swedencentral' | 'gpt-5:eastus2' | 'gpt-5:swedencentral' | 'gpt-5:canadaeast' | 'gpt-5:eastus' | 'gpt-5:westeurope' | 'gpt-5:germanywestcentral' | 'gpt-5:polandcentral' | 'gpt-5:spaincentral' | 'gpt-5-mini:eastus2' | 'gpt-5-mini:swedencentral' | 'gpt-5-mini:westeurope' | 'gpt-5-mini:germanywestcentral' | 'gpt-5-mini:polandcentral' | 'gpt-5-mini:spaincentral' | 'gpt-5-nano:eastus2' | 'gpt-5-nano:swedencentral' | 'gpt-4.1-2025-04-14:westus' | 'gpt-4.1-2025-04-14:eastus2' | 'gpt-4.1-2025-04-14:eastus' | 'gpt-4.1-2025-04-14:westus3' | 'gpt-4.1-2025-04-14:northcentralus' | 'gpt-4.1-2025-04-14:southcentralus' | 'gpt-4.1-2025-04-14:westeurope' | 'gpt-4.1-2025-04-14:germanywestcentral' | 'gpt-4.1-2025-04-14:polandcentral' | 'gpt-4.1-2025-04-14:spaincentral' | 'gpt-4.1-mini-2025-04-14:westus' | 'gpt-4.1-mini-2025-04-14:eastus2' | 'gpt-4.1-mini-2025-04-14:eastus' | 'gpt-4.1-mini-2025-04-14:westus3' | 'gpt-4.1-mini-2025-04-14:northcentralus' | 'gpt-4.1-mini-2025-04-14:southcentralus' | 'gpt-4.1-mini-2025-04-14:westeurope' | 'gpt-4.1-mini-2025-04-14:germanywestcentral' | 'gpt-4.1-mini-2025-04-14:polandcentral' | 'gpt-4.1-mini-2025-04-14:spaincentral' | 'gpt-4.1-nano-2025-04-14:westus' | 'gpt-4.1-nano-2025-04-14:eastus2' | 'gpt-4.1-nano-2025-04-14:westus3' | 'gpt-4.1-nano-2025-04-14:northcentralus' | 'gpt-4.1-nano-2025-04-14:southcentralus' | 'gpt-4o-2024-11-20:swedencentral' | 'gpt-4o-2024-11-20:westus' | 'gpt-4o-2024-11-20:eastus2' | 'gpt-4o-2024-11-20:eastus' | 'gpt-4o-2024-11-20:westus3' | 'gpt-4o-2024-11-20:southcentralus' | 'gpt-4o-2024-11-20:westeurope' | 'gpt-4o-2024-11-20:germanywestcentral' | 'gpt-4o-2024-11-20:polandcentral' | 'gpt-4o-2024-11-20:spaincentral' | 'gpt-4o-2024-08-06:westus' | 'gpt-4o-2024-08-06:westus3' | 'gpt-4o-2024-08-06:eastus' | 'gpt-4o-2024-08-06:eastus2' | 'gpt-4o-2024-08-06:northcentralus' | 'gpt-4o-2024-08-06:southcentralus' | 'gpt-4o-mini-2024-07-18:westus' | 'gpt-4o-mini-2024-07-18:westus3' | 'gpt-4o-mini-2024-07-18:eastus' | 'gpt-4o-mini-2024-07-18:eastus2' | 'gpt-4o-mini-2024-07-18:northcentralus' | 'gpt-4o-mini-2024-07-18:southcentralus' | 'gpt-4o-2024-05-13:eastus2' | 'gpt-4o-2024-05-13:eastus' | 'gpt-4o-2024-05-13:northcentralus' | 'gpt-4o-2024-05-13:southcentralus' | 'gpt-4o-2024-05-13:westus3' | 'gpt-4o-2024-05-13:westus' | 'gpt-4-turbo-2024-04-09:eastus2' | 'gpt-4-0125-preview:eastus' | 'gpt-4-0125-preview:northcentralus' | 'gpt-4-0125-preview:southcentralus' | 'gpt-4-1106-preview:australiaeast' | 'gpt-4-1106-preview:canadaeast' | 'gpt-4-1106-preview:france' | 'gpt-4-1106-preview:india' | 'gpt-4-1106-preview:norway' | 'gpt-4-1106-preview:swedencentral' | 'gpt-4-1106-preview:uk' | 'gpt-4-1106-preview:westus' | 'gpt-4-1106-preview:westus3' | 'gpt-4-0613:canadaeast' | 'gpt-3.5-turbo-0125:canadaeast' | 'gpt-3.5-turbo-0125:northcentralus' | 'gpt-3.5-turbo-0125:southcentralus' | 'gpt-3.5-turbo-1106:canadaeast' | 'gpt-3.5-turbo-1106:westus' | 'gpt-4.1:australiaeast' | 'gpt-4o:australiaeast' | 'gpt-5.4-mini:australiaeast', required — This is the OpenAI model that will be used. When using Vapi OpenAI or your own Azure Credentials, you have the option to specify the region for the selected model. This shouldn't be specified unless you have a specific reason to do so. Vapi will automatically find the fastest region that make sense. This is helpful when you are required to comply with Data Residency rules. Learn more about Azure regions here https://azure.microsoft.com/en-us/explore/global-infrastructure/data-residency/. @default undefined
          - `fallbackModels` string[] — These are the fallback models that will be used if the primary model fails. This shouldn't be specified unless you have a specific reason to do so. Vapi will automatically find the fastest fallbacks that make sense.
          - `toolStrictCompatibilityMode` 'strip-parameters-with-unsupported-validation' | 'strip-unsupported-validation' — Azure OpenAI doesn't support `maxLength` right now https://learn.microsoft.com/en-us/azure/ai-services/openai/how-to/structured-outputs?tabs=python-secure%2Cdotnet-entra-id&pivots=programming-language-csharp#unsupported-type-specific-keywords. Need to strip. - `strip-parameters-with-unsupported-validation` will strip parameters with unsupported validation. - `strip-unsupported-validation` will keep the parameters but strip unsupported validation. @default `strip-unsupported-validation`
          - `promptCacheRetention` 'in_memory' | '24h' — This controls the prompt cache retention policy for models that support extended caching (GPT-4.1, GPT-5 series). - `in_memory`: Default behavior, cache retained in GPU memory only - `24h`: Extended caching, keeps cached prefixes active for up to 24 hours by offloading to GPU-local storage Only applies to models: gpt-5.6-sol, gpt-5.6-terra, gpt-5.6-luna, gpt-5.5, chat-latest, gpt-5.4, gpt-5.4-mini, gpt-5.4-nano, gpt-5.2, gpt-5.1, gpt-5.1-codex, gpt-5.1-codex-mini, gpt-5.1-chat-latest, gpt-5, gpt-5-codex, gpt-4.1 @default undefined (uses API default which is 'in_memory')
          - `promptCacheKey` string — This is the prompt cache key for models that support extended caching (GPT-4.1, GPT-5 series). Providing a cache key allows you to share cached prefixes across requests. @default undefined
          - `reasoningEffort` 'minimal' | 'none' | 'low' | 'medium' | 'high' | 'xhigh' — Reasoning effort for reasoning-capable OpenAI models. For `gpt-realtime-2`: forwarded to V2 stream's session.update as `reasoning.effort`. For non-realtime OpenAI models, model-aware validation limits newly public values while preserving the existing four-value storage contract.
          - `temperature` number — This is the temperature that will be used for calls. Default is 0.5.
          - `maxTokens` number — This is the max number of tokens that the assistant will be allowed to generate in each turn of the conversation. Default is 250.
          - `emotionRecognitionEnabled` boolean — This determines whether we detect user's emotion while they speak and send it as an additional info to model. Default `false` because the model is usually are good at understanding the user's emotion from text. @default false
          - `numFastTurns` number — This sets how many turns at the start of the conversation to use a smaller, faster model from the same provider before switching to the primary model. Example, gpt-3.5-turbo if provider is openai. Default is 0. @default 0
        - OpenRouterModel
          - `messages` OpenAIMessage[] — This is the starting state for the conversation.
            - `content` string, nullable, required
            - `role` 'assistant' | 'function' | 'user' | 'system' | 'tool', required
          - `tools` union[] — These are the tools that the assistant can use during the call. To use existing tools, use `toolIds`. Both `tools` and `toolIds` can be used together.
            - union
              - …
          - `toolIds` string[] — These are the tools that the assistant can use during the call. To use transient tools, use `tools`. Both `tools` and `toolIds` can be used together.
          - `toolRefs` ToolRef[] — These are version-pinned references to tools. Each entry pins a specific version of a tool by `(toolId, version)`. When the same `toolId` appears in both `toolIds` and `toolRefs[]`, the `toolRefs` pin wins (the `toolIds` entry is dropped at write time).
            - `toolId` string, uuid, required — This is the unique identifier of the tool whose version is being pinned.
            - `version` string, required — Public version label of the tool, e.g. "v3"
          - `knowledgeBase` CreateCustomKnowledgeBaseDTO
            - `provider` 'custom-knowledge-base', required — This knowledge base is bring your own knowledge base implementation.
            - `server` Server, required
              - …
          - `provider` 'openrouter', required
          - `model` string, required — This is the name of the model. Ex. cognitivecomputations/dolphin-mixtral-8x7b
          - `temperature` number — This is the temperature that will be used for calls. Default is 0.5.
          - `maxTokens` number — This is the max number of tokens that the assistant will be allowed to generate in each turn of the conversation. Default is 250.
          - `emotionRecognitionEnabled` boolean — This determines whether we detect user's emotion while they speak and send it as an additional info to model. Default `false` because the model is usually are good at understanding the user's emotion from text. @default false
          - `numFastTurns` number — This sets how many turns at the start of the conversation to use a smaller, faster model from the same provider before switching to the primary model. Example, gpt-3.5-turbo if provider is openai. Default is 0. @default 0
        - PerplexityAIModel
          - `messages` OpenAIMessage[] — This is the starting state for the conversation.
            - `content` string, nullable, required
            - `role` 'assistant' | 'function' | 'user' | 'system' | 'tool', required
          - `tools` union[] — These are the tools that the assistant can use during the call. To use existing tools, use `toolIds`. Both `tools` and `toolIds` can be used together.
            - union
              - …
          - `toolIds` string[] — These are the tools that the assistant can use during the call. To use transient tools, use `tools`. Both `tools` and `toolIds` can be used together.
          - `toolRefs` ToolRef[] — These are version-pinned references to tools. Each entry pins a specific version of a tool by `(toolId, version)`. When the same `toolId` appears in both `toolIds` and `toolRefs[]`, the `toolRefs` pin wins (the `toolIds` entry is dropped at write time).
            - `toolId` string, uuid, required — This is the unique identifier of the tool whose version is being pinned.
            - `version` string, required — Public version label of the tool, e.g. "v3"
          - `knowledgeBase` CreateCustomKnowledgeBaseDTO
            - `provider` 'custom-knowledge-base', required — This knowledge base is bring your own knowledge base implementation.
            - `server` Server, required
              - …
          - `provider` 'perplexity-ai', required
          - `model` string, required — This is the name of the model. Ex. cognitivecomputations/dolphin-mixtral-8x7b
          - `temperature` number — This is the temperature that will be used for calls. Default is 0.5.
          - `maxTokens` number — This is the max number of tokens that the assistant will be allowed to generate in each turn of the conversation. Default is 250.
          - `emotionRecognitionEnabled` boolean — This determines whether we detect user's emotion while they speak and send it as an additional info to model. Default `false` because the model is usually are good at understanding the user's emotion from text. @default false
          - `numFastTurns` number — This sets how many turns at the start of the conversation to use a smaller, faster model from the same provider before switching to the primary model. Example, gpt-3.5-turbo if provider is openai. Default is 0. @default 0
        - TogetherAIModel
          - `messages` OpenAIMessage[] — This is the starting state for the conversation.
            - `content` string, nullable, required
            - `role` 'assistant' | 'function' | 'user' | 'system' | 'tool', required
          - `tools` union[] — These are the tools that the assistant can use during the call. To use existing tools, use `toolIds`. Both `tools` and `toolIds` can be used together.
            - union
              - …
          - `toolIds` string[] — These are the tools that the assistant can use during the call. To use transient tools, use `tools`. Both `tools` and `toolIds` can be used together.
          - `toolRefs` ToolRef[] — These are version-pinned references to tools. Each entry pins a specific version of a tool by `(toolId, version)`. When the same `toolId` appears in both `toolIds` and `toolRefs[]`, the `toolRefs` pin wins (the `toolIds` entry is dropped at write time).
            - `toolId` string, uuid, required — This is the unique identifier of the tool whose version is being pinned.
            - `version` string, required — Public version label of the tool, e.g. "v3"
          - `knowledgeBase` CreateCustomKnowledgeBaseDTO
            - `provider` 'custom-knowledge-base', required — This knowledge base is bring your own knowledge base implementation.
            - `server` Server, required
              - …
          - `provider` 'together-ai', required
          - `model` string, required — This is the name of the model. Ex. cognitivecomputations/dolphin-mixtral-8x7b
          - `temperature` number — This is the temperature that will be used for calls. Default is 0.5.
          - `maxTokens` number — This is the max number of tokens that the assistant will be allowed to generate in each turn of the conversation. Default is 250.
          - `emotionRecognitionEnabled` boolean — This determines whether we detect user's emotion while they speak and send it as an additional info to model. Default `false` because the model is usually are good at understanding the user's emotion from text. @default false
          - `numFastTurns` number — This sets how many turns at the start of the conversation to use a smaller, faster model from the same provider before switching to the primary model. Example, gpt-3.5-turbo if provider is openai. Default is 0. @default 0
        - XaiModel
          - `messages` OpenAIMessage[] — This is the starting state for the conversation.
            - `content` string, nullable, required
            - `role` 'assistant' | 'function' | 'user' | 'system' | 'tool', required
          - `tools` union[] — These are the tools that the assistant can use during the call. To use existing tools, use `toolIds`. Both `tools` and `toolIds` can be used together.
            - union
              - …
          - `toolIds` string[] — These are the tools that the assistant can use during the call. To use transient tools, use `tools`. Both `tools` and `toolIds` can be used together.
          - `toolRefs` ToolRef[] — These are version-pinned references to tools. Each entry pins a specific version of a tool by `(toolId, version)`. When the same `toolId` appears in both `toolIds` and `toolRefs[]`, the `toolRefs` pin wins (the `toolIds` entry is dropped at write time).
            - `toolId` string, uuid, required — This is the unique identifier of the tool whose version is being pinned.
            - `version` string, required — Public version label of the tool, e.g. "v3"
          - `knowledgeBase` CreateCustomKnowledgeBaseDTO
            - `provider` 'custom-knowledge-base', required — This knowledge base is bring your own knowledge base implementation.
            - `server` Server, required
              - …
          - `model` 'grok-beta' | 'grok-2' | 'grok-3' | 'grok-4-fast-reasoning' | 'grok-4-fast-non-reasoning' | 'grok-4.20-0309-reasoning' | 'grok-4.20-0309-non-reasoning' | 'grok-4.3', required — This is the name of the model. Ex. cognitivecomputations/dolphin-mixtral-8x7b
          - `provider` 'xai', required
          - `temperature` number — This is the temperature that will be used for calls. Default is 0.5.
          - `maxTokens` number — This is the max number of tokens that the assistant will be allowed to generate in each turn of the conversation. Default is 250.
          - `emotionRecognitionEnabled` boolean — This determines whether we detect user's emotion while they speak and send it as an additional info to model. Default `false` because the model is usually are good at understanding the user's emotion from text. @default false
          - `numFastTurns` number — This sets how many turns at the start of the conversation to use a smaller, faster model from the same provider before switching to the primary model. Example, gpt-3.5-turbo if provider is openai. Default is 0. @default 0
        - VapiModel
          - `messages` OpenAIMessage[] — This is the starting state for the conversation.
            - `content` string, nullable, required
            - `role` 'assistant' | 'function' | 'user' | 'system' | 'tool', required
          - `tools` union[] — These are the tools that the assistant can use during the call. To use existing tools, use `toolIds`. Both `tools` and `toolIds` can be used together.
            - union
              - …
          - `toolIds` string[] — These are the tools that the assistant can use during the call. To use transient tools, use `tools`. Both `tools` and `toolIds` can be used together.
          - `toolRefs` ToolRef[] — These are version-pinned references to tools. Each entry pins a specific version of a tool by `(toolId, version)`. When the same `toolId` appears in both `toolIds` and `toolRefs[]`, the `toolRefs` pin wins (the `toolIds` entry is dropped at write time).
            - `toolId` string, uuid, required — This is the unique identifier of the tool whose version is being pinned.
            - `version` string, required — Public version label of the tool, e.g. "v3"
          - `knowledgeBase` CreateCustomKnowledgeBaseDTO
            - `provider` 'custom-knowledge-base', required — This knowledge base is bring your own knowledge base implementation.
            - `server` Server, required
              - …
          - `model` string — White-label Vapi models are selected by `version`, not a model name, so `model` is optional here (the runtime already accepts a version-only Vapi payload). Overriding the required `ModelBase.model`: the declared type stays `string` to match the base (avoids TS2416) and the `= undefined!` initializer satisfies TS2612 for the field override, while `@IsOptional` + `@ApiPropertyOptional` make validation and the generated OpenAPI schema treat it as optional (so `VapiModel.required` is `['provider']`).
          - `version` 'latest' | '1' — Vapi-managed model version (update channel). When set, this is a Vapi-managed LLM routed by the registry; when absent, this is the legacy workflow form below (`steps` / `workflow`).
          - `provider` 'vapi', required
          - `workflowId` string — This is the workflow that will be used for the call. To use a transient workflow, use `workflow` instead.
          - `workflow` WorkflowUserEditable
            - `nodes` union[], required
              - …
            - `model` union — This is the model for the workflow. This can be overridden at node level using `nodes[n].model`.
              - …
            - `transcriber` union — This is the transcriber for the workflow. This can be overridden at node level using `nodes[n].transcriber`.
              - …
            - `voice` union — This is the voice for the workflow. This can be overridden at node level using `nodes[n].voice`.
              - …
            - `observabilityPlan` LangfuseObservabilityPlan
              - …
            - `backgroundSound` union — This is the background sound in the call. Default for phone calls is 'office' and default for web calls is 'off'. You can also provide a custom sound by providing a URL to an audio file.
              - …
            - `hooks` union[] — This is a set of actions that will be performed on certain events.
              - …
            - `credentials` union[] — These are dynamic credentials that will be used for the workflow calls. By default, all the credentials are available for use in the call but you can supplement an additional credentials using this. Dynamic credentials override existing credentials.
              - …
            - `voicemailDetection` union — This is the voicemail detection plan for the workflow.
              - …
            - `maxDurationSeconds` number — This is the maximum duration of the call in seconds. After this duration, the call will automatically end. Default is 1800 (30 minutes), max is 43200 (12 hours), and min is 10 seconds.
            - `name` string, required
            - `edges` Edge[], required
              - …
            - `globalPrompt` string
            - `server` Server
              - …
            - `compliancePlan` CompliancePlan
              - …
            - `analysisPlan` AnalysisPlan
              - …
            - `artifactPlan` ArtifactPlan
              - …
            - `startSpeakingPlan` StartSpeakingPlan
              - …
            - `stopSpeakingPlan` StopSpeakingPlan
              - …
            - `monitorPlan` MonitorPlan
              - …
            - `backgroundSpeechDenoisingPlan` BackgroundSpeechDenoisingPlan
              - …
            - `credentialIds` string[] — These are the credentials that will be used for the workflow calls. By default, all the credentials are available for use in the call but you can provide a subset using this.
            - `keypadInputPlan` KeypadInputPlan
              - …
            - `voicemailMessage` string — This is the message that the assistant will say if the call is forwarded to voicemail. If unspecified, it will hang up.
          - `temperature` number — This is the temperature that will be used for calls. Default is 0.5.
          - `emotionRecognitionEnabled` boolean — This determines whether we detect user's emotion while they speak and send it as an additional info to model. Default `false` because the model is usually are good at understanding the user's emotion from text. @default false
          - `numFastTurns` number — This sets how many turns at the start of the conversation to use a smaller, faster model from the same provider before switching to the primary model. Example, gpt-3.5-turbo if provider is openai. Default is 0. @default 0
      - `voice` union — These are the options for the assistant's voice.
        - AzureVoice
          - `cachingEnabled` boolean — This is the flag to toggle voice caching for the assistant.
          - `provider` 'azure', required — This is the voice provider that will be used.
          - `voiceId` union, required — This is the provider-specific ID that will be used.
            - 'andrew' | 'brian' | 'emma'
            - string
          - `chunkPlan` ChunkPlan
            - `enabled` boolean — This determines whether the model output is chunked before being sent to the voice provider. Default `true`. Usage: - To rely on the voice provider's audio generation logic, set this to `false`. - If seeing issues with quality, set this to `true`. If disabled, Vapi-provided audio control tokens like <flush /> will not work. @default true
            - `minCharacters` number — This is the minimum number of characters in a chunk. Usage: - To increase quality, set this to a higher value. - To decrease latency, set this to a lower value. @default 30
            - `punctuationBoundaries` string[] — These are the punctuations that are considered valid boundaries for a chunk to be created. Usage: - To increase quality, constrain to fewer boundaries. - To decrease latency, enable all. Default is automatically set to balance the trade-off between quality and latency based on the provider.
            - `formatPlan` FormatPlan
              - …
          - `speed` number — This is the speed multiplier that will be used.
          - `fallbackPlan` FallbackPlan
            - `voices` union[], required — This is the list of voices to fallback to in the event that the primary voice provider fails.
              - …
        - CartesiaVoice
          - `cachingEnabled` boolean — This is the flag to toggle voice caching for the assistant.
          - `provider` 'cartesia', required — This is the voice provider that will be used.
          - `voiceId` string, required — The ID of the particular voice you want to use.
          - `model` 'sonic-3.5' | 'sonic-3.5-2026-05-04' | 'sonic-3' | 'sonic-3-2026-01-12' | 'sonic-3-2025-10-27' | 'sonic-2' | 'sonic-2-2025-06-11' | 'sonic-english' | 'sonic-multilingual' | 'sonic-preview' | 'sonic' — This is the model that will be used. This is optional and will default to the correct model for the voiceId.
          - `language` 'ar' | 'bg' | 'bn' | 'cs' | 'da' | 'de' | 'el' | 'en' | 'es' | 'fi' | 'fr' | 'gu' | 'he' | 'hi' | 'hr' | 'hu' | 'id' | 'it' | 'ja' | 'ka' | 'kn' | 'ko' | 'ml' | 'mr' | 'ms' | 'nl' | 'no' | 'pa' | 'pl' | 'pt' | 'ro' | 'ru' | 'sk' | 'sv' | 'ta' | 'te' | 'th' | 'tl' | 'tr' | 'uk' | 'vi' | 'zh' — This is the language that will be used. This is optional and will default to the correct language for the voiceId.
          - `experimentalControls` CartesiaExperimentalControls
            - `speed` union
              - …
            - `emotion` 'anger:lowest' | 'anger:low' | 'anger:high' | 'anger:highest' | 'positivity:lowest' | 'positivity:low' | 'positivity:high' | 'positivity:highest' | 'surprise:lowest' | 'surprise:low' | 'surprise:high' | 'surprise:highest' | 'sadness:lowest' | 'sadness:low' | 'sadness:high' | 'sadness:highest' | 'curiosity:lowest' | 'curiosity:low' | 'curiosity:high' | 'curiosity:highest'
          - `generationConfig` CartesiaGenerationConfig
            - `speed` number — Fine-grained speed control for sonic-3. Only available for sonic-3 model.
            - `volume` number — Fine-grained volume control for sonic-3. Only available for sonic-3 model.
            - `experimental` CartesiaGenerationConfigExperimental
              - …
          - `pronunciationDictId` string — Pronunciation dictionary ID for sonic-3. Allows custom pronunciations for specific words. Only available for sonic-3 model.
          - `chunkPlan` ChunkPlan
            - `enabled` boolean — This determines whether the model output is chunked before being sent to the voice provider. Default `true`. Usage: - To rely on the voice provider's audio generation logic, set this to `false`. - If seeing issues with quality, set this to `true`. If disabled, Vapi-provided audio control tokens like <flush /> will not work. @default true
            - `minCharacters` number — This is the minimum number of characters in a chunk. Usage: - To increase quality, set this to a higher value. - To decrease latency, set this to a lower value. @default 30
            - `punctuationBoundaries` string[] — These are the punctuations that are considered valid boundaries for a chunk to be created. Usage: - To increase quality, constrain to fewer boundaries. - To decrease latency, enable all. Default is automatically set to balance the trade-off between quality and latency based on the provider.
            - `formatPlan` FormatPlan
              - …
          - `fallbackPlan` FallbackPlan
            - `voices` union[], required — This is the list of voices to fallback to in the event that the primary voice provider fails.
              - …
        - CustomVoice
          - `cachingEnabled` boolean — This is the flag to toggle voice caching for the assistant.
          - `provider` 'custom-voice', required — This is the voice provider that will be used. Use `custom-voice` for providers that are not natively supported.
          - `voiceId` string — This is the provider-specific ID that will be used. This is passed in the voice request payload to identify the voice to use.
          - `chunkPlan` ChunkPlan
            - `enabled` boolean — This determines whether the model output is chunked before being sent to the voice provider. Default `true`. Usage: - To rely on the voice provider's audio generation logic, set this to `false`. - If seeing issues with quality, set this to `true`. If disabled, Vapi-provided audio control tokens like <flush /> will not work. @default true
            - `minCharacters` number — This is the minimum number of characters in a chunk. Usage: - To increase quality, set this to a higher value. - To decrease latency, set this to a lower value. @default 30
            - `punctuationBoundaries` string[] — These are the punctuations that are considered valid boundaries for a chunk to be created. Usage: - To increase quality, constrain to fewer boundaries. - To decrease latency, enable all. Default is automatically set to balance the trade-off between quality and latency based on the provider.
            - `formatPlan` FormatPlan
              - …
          - `server` Server, required
            - `timeoutSeconds` number — This is the timeout in seconds for the request. Defaults to 20 seconds. @default 20
            - `credentialId` string — The credential ID for server authentication
            - `staticIpAddressesEnabled` boolean — If enabled, requests will originate from a static set of IPs owned and managed by Vapi. @default false
            - `encryptedPaths` string[] — This is the paths to encrypt in the request body if credentialId and encryptionPlan are defined.
            - `url` string — This is where the request will be sent.
            - `headers` object — These are the headers to include in the request. Each key-value pair represents a header name and its value. Note: Specifying an Authorization header here will override the authorization provided by the `credentialId` (if provided). This is an anti-pattern and should be avoided outside of edge case scenarios.
            - `backoffPlan` BackoffPlan
              - …
          - `fallbackPlan` FallbackPlan
            - `voices` union[], required — This is the list of voices to fallback to in the event that the primary voice provider fails.
              - …
        - DeepgramVoice
          - `cachingEnabled` boolean — This is the flag to toggle voice caching for the assistant.
          - `provider` 'deepgram', required — This is the voice provider that will be used.
          - `voiceId` 'asteria' | 'luna' | 'stella' | 'athena' | 'hera' | 'orion' | 'arcas' | 'perseus' | 'angus' | 'orpheus' | 'helios' | 'zeus' | 'thalia' | 'andromeda' | 'helena' | 'apollo' | 'arcas' | 'aries' | 'amalthea' | 'asteria' | 'athena' | 'atlas' | 'aurora' | 'callista' | 'cora' | 'cordelia' | 'delia' | 'draco' | 'electra' | 'harmonia' | 'hera' | 'hermes' | 'hyperion' | 'iris' | 'janus' | 'juno' | 'jupiter' | 'luna' | 'mars' | 'minerva' | 'neptune' | 'odysseus' | 'ophelia' | 'orion' | 'orpheus' | 'pandora' | 'phoebe' | 'pluto' | 'saturn' | 'selene' | 'theia' | 'vesta' | 'zeus' | 'celeste' | 'estrella' | 'nestor' | 'sirio' | 'carina' | 'alvaro' | 'diana' | 'aquila' | 'selena' | 'javier' | 'viktoria' | 'kara' | 'fabian' | 'julius' | 'lara' | 'elara' | 'aurelia', required — This is the provider-specific ID that will be used.
          - `model` 'aura' | 'aura-2' — This is the model that will be used. Defaults to 'aura-2' when not specified.
          - `mipOptOut` boolean — If set to true, this will add mip_opt_out=true as a query parameter of all API requests. See https://developers.deepgram.com/docs/the-deepgram-model-improvement-partnership-program#want-to-opt-out This will only be used if you are using your own Deepgram API key. @default false
          - `chunkPlan` ChunkPlan
            - `enabled` boolean — This determines whether the model output is chunked before being sent to the voice provider. Default `true`. Usage: - To rely on the voice provider's audio generation logic, set this to `false`. - If seeing issues with quality, set this to `true`. If disabled, Vapi-provided audio control tokens like <flush /> will not work. @default true
            - `minCharacters` number — This is the minimum number of characters in a chunk. Usage: - To increase quality, set this to a higher value. - To decrease latency, set this to a lower value. @default 30
            - `punctuationBoundaries` string[] — These are the punctuations that are considered valid boundaries for a chunk to be created. Usage: - To increase quality, constrain to fewer boundaries. - To decrease latency, enable all. Default is automatically set to balance the trade-off between quality and latency based on the provider.
            - `formatPlan` FormatPlan
              - …
          - `fallbackPlan` FallbackPlan
            - `voices` union[], required — This is the list of voices to fallback to in the event that the primary voice provider fails.
              - …
        - ElevenLabsVoice
          - `cachingEnabled` boolean — This is the flag to toggle voice caching for the assistant.
          - `provider` '11labs', required — This is the voice provider that will be used.
          - `voiceId` union, required — This is the provider-specific ID that will be used. Ensure the Voice is present in your 11Labs Voice Library.
            - 'burt' | 'marissa' | 'andrea' | 'sarah' | 'phillip' | 'steve' | 'joseph' | 'myra' | 'paula' | 'ryan' | 'drew' | 'paul' | 'mrb' | 'matilda' | 'mark'
            - string
          - `stability` number — Defines the stability for voice settings.
          - `similarityBoost` number — Defines the similarity boost for voice settings.
          - `style` number — Defines the style for voice settings.
          - `useSpeakerBoost` boolean — Defines the use speaker boost for voice settings.
          - `speed` number — Defines the speed for voice settings.
          - `optimizeStreamingLatency` number — Defines the optimize streaming latency for voice settings. Defaults to 3.
          - `enableSsmlParsing` boolean — This enables the use of https://elevenlabs.io/docs/speech-synthesis/prompting#pronunciation. Defaults to false to save latency. @default false
          - `autoMode` boolean — Defines the auto mode for voice settings. Defaults to false.
          - `model` 'eleven_multilingual_v2' | 'eleven_turbo_v2' | 'eleven_turbo_v2_5' | 'eleven_flash_v2' | 'eleven_flash_v2_5' | 'eleven_monolingual_v1' | 'eleven_v3' — This is the model that will be used. Defaults to 'eleven_turbo_v2' if not specified.
          - `language` string — This is the language (ISO 639-1) that is enforced for the model. Currently only Turbo v2.5 supports language enforcement. For other models, an error will be returned if language code is provided.
          - `chunkPlan` ChunkPlan
            - `enabled` boolean — This determines whether the model output is chunked before being sent to the voice provider. Default `true`. Usage: - To rely on the voice provider's audio generation logic, set this to `false`. - If seeing issues with quality, set this to `true`. If disabled, Vapi-provided audio control tokens like <flush /> will not work. @default true
            - `minCharacters` number — This is the minimum number of characters in a chunk. Usage: - To increase quality, set this to a higher value. - To decrease latency, set this to a lower value. @default 30
            - `punctuationBoundaries` string[] — These are the punctuations that are considered valid boundaries for a chunk to be created. Usage: - To increase quality, constrain to fewer boundaries. - To decrease latency, enable all. Default is automatically set to balance the trade-off between quality and latency based on the provider.
            - `formatPlan` FormatPlan
              - …
          - `pronunciationDictionaryLocators` ElevenLabsPronunciationDictionaryLocator[] — This is the pronunciation dictionary locators to use.
            - `pronunciationDictionaryId` string, required — This is the ID of the pronunciation dictionary to use.
            - `versionId` string — This is the version ID of the pronunciation dictionary to use. Omit to use the dictionary's latest version.
          - `fallbackPlan` FallbackPlan
            - `voices` union[], required — This is the list of voices to fallback to in the event that the primary voice provider fails.
              - …
        - HumeVoice
          - `cachingEnabled` boolean — This is the flag to toggle voice caching for the assistant.
          - `provider` 'hume', required — This is the voice provider that will be used.
          - `model` 'octave' | 'octave2' — This is the model that will be used.
          - `voiceId` string, required — The ID of the particular voice you want to use.
          - `isCustomHumeVoice` boolean — Indicates whether the chosen voice is a preset Hume AI voice or a custom voice.
          - `chunkPlan` ChunkPlan
            - `enabled` boolean — This determines whether the model output is chunked before being sent to the voice provider. Default `true`. Usage: - To rely on the voice provider's audio generation logic, set this to `false`. - If seeing issues with quality, set this to `true`. If disabled, Vapi-provided audio control tokens like <flush /> will not work. @default true
            - `minCharacters` number — This is the minimum number of characters in a chunk. Usage: - To increase quality, set this to a higher value. - To decrease latency, set this to a lower value. @default 30
            - `punctuationBoundaries` string[] — These are the punctuations that are considered valid boundaries for a chunk to be created. Usage: - To increase quality, constrain to fewer boundaries. - To decrease latency, enable all. Default is automatically set to balance the trade-off between quality and latency based on the provider.
            - `formatPlan` FormatPlan
              - …
          - `description` string — Natural language instructions describing how the synthesized speech should sound, including but not limited to tone, intonation, pacing, and accent (e.g., 'a soft, gentle voice with a strong British accent'). If a Voice is specified in the request, this description serves as acting instructions. If no Voice is specified, a new voice is generated based on this description.
          - `fallbackPlan` FallbackPlan
            - `voices` union[], required — This is the list of voices to fallback to in the event that the primary voice provider fails.
              - …
        - LMNTVoice
          - `cachingEnabled` boolean — This is the flag to toggle voice caching for the assistant.
          - `provider` 'lmnt', required — This is the voice provider that will be used.
          - `voiceId` union, required — This is the provider-specific ID that will be used.
            - 'amy' | 'ansel' | 'autumn' | 'ava' | 'brandon' | 'caleb' | 'cassian' | 'chloe' | 'dalton' | 'daniel' | 'dustin' | 'elowen' | 'evander' | 'huxley' | 'james' | 'juniper' | 'kennedy' | 'lauren' | 'leah' | 'lily' | 'lucas' | 'magnus' | 'miles' | 'morgan' | 'natalie' | 'nathan' | 'noah' | 'nyssa' | 'oliver' | 'paige' | 'ryan' | 'sadie' | 'sophie' | 'stella' | 'terrence' | 'tyler' | 'vesper' | 'violet' | 'warrick' | 'zain' | 'zeke' | 'zoe'
            - string
          - `speed` number — This is the speed multiplier that will be used.
          - `language` union — Two letter ISO 639-1 language code. Use "auto" for auto-detection.
            - 'aa' | 'ab' | 'ae' | 'af' | 'ak' | 'am' | 'an' | 'ar' | 'as' | 'av' | 'ay' | 'az' | 'ba' | 'be' | 'bg' | 'bh' | 'bi' | 'bm' | 'bn' | 'bo' | 'br' | 'bs' | 'ca' | 'ce' | 'ch' | 'co' | 'cr' | 'cs' | 'cu' | 'cv' | 'cy' | 'da' | 'de' | 'dv' | 'dz' | 'ee' | 'el' | 'en' | 'eo' | 'es' | 'et' | 'eu' | 'fa' | 'ff' | 'fi' | 'fj' | 'fo' | 'fr' | 'fy' | 'ga' | 'gd' | 'gl' | 'gn' | 'gu' | 'gv' | 'ha' | 'he' | 'hi' | 'ho' | 'hr' | 'ht' | 'hu' | 'hy' | 'hz' | 'ia' | 'id' | 'ie' | 'ig' | 'ii' | 'ik' | 'io' | 'is' | 'it' | 'iu' | 'ja' | 'jv' | 'ka' | 'kg' | 'ki' | 'kj' | 'kk' | 'kl' | 'km' | 'kn' | 'ko' | 'kr' | 'ks' | 'ku' | 'kv' | 'kw' | 'ky' | 'la' | 'lb' | 'lg' | 'li' | 'ln' | 'lo' | 'lt' | 'lu' | 'lv' | 'mg' | 'mh' | 'mi' | 'mk' | 'ml' | 'mn' | 'mr' | 'ms' | 'mt' | 'my' | 'na' | 'nb' | 'nd' | 'ne' | 'ng' | 'nl' | 'nn' | 'no' | 'nr' | 'nv' | 'ny' | 'oc' | 'oj' | 'om' | 'or' | 'os' | 'pa' | 'pi' | 'pl' | 'ps' | 'pt' | 'qu' | 'rm' | 'rn' | 'ro' | 'ru' | 'rw' | 'sa' | 'sc' | 'sd' | 'se' | 'sg' | 'si' | 'sk' | 'sl' | 'sm' | 'sn' | 'so' | 'sq' | 'sr' | 'ss' | 'st' | 'su' | 'sv' | 'sw' | 'ta' | 'te' | 'tg' | 'th' | 'ti' | 'tk' | 'tl' | 'tn' | 'to' | 'tr' | 'ts' | 'tt' | 'tw' | 'ty' | 'ug' | 'uk' | 'ur' | 'uz' | 've' | 'vi' | 'vo' | 'wa' | 'wo' | 'xh' | 'yi' | 'yue' | 'yo' | 'za' | 'zh' | 'zu'
            - 'auto'
          - `chunkPlan` ChunkPlan
            - `enabled` boolean — This determines whether the model output is chunked before being sent to the voice provider. Default `true`. Usage: - To rely on the voice provider's audio generation logic, set this to `false`. - If seeing issues with quality, set this to `true`. If disabled, Vapi-provided audio control tokens like <flush /> will not work. @default true
            - `minCharacters` number — This is the minimum number of characters in a chunk. Usage: - To increase quality, set this to a higher value. - To decrease latency, set this to a lower value. @default 30
            - `punctuationBoundaries` string[] — These are the punctuations that are considered valid boundaries for a chunk to be created. Usage: - To increase quality, constrain to fewer boundaries. - To decrease latency, enable all. Default is automatically set to balance the trade-off between quality and latency based on the provider.
            - `formatPlan` FormatPlan
              - …
          - `fallbackPlan` FallbackPlan
            - `voices` union[], required — This is the list of voices to fallback to in the event that the primary voice provider fails.
              - …
        - NeuphonicVoice
          - `cachingEnabled` boolean — This is the flag to toggle voice caching for the assistant.
          - `provider` 'neuphonic', required — This is the voice provider that will be used.
          - `voiceId` union, required — This is the provider-specific ID that will be used.
            - string
            - string
          - `model` 'neu_hq' | 'neu_fast' — This is the model that will be used. Defaults to 'neu_fast' if not specified.
          - `language` object, required — This is the language (ISO 639-1) that is enforced for the model.
          - `speed` number — This is the speed multiplier that will be used.
          - `chunkPlan` ChunkPlan
            - `enabled` boolean — This determines whether the model output is chunked before being sent to the voice provider. Default `true`. Usage: - To rely on the voice provider's audio generation logic, set this to `false`. - If seeing issues with quality, set this to `true`. If disabled, Vapi-provided audio control tokens like <flush /> will not work. @default true
            - `minCharacters` number — This is the minimum number of characters in a chunk. Usage: - To increase quality, set this to a higher value. - To decrease latency, set this to a lower value. @default 30
            - `punctuationBoundaries` string[] — These are the punctuations that are considered valid boundaries for a chunk to be created. Usage: - To increase quality, constrain to fewer boundaries. - To decrease latency, enable all. Default is automatically set to balance the trade-off between quality and latency based on the provider.
            - `formatPlan` FormatPlan
              - …
          - `fallbackPlan` FallbackPlan
            - `voices` union[], required — This is the list of voices to fallback to in the event that the primary voice provider fails.
              - …
        - OpenAIVoice
          - `cachingEnabled` boolean — This is the flag to toggle voice caching for the assistant.
          - `provider` 'openai', required — This is the voice provider that will be used.
          - `voiceId` union, required — This is the provider-specific ID that will be used. Please note that ash, ballad, coral, sage, and verse may only be used with realtime models.
            - 'alloy' | 'echo' | 'fable' | 'onyx' | 'nova' | 'shimmer' | 'marin' | 'cedar'
            - string
          - `model` 'tts-1' | 'tts-1-hd' | 'gpt-4o-mini-tts' — This is the model that will be used for text-to-speech.
          - `instructions` string — This is a prompt that allows you to control the voice of your generated audio. Does not work with 'tts-1' or 'tts-1-hd' models.
          - `speed` number — This is the speed multiplier that will be used.
          - `chunkPlan` ChunkPlan
            - `enabled` boolean — This determines whether the model output is chunked before being sent to the voice provider. Default `true`. Usage: - To rely on the voice provider's audio generation logic, set this to `false`. - If seeing issues with quality, set this to `true`. If disabled, Vapi-provided audio control tokens like <flush /> will not work. @default true
            - `minCharacters` number — This is the minimum number of characters in a chunk. Usage: - To increase quality, set this to a higher value. - To decrease latency, set this to a lower value. @default 30
            - `punctuationBoundaries` string[] — These are the punctuations that are considered valid boundaries for a chunk to be created. Usage: - To increase quality, constrain to fewer boundaries. - To decrease latency, enable all. Default is automatically set to balance the trade-off between quality and latency based on the provider.
            - `formatPlan` FormatPlan
              - …
          - `fallbackPlan` FallbackPlan
            - `voices` union[], required — This is the list of voices to fallback to in the event that the primary voice provider fails.
              - …
        - PlayHTVoice
          - `cachingEnabled` boolean — This is the flag to toggle voice caching for the assistant.
          - `provider` 'playht', required — This is the voice provider that will be used.
          - `voiceId` union, required — This is the provider-specific ID that will be used.
            - 'jennifer' | 'melissa' | 'will' | 'chris' | 'matt' | 'jack' | 'ruby' | 'davis' | 'donna' | 'michael'
            - string
          - `speed` number — This is the speed multiplier that will be used.
          - `temperature` number — A floating point number between 0, exclusive, and 2, inclusive. If equal to null or not provided, the model's default temperature will be used. The temperature parameter controls variance. Lower temperatures result in more predictable results, higher temperatures allow each run to vary more, so the voice may sound less like the baseline voice.
          - `emotion` 'female_happy' | 'female_sad' | 'female_angry' | 'female_fearful' | 'female_disgust' | 'female_surprised' | 'male_happy' | 'male_sad' | 'male_angry' | 'male_fearful' | 'male_disgust' | 'male_surprised' — An emotion to be applied to the speech.
          - `voiceGuidance` number — A number between 1 and 6. Use lower numbers to reduce how unique your chosen voice will be compared to other voices.
          - `styleGuidance` number — A number between 1 and 30. Use lower numbers to to reduce how strong your chosen emotion will be. Higher numbers will create a very emotional performance.
          - `textGuidance` number — A number between 1 and 2. This number influences how closely the generated speech adheres to the input text. Use lower values to create more fluid speech, but with a higher chance of deviating from the input text. Higher numbers will make the generated speech more accurate to the input text, ensuring that the words spoken align closely with the provided text.
          - `model` 'PlayHT2.0' | 'PlayHT2.0-turbo' | 'Play3.0-mini' | 'PlayDialog' — Playht voice model/engine to use.
          - `language` 'afrikaans' | 'albanian' | 'amharic' | 'arabic' | 'bengali' | 'bulgarian' | 'catalan' | 'croatian' | 'czech' | 'danish' | 'dutch' | 'english' | 'french' | 'galician' | 'german' | 'greek' | 'hebrew' | 'hindi' | 'hungarian' | 'indonesian' | 'italian' | 'japanese' | 'korean' | 'malay' | 'mandarin' | 'polish' | 'portuguese' | 'russian' | 'serbian' | 'spanish' | 'swedish' | 'tagalog' | 'thai' | 'turkish' | 'ukrainian' | 'urdu' | 'xhosa' — The language to use for the speech.
          - `chunkPlan` ChunkPlan
            - `enabled` boolean — This determines whether the model output is chunked before being sent to the voice provider. Default `true`. Usage: - To rely on the voice provider's audio generation logic, set this to `false`. - If seeing issues with quality, set this to `true`. If disabled, Vapi-provided audio control tokens like <flush /> will not work. @default true
            - `minCharacters` number — This is the minimum number of characters in a chunk. Usage: - To increase quality, set this to a higher value. - To decrease latency, set this to a lower value. @default 30
            - `punctuationBoundaries` string[] — These are the punctuations that are considered valid boundaries for a chunk to be created. Usage: - To increase quality, constrain to fewer boundaries. - To decrease latency, enable all. Default is automatically set to balance the trade-off between quality and latency based on the provider.
            - `formatPlan` FormatPlan
              - …
          - `fallbackPlan` FallbackPlan
            - `voices` union[], required — This is the list of voices to fallback to in the event that the primary voice provider fails.
              - …
        - WellSaidVoice
          - `cachingEnabled` boolean — This is the flag to toggle voice caching for the assistant.
          - `provider` 'wellsaid', required — This is the voice provider that will be used.
          - `voiceId` string, required — The WellSaid speaker ID to synthesize.
          - `model` 'caruso' | 'legacy' — This is the model that will be used.
          - `enableSsml` boolean — Enables limited SSML translation for input text.
          - `libraryIds` string[] — Array of library IDs to use for voice synthesis.
          - `chunkPlan` ChunkPlan
            - `enabled` boolean — This determines whether the model output is chunked before being sent to the voice provider. Default `true`. Usage: - To rely on the voice provider's audio generation logic, set this to `false`. - If seeing issues with quality, set this to `true`. If disabled, Vapi-provided audio control tokens like <flush /> will not work. @default true
            - `minCharacters` number — This is the minimum number of characters in a chunk. Usage: - To increase quality, set this to a higher value. - To decrease latency, set this to a lower value. @default 30
            - `punctuationBoundaries` string[] — These are the punctuations that are considered valid boundaries for a chunk to be created. Usage: - To increase quality, constrain to fewer boundaries. - To decrease latency, enable all. Default is automatically set to balance the trade-off between quality and latency based on the provider.
            - `formatPlan` FormatPlan
              - …
          - `fallbackPlan` FallbackPlan
            - `voices` union[], required — This is the list of voices to fallback to in the event that the primary voice provider fails.
              - …
        - RimeAIVoice
          - `cachingEnabled` boolean — This is the flag to toggle voice caching for the assistant.
          - `provider` 'rime-ai', required — This is the voice provider that will be used.
          - `voiceId` union, required — This is the provider-specific ID that will be used.
            - 'cove' | 'moon' | 'wildflower' | 'eva' | 'amber' | 'maya' | 'lagoon' | 'breeze' | 'helen' | 'joy' | 'marsh' | 'creek' | 'cedar' | 'alpine' | 'summit' | 'nicholas' | 'tyler' | 'colin' | 'hank' | 'thunder' | 'astra' | 'eucalyptus' | 'moraine' | 'peak' | 'tundra' | 'mesa_extra' | 'talon' | 'marlu' | 'glacier' | 'falcon' | 'luna' | 'celeste' | 'estelle' | 'andromeda' | 'esther' | 'lyra' | 'lintel' | 'oculus' | 'vespera' | 'transom' | 'bond' | 'arcade' | 'atrium' | 'cupola' | 'fern' | 'sirius' | 'orion' | 'masonry' | 'albion' | 'parapet' — Popular Rime AI voices across mist, mistv2, and arcana models. Any valid Rime AI voice ID is accepted, not just these suggestions.
            - string — Any valid Rime AI voice ID. See https://docs.rime.ai/docs/voices for the full catalog.
          - `model` 'arcana' | 'mistv2' | 'mist' — This is the model that will be used. Defaults to 'arcana' when not specified.
          - `speed` number — This is the speed multiplier that will be used.
          - `pauseBetweenBrackets` boolean — This is a flag that controls whether to add slight pauses using angle brackets. Example: "Hi. <200> I'd love to have a conversation with you." adds a 200ms pause between the first and second sentences.
          - `phonemizeBetweenBrackets` boolean — This is a flag that controls whether text inside brackets should be phonemized (converted to phonetic pronunciation) - Example: "{h'El.o} World" will pronounce "Hello" as expected.
          - `reduceLatency` boolean — This is a flag that controls whether to optimize for reduced latency in streaming. https://docs.rime.ai/api-reference/endpoint/websockets#param-reduce-latency
          - `inlineSpeedAlpha` string — This is a string that allows inline speed control using alpha notation. https://docs.rime.ai/api-reference/endpoint/websockets#param-inline-speed-alpha
          - `language` 'en' | 'es' | 'de' | 'fr' | 'ar' | 'hi' | 'ja' | 'he' | 'pt' | 'ta' | 'si' — Language for speech synthesis. Uses ISO 639 codes. Supported: en, es, de, fr, ar, hi, ja, he, pt, ta, si.
          - `chunkPlan` ChunkPlan
            - `enabled` boolean — This determines whether the model output is chunked before being sent to the voice provider. Default `true`. Usage: - To rely on the voice provider's audio generation logic, set this to `false`. - If seeing issues with quality, set this to `true`. If disabled, Vapi-provided audio control tokens like <flush /> will not work. @default true
            - `minCharacters` number — This is the minimum number of characters in a chunk. Usage: - To increase quality, set this to a higher value. - To decrease latency, set this to a lower value. @default 30
            - `punctuationBoundaries` string[] — These are the punctuations that are considered valid boundaries for a chunk to be created. Usage: - To increase quality, constrain to fewer boundaries. - To decrease latency, enable all. Default is automatically set to balance the trade-off between quality and latency based on the provider.
            - `formatPlan` FormatPlan
              - …
          - `fallbackPlan` FallbackPlan
            - `voices` union[], required — This is the list of voices to fallback to in the event that the primary voice provider fails.
              - …
        - SmallestAIVoice
          - `cachingEnabled` boolean — This is the flag to toggle voice caching for the assistant.
          - `provider` 'smallest-ai', required — This is the voice provider that will be used.
          - `voiceId` union, required — This is the provider-specific ID that will be used.
            - 'emily' | 'jasmine' | 'arman' | 'james' | 'mithali' | 'aravind' | 'raj' | 'diya' | 'raman' | 'ananya' | 'isha' | 'william' | 'aarav' | 'monika' | 'niharika' | 'deepika' | 'raghav' | 'kajal' | 'radhika' | 'mansi' | 'nisha' | 'saurabh' | 'pooja' | 'saina' | 'sanya'
            - string
          - `model` 'lightning' — Smallest AI voice model to use. Defaults to 'lightning' when not specified.
          - `speed` number — This is the speed multiplier that will be used.
          - `chunkPlan` ChunkPlan
            - `enabled` boolean — This determines whether the model output is chunked before being sent to the voice provider. Default `true`. Usage: - To rely on the voice provider's audio generation logic, set this to `false`. - If seeing issues with quality, set this to `true`. If disabled, Vapi-provided audio control tokens like <flush /> will not work. @default true
            - `minCharacters` number — This is the minimum number of characters in a chunk. Usage: - To increase quality, set this to a higher value. - To decrease latency, set this to a lower value. @default 30
            - `punctuationBoundaries` string[] — These are the punctuations that are considered valid boundaries for a chunk to be created. Usage: - To increase quality, constrain to fewer boundaries. - To decrease latency, enable all. Default is automatically set to balance the trade-off between quality and latency based on the provider.
            - `formatPlan` FormatPlan
              - …
          - `fallbackPlan` FallbackPlan
            - `voices` union[], required — This is the list of voices to fallback to in the event that the primary voice provider fails.
              - …
        - TavusVoice
          - `cachingEnabled` boolean — This is the flag to toggle voice caching for the assistant.
          - `provider` 'tavus', required — This is the voice provider that will be used.
          - `voiceId` union, required — This is the provider-specific ID that will be used.
            - 'r52da2535a'
            - string
          - `chunkPlan` ChunkPlan
            - `enabled` boolean — This determines whether the model output is chunked before being sent to the voice provider. Default `true`. Usage: - To rely on the voice provider's audio generation logic, set this to `false`. - If seeing issues with quality, set this to `true`. If disabled, Vapi-provided audio control tokens like <flush /> will not work. @default true
            - `minCharacters` number — This is the minimum number of characters in a chunk. Usage: - To increase quality, set this to a higher value. - To decrease latency, set this to a lower value. @default 30
            - `punctuationBoundaries` string[] — These are the punctuations that are considered valid boundaries for a chunk to be created. Usage: - To increase quality, constrain to fewer boundaries. - To decrease latency, enable all. Default is automatically set to balance the trade-off between quality and latency based on the provider.
            - `formatPlan` FormatPlan
              - …
          - `personaId` string — This is the unique identifier for the persona that the replica will use in the conversation.
          - `callbackUrl` string — This is the url that will receive webhooks with updates regarding the conversation state.
          - `conversationName` string — This is the name for the conversation.
          - `conversationalContext` string — This is the context that will be appended to any context provided in the persona, if one is provided.
          - `customGreeting` string — This is the custom greeting that the replica will give once a participant joines the conversation.
          - `properties` TavusConversationProperties
            - `maxCallDuration` number — The maximum duration of the call in seconds. The default `maxCallDuration` is 3600 seconds (1 hour). Once the time limit specified by this parameter has been reached, the conversation will automatically shut down.
            - `participantLeftTimeout` number — The duration in seconds after which the call will be automatically shut down once the last participant leaves.
            - `participantAbsentTimeout` number — Starting from conversation creation, the duration in seconds after which the call will be automatically shut down if no participant joins the call. Default is 300 seconds (5 minutes).
            - `enableRecording` boolean — If true, the user will be able to record the conversation.
            - `enableTranscription` boolean — If true, the user will be able to transcribe the conversation. You can find more instructions on displaying transcriptions if you are using your custom DailyJS components here. You need to have an event listener on Daily that listens for `app-messages`.
            - `applyGreenscreen` boolean — If true, the background will be replaced with a greenscreen (RGB values: `[0, 255, 155]`). You can use WebGL on the frontend to make the greenscreen transparent or change its color.
            - `language` string — The language of the conversation. Please provide the **full language name**, not the two-letter code. If you are using your own TTS voice, please ensure it supports the language you provide. If you are using a stock replica or default persona, please note that only ElevenLabs and Cartesia supported languages are available. You can find a full list of supported languages for Cartesia here, for ElevenLabs here, and for PlayHT here.
            - `recordingS3BucketName` string — The name of the S3 bucket where the recording will be stored.
            - `recordingS3BucketRegion` string — The region of the S3 bucket where the recording will be stored.
            - `awsAssumeRoleArn` string — The ARN of the role that will be assumed to access the S3 bucket.
          - `fallbackPlan` FallbackPlan
            - `voices` union[], required — This is the list of voices to fallback to in the event that the primary voice provider fails.
              - …
        - VapiVoice
          - `cachingEnabled` boolean — This is the flag to toggle voice caching for the assistant.
          - `provider` 'vapi', required — This is the voice provider that will be used.
          - `voiceId` string, required — The voice to use: a built-in Vapi voice name, or a cloned voice id (used with version 2).
          - `version` '1' | '2' | 'latest' — The Vapi voice routing generation. `latest` auto-updates to the newest generation; version 1 uses legacy mappings; version 2 can use xAI-backed voices when available. When omitted, Version 1 is used. Accepts the string channel ('latest', '1', '2'); legacy numeric values (1, 2) are also accepted and coerced to their string form.
          - `speed` number — This is the speed multiplier that will be used. @default 1
          - `language` 'en-US' | 'en-GB' | 'en-AU' | 'en-CA' | 'ja' | 'zh' | 'de' | 'hi' | 'fr-FR' | 'fr-CA' | 'ko' | 'pt-BR' | 'pt-PT' | 'it' | 'es-ES' | 'es-MX' | 'id' | 'nl' | 'tr' | 'fil' | 'pl' | 'sv' | 'bg' | 'ro' | 'ar-SA' | 'ar-AE' | 'cs' | 'el' | 'fi' | 'hr' | 'ms' | 'sk' | 'da' | 'ta' | 'uk' | 'ru' | 'hu' | 'no' | 'vi' | 'auto' | 'en' | 'ar' | 'ar-EG' | 'bn' | 'es' | 'fr' | 'gu' | 'he' | 'ka' | 'kn' | 'ml' | 'mr' | 'pa' | 'pt' | 'te' | 'th' | 'tl' — Language for Vapi voice synthesis. For Version 2, omit this field or set `auto` for automatic language detection. Version 1 supports legacy Vapi language values.
          - `pronunciationDictionary` VapiPronunciationDictionaryLocator[] — List of pronunciation dictionary locators for custom word pronunciations.
            - `pronunciationDictId` string, required — The pronunciation dictionary ID
            - `versionId` string — Version ID (only used by ElevenLabs, ignored for Cartesia)
            - `provider` 'cartesia' | '11labs' — Provider that hosts this pronunciation dictionary
          - `chunkPlan` ChunkPlan
            - `enabled` boolean — This determines whether the model output is chunked before being sent to the voice provider. Default `true`. Usage: - To rely on the voice provider's audio generation logic, set this to `false`. - If seeing issues with quality, set this to `true`. If disabled, Vapi-provided audio control tokens like <flush /> will not work. @default true
            - `minCharacters` number — This is the minimum number of characters in a chunk. Usage: - To increase quality, set this to a higher value. - To decrease latency, set this to a lower value. @default 30
            - `punctuationBoundaries` string[] — These are the punctuations that are considered valid boundaries for a chunk to be created. Usage: - To increase quality, constrain to fewer boundaries. - To decrease latency, enable all. Default is automatically set to balance the trade-off between quality and latency based on the provider.
            - `formatPlan` FormatPlan
              - …
        - SesameVoice
          - `cachingEnabled` boolean — This is the flag to toggle voice caching for the assistant.
          - `provider` 'sesame', required — This is the voice provider that will be used.
          - `voiceId` string, required — This is the provider-specific ID that will be used.
          - `model` 'csm-1b', required — This is the model that will be used.
          - `chunkPlan` ChunkPlan
            - `enabled` boolean — This determines whether the model output is chunked before being sent to the voice provider. Default `true`. Usage: - To rely on the voice provider's audio generation logic, set this to `false`. - If seeing issues with quality, set this to `true`. If disabled, Vapi-provided audio control tokens like <flush /> will not work. @default true
            - `minCharacters` number — This is the minimum number of characters in a chunk. Usage: - To increase quality, set this to a higher value. - To decrease latency, set this to a lower value. @default 30
            - `punctuationBoundaries` string[] — These are the punctuations that are considered valid boundaries for a chunk to be created. Usage: - To increase quality, constrain to fewer boundaries. - To decrease latency, enable all. Default is automatically set to balance the trade-off between quality and latency based on the provider.
            - `formatPlan` FormatPlan
              - …
          - `fallbackPlan` FallbackPlan
            - `voices` union[], required — This is the list of voices to fallback to in the event that the primary voice provider fails.
              - …
        - InworldVoice
          - `cachingEnabled` boolean — This is the flag to toggle voice caching for the assistant.
          - `provider` 'inworld', required — This is the voice provider that will be used.
          - `voiceId` 'Alex' | 'Ashley' | 'Craig' | 'Deborah' | 'Dennis' | 'Edward' | 'Elizabeth' | 'Hades' | 'Julia' | 'Pixie' | 'Mark' | 'Olivia' | 'Priya' | 'Ronald' | 'Sarah' | 'Shaun' | 'Theodore' | 'Timothy' | 'Wendy' | 'Dominus' | 'Hana' | 'Clive' | 'Carter' | 'Blake' | 'Luna' | 'Yichen' | 'Xiaoyin' | 'Xinyi' | 'Jing' | 'Erik' | 'Katrien' | 'Lennart' | 'Lore' | 'Alain' | 'Hélène' | 'Mathieu' | 'Étienne' | 'Johanna' | 'Josef' | 'Gianni' | 'Orietta' | 'Asuka' | 'Satoshi' | 'Hyunwoo' | 'Minji' | 'Seojun' | 'Yoona' | 'Szymon' | 'Wojciech' | 'Heitor' | 'Maitê' | 'Diego' | 'Lupita' | 'Miguel' | 'Rafael' | 'Svetlana' | 'Elena' | 'Dmitry' | 'Nikolai' | 'Riya' | 'Manoj' | 'Yael' | 'Oren' | 'Nour' | 'Omar', required — Available voices by language: • en: Alex, Ashley, Craig, Deborah, Dennis, Edward, Elizabeth, Hades, Julia, Pixie, Mark, Olivia, Priya, Ronald, Sarah, Shaun, Theodore, Timothy, Wendy, Dominus, Hana, Clive, Carter, Blake, Luna • zh: Yichen, Xiaoyin, Xinyi, Jing • nl: Erik, Katrien, Lennart, Lore • fr: Alain, Hélène, Mathieu, Étienne • de: Johanna, Josef • it: Gianni, Orietta • ja: Asuka, Satoshi • ko: Hyunwoo, Minji, Seojun, Yoona • pl: Szymon, Wojciech • pt: Heitor, Maitê • es: Diego, Lupita, Miguel, Rafael • ru: Svetlana, Elena, Dmitry, Nikolai • hi: Riya, Manoj • he: Yael, Oren • ar: Nour, Omar
          - `model` 'inworld-tts-1' — This is the model that will be used.
          - `languageCode` 'en' | 'zh' | 'ko' | 'nl' | 'fr' | 'es' | 'ja' | 'de' | 'it' | 'pl' | 'pt' | 'ru' | 'hi' | 'he' | 'ar' — Language code for Inworld TTS synthesis
          - `temperature` number — A floating point number between 0, exclusive, and 2, inclusive. If equal to null or not provided, the model's default temperature of 1.1 will be used. The temperature parameter controls variance. Higher values will make the output more random and can lead to more expressive results. Lower values will make it more deterministic. See https://docs.inworld.ai/docs/tts/capabilities/generating-audio#additional-configurations for more details.
          - `speakingRate` number — A floating point number between 0.5, inclusive, and 1.5, inclusive. If equal to null or not provided, the model's default speaking speed of 1.0 will be used. Values above 0.8 are recommended for higher quality. See https://docs.inworld.ai/docs/tts/capabilities/generating-audio#additional-configurations for more details.
          - `chunkPlan` ChunkPlan
            - `enabled` boolean — This determines whether the model output is chunked before being sent to the voice provider. Default `true`. Usage: - To rely on the voice provider's audio generation logic, set this to `false`. - If seeing issues with quality, set this to `true`. If disabled, Vapi-provided audio control tokens like <flush /> will not work. @default true
            - `minCharacters` number — This is the minimum number of characters in a chunk. Usage: - To increase quality, set this to a higher value. - To decrease latency, set this to a lower value. @default 30
            - `punctuationBoundaries` string[] — These are the punctuations that are considered valid boundaries for a chunk to be created. Usage: - To increase quality, constrain to fewer boundaries. - To decrease latency, enable all. Default is automatically set to balance the trade-off between quality and latency based on the provider.
            - `formatPlan` FormatPlan
              - …
          - `fallbackPlan` FallbackPlan
            - `voices` union[], required — This is the list of voices to fallback to in the event that the primary voice provider fails.
              - …
        - MinimaxVoice
          - `cachingEnabled` boolean — This is the flag to toggle voice caching for the assistant.
          - `provider` 'minimax', required — This is the voice provider that will be used.
          - `voiceId` string, required — This is the provider-specific ID that will be used. Use a voice from MINIMAX_PREDEFINED_VOICES or a custom cloned voice ID.
          - `model` 'speech-02-hd' | 'speech-02-turbo' | 'speech-2.5-turbo-preview' — This is the model that will be used. Options are 'speech-02-hd' and 'speech-02-turbo'. speech-02-hd is optimized for high-fidelity applications like voiceovers and audiobooks. speech-02-turbo is designed for real-time applications with low latency. @default "speech-02-turbo"
          - `emotion` string — The emotion to use for the voice. If not provided, will use auto-detect mode. Options include: 'happy', 'sad', 'angry', 'fearful', 'surprised', 'disgusted', 'neutral'
          - `subtitleType` 'word' | 'sentence' — Controls the granularity of subtitle/timing data returned by Minimax during synthesis. Set to 'word' to receive per-word timestamps in assistant.speechStarted events for karaoke-style caption rendering. @default "sentence"
          - `pitch` number — Voice pitch adjustment. Range from -12 to 12 semitones. @default 0
          - `speed` number — Voice speed adjustment. Range from 0.5 to 2.0. @default 1.0
          - `volume` number — Voice volume adjustment. Range from 0.5 to 2.0. @default 1.0
          - `region` 'worldwide' | 'china' — The region for Minimax API. Defaults to "worldwide".
          - `languageBoost` 'Chinese' | 'Chinese,Yue' | 'English' | 'Arabic' | 'Russian' | 'Spanish' | 'French' | 'Portuguese' | 'German' | 'Turkish' | 'Dutch' | 'Ukrainian' | 'Vietnamese' | 'Indonesian' | 'Japanese' | 'Italian' | 'Korean' | 'Thai' | 'Polish' | 'Romanian' | 'Greek' | 'Czech' | 'Finnish' | 'Hindi' | 'Bulgarian' | 'Danish' | 'Hebrew' | 'Malay' | 'Persian' | 'Slovak' | 'Swedish' | 'Croatian' | 'Filipino' | 'Hungarian' | 'Norwegian' | 'Slovenian' | 'Catalan' | 'Nynorsk' | 'Tamil' | 'Afrikaans' | 'auto' — Language hint for MiniMax T2A. Example: yue (Cantonese), zh (Chinese), en (English).
          - `textNormalizationEnabled` boolean — Enable MiniMax text normalization to improve number reading and formatting.
          - `chunkPlan` ChunkPlan
            - `enabled` boolean — This determines whether the model output is chunked before being sent to the voice provider. Default `true`. Usage: - To rely on the voice provider's audio generation logic, set this to `false`. - If seeing issues with quality, set this to `true`. If disabled, Vapi-provided audio control tokens like <flush /> will not work. @default true
            - `minCharacters` number — This is the minimum number of characters in a chunk. Usage: - To increase quality, set this to a higher value. - To decrease latency, set this to a lower value. @default 30
            - `punctuationBoundaries` string[] — These are the punctuations that are considered valid boundaries for a chunk to be created. Usage: - To increase quality, constrain to fewer boundaries. - To decrease latency, enable all. Default is automatically set to balance the trade-off between quality and latency based on the provider.
            - `formatPlan` FormatPlan
              - …
          - `fallbackPlan` FallbackPlan
            - `voices` union[], required — This is the list of voices to fallback to in the event that the primary voice provider fails.
              - …
        - XaiVoice
          - `cachingEnabled` boolean — This is the flag to toggle voice caching for the assistant.
          - `provider` 'xai', required — This is the voice provider that will be used.
          - `voiceId` 'eve' | 'ara' | 'rex' | 'sal' | 'leo', required — Built-in voices: eve, ara, rex, sal, leo. Cloned voice IDs are also accepted.
          - `language` 'auto' | 'en' | 'ar-EG' | 'ar-SA' | 'ar-AE' | 'bn' | 'zh' | 'fr' | 'de' | 'hi' | 'id' | 'it' | 'ja' | 'ko' | 'pt-BR' | 'pt-PT' | 'ru' | 'es-MX' | 'es-ES' | 'tr' | 'vi' — BCP-47 language code for xAI TTS synthesis.
          - `speed` number — Speed multiplier for xAI TTS synthesis.
          - `chunkPlan` ChunkPlan
            - `enabled` boolean — This determines whether the model output is chunked before being sent to the voice provider. Default `true`. Usage: - To rely on the voice provider's audio generation logic, set this to `false`. - If seeing issues with quality, set this to `true`. If disabled, Vapi-provided audio control tokens like <flush /> will not work. @default true
            - `minCharacters` number — This is the minimum number of characters in a chunk. Usage: - To increase quality, set this to a higher value. - To decrease latency, set this to a lower value. @default 30
            - `punctuationBoundaries` string[] — These are the punctuations that are considered valid boundaries for a chunk to be created. Usage: - To increase quality, constrain to fewer boundaries. - To decrease latency, enable all. Default is automatically set to balance the trade-off between quality and latency based on the provider.
            - `formatPlan` FormatPlan
              - …
          - `fallbackPlan` FallbackPlan
            - `voices` union[], required — This is the list of voices to fallback to in the event that the primary voice provider fails.
              - …
        - MicrosoftVoice
          - `cachingEnabled` boolean — This is the flag to toggle voice caching for the assistant.
          - `provider` 'microsoft', required — This is the voice provider that will be used.
          - `voiceId` 'de-DE-Klaus:MAI-Voice-2' | 'de-DE-Mia:MAI-Voice-2' | 'en-AU-Lisa:MAI-Voice-2' | 'en-US-Ethan:MAI-Voice-2' | 'en-US-Grant:MAI-Voice-2' | 'en-US-Harper:MAI-Voice-2' | 'en-US-Iris:MAI-Voice-2' | 'en-US-Jasper:MAI-Voice-2' | 'en-US-Olivia:MAI-Voice-2' | 'es-ES-Marta:MAI-Voice-2' | 'es-MX-Alejo:MAI-Voice-2' | 'es-MX-Valeria:MAI-Voice-2' | 'fr-FR-Marc:MAI-Voice-2' | 'fr-FR-Soleil:MAI-Voice-2' | 'hi-IN-Arjun:MAI-Voice-2' | 'hi-IN-Dhruv:MAI-Voice-2' | 'hi-IN-Kavya:MAI-Voice-2' | 'hi-IN-Priya:MAI-Voice-2' | 'hu-HU-Bence:MAI-Voice-2' | 'hu-HU-Levente:MAI-Voice-2' | 'hu-HU-Lilla:MAI-Voice-2' | 'hu-HU-Réka:MAI-Voice-2' | 'it-IT-Luca:MAI-Voice-2' | 'it-IT-Rosa:MAI-Voice-2' | 'ko-KR-Hana:MAI-Voice-2' | 'ko-KR-Junho:MAI-Voice-2' | 'nl-NL-Fleur:MAI-Voice-2' | 'nl-NL-Sander:MAI-Voice-2' | 'pt-BR-Caio:MAI-Voice-2' | 'pt-BR-Luana:MAI-Voice-2' | 'pt-BR-Pedro:MAI-Voice-2' | 'pt-BR-Rafael:MAI-Voice-2' | 'pt-PT-Rui:MAI-Voice-2' | 'ro-RO-Andrei:MAI-Voice-2' | 'ro-RO-Elena:MAI-Voice-2' | 'ro-RO-Ioana:MAI-Voice-2' | 'ro-RO-Radu:MAI-Voice-2' | 'ru-RU-Lev:MAI-Voice-2' | 'ru-RU-Masha:MAI-Voice-2' | 'th-TH-Krit:MAI-Voice-2' | 'th-TH-Nattapong:MAI-Voice-2' | 'tr-TR-Aydin:MAI-Voice-2' | 'tr-TR-Elif:MAI-Voice-2' | 'zh-CN-Bo:MAI-Voice-2' | 'zh-CN-Lan:MAI-Voice-2' | 'zh-CN-Mei:MAI-Voice-2', required — MAI-Voice-2 voice ID. Built-in voices listed in enum.
          - `style` 'adventurous' | 'angry' | 'caring' | 'cheerful' | 'confused' | 'curious' | 'determined' | 'disappointed' | 'disgusted' | 'embarrassed' | 'empathy' | 'encouraging' | 'excited' | 'fearful' | 'friendly' | 'happy' | 'hopeful' | 'jealous' | 'joyful' | 'nostalgic' | 'reflective' | 'regretful' | 'relieved' | 'sad' | 'serious' | 'shouting' | 'softvoice' | 'surprised' | 'whispering' — Speaking style applied via mstts:express-as on every request. Unknown styles are ignored by Azure and fall back to neutral.
          - `styleDegree` number — Style intensity (0.01–2). Default 1 = the predefined style strength. Only applies when `style` is set.
          - `role` 'Girl' | 'Boy' | 'YoungAdultFemale' | 'YoungAdultMale' | 'OlderAdultFemale' | 'OlderAdultMale' | 'SeniorFemale' | 'SeniorMale' — Role-play (age/gender imitation). Requires `style` to be set; ignored otherwise.
          - `chunkPlan` ChunkPlan
            - `enabled` boolean — This determines whether the model output is chunked before being sent to the voice provider. Default `true`. Usage: - To rely on the voice provider's audio generation logic, set this to `false`. - If seeing issues with quality, set this to `true`. If disabled, Vapi-provided audio control tokens like <flush /> will not work. @default true
            - `minCharacters` number — This is the minimum number of characters in a chunk. Usage: - To increase quality, set this to a higher value. - To decrease latency, set this to a lower value. @default 30
            - `punctuationBoundaries` string[] — These are the punctuations that are considered valid boundaries for a chunk to be created. Usage: - To increase quality, constrain to fewer boundaries. - To decrease latency, enable all. Default is automatically set to balance the trade-off between quality and latency based on the provider.
            - `formatPlan` FormatPlan
              - …
          - `speed` number — This is the speed multiplier that will be used.
          - `fallbackPlan` FallbackPlan
            - `voices` union[], required — This is the list of voices to fallback to in the event that the primary voice provider fails.
              - …
      - `firstMessage` string — This is the first message that the assistant will say. This can also be a URL to a containerized audio file (mp3, wav, etc.). If unspecified, assistant will wait for user to speak and use the model to respond once they speak.
      - `firstMessageInterruptionsEnabled` boolean
      - `firstMessageMode` 'assistant-speaks-first' | 'assistant-speaks-first-with-model-generated-message' | 'assistant-waits-for-user' — This is the mode for the first message. Default is 'assistant-speaks-first'. Use: - 'assistant-speaks-first' to have the assistant speak first. - 'assistant-waits-for-user' to have the assistant wait for the user to speak first. - 'assistant-speaks-first-with-model-generated-message' to have the assistant speak first with a message generated by the model based on the conversation state. (`assistant.model.messages` at call start, `call.messages` at squad transfer points). @default 'assistant-speaks-first'
      - `voicemailDetection` union — These are the settings to configure or disable voicemail detection. Alternatively, voicemail detection can be configured using the model.tools=[VoicemailTool]. By default, voicemail detection is disabled.
        - 'off'
        - GoogleVoicemailDetectionPlan
          - `beepMaxAwaitSeconds` number — This is the maximum duration from the start of the call that we will wait for a voicemail beep, before speaking our message - If we detect a voicemail beep before this, we will speak the message at that point. - Setting too low a value means that the bot will start speaking its voicemail message too early. If it does so before the actual beep, it will get cut off. You should definitely tune this to your use case. @default 30 @min 0 @max 60
          - `provider` 'google', required — This is the provider to use for voicemail detection.
          - `backoffPlan` VoicemailDetectionBackoffPlan
            - `startAtSeconds` number — This is the number of seconds to wait before starting the first retry attempt.
            - `frequencySeconds` number — This is the interval in seconds between retry attempts.
            - `maxRetries` number — This is the maximum number of retry attempts before giving up.
          - `type` 'audio' | 'transcript' — This is the detection type to use for voicemail detection. - 'audio': Uses native audio models (default) - 'transcript': Uses ASR/transcript-based detection @default 'audio' (audio detection)
        - OpenAIVoicemailDetectionPlan
          - `beepMaxAwaitSeconds` number — This is the maximum duration from the start of the call that we will wait for a voicemail beep, before speaking our message - If we detect a voicemail beep before this, we will speak the message at that point. - Setting too low a value means that the bot will start speaking its voicemail message too early. If it does so before the actual beep, it will get cut off. You should definitely tune this to your use case. @default 30 @min 0 @max 60
          - `provider` 'openai', required — This is the provider to use for voicemail detection.
          - `backoffPlan` VoicemailDetectionBackoffPlan
            - `startAtSeconds` number — This is the number of seconds to wait before starting the first retry attempt.
            - `frequencySeconds` number — This is the interval in seconds between retry attempts.
            - `maxRetries` number — This is the maximum number of retry attempts before giving up.
          - `type` 'audio' | 'transcript' — This is the detection type to use for voicemail detection. - 'audio': Uses native audio models (default) - 'transcript': Uses ASR/transcript-based detection @default 'audio' (audio detection)
        - TwilioVoicemailDetectionPlan
          - `provider` 'twilio', required — This is the provider to use for voicemail detection.
          - `voicemailDetectionTypes` string[] — These are the AMD messages from Twilio that are considered as voicemail. Default is ['machine_end_beep', 'machine_end_silence']. @default {Array} ['machine_end_beep', 'machine_end_silence']
          - `enabled` boolean — This sets whether the assistant should detect voicemail. Defaults to true. @default true
          - `machineDetectionTimeout` number — The number of seconds that Twilio should attempt to perform answering machine detection before timing out and returning AnsweredBy as unknown. Default is 30 seconds. Increasing this value will provide the engine more time to make a determination. This can be useful when DetectMessageEnd is provided in the MachineDetection parameter and there is an expectation of long answering machine greetings that can exceed 30 seconds. Decreasing this value will reduce the amount of time the engine has to make a determination. This can be particularly useful when the Enable option is provided in the MachineDetection parameter and you want to limit the time for initial detection. Check the [Twilio docs](https://www.twilio.com/docs/voice/answering-machine-detection#optional-api-tuning-parameters) for more info. @default 30
          - `machineDetectionSpeechThreshold` number — The number of milliseconds that is used as the measuring stick for the length of the speech activity. Durations lower than this value will be interpreted as a human, longer as a machine. Default is 2400 milliseconds. Increasing this value will reduce the chance of a False Machine (detected machine, actually human) for a long human greeting (e.g., a business greeting) but increase the time it takes to detect a machine. Decreasing this value will reduce the chances of a False Human (detected human, actually machine) for short voicemail greetings. The value of this parameter may need to be reduced by more than 1000ms to detect very short voicemail greetings. A reduction of that significance can result in increased False Machine detections. Adjusting the MachineDetectionSpeechEndThreshold is likely the better approach for short voicemails. Decreasing MachineDetectionSpeechThreshold will also reduce the time it takes to detect a machine. Check the [Twilio docs](https://www.twilio.com/docs/voice/answering-machine-detection#optional-api-tuning-parameters) for more info. @default 2400
          - `machineDetectionSpeechEndThreshold` number — The number of milliseconds of silence after speech activity at which point the speech activity is considered complete. Default is 1200 milliseconds. Increasing this value will typically be used to better address the short voicemail greeting scenarios. For short voicemails, there is typically 1000-2000ms of audio followed by 1200-2400ms of silence and then additional audio before the beep. Increasing the MachineDetectionSpeechEndThreshold to ~2500ms will treat the 1200-2400ms of silence as a gap in the greeting but not the end of the greeting and will result in a machine detection. The downsides of such a change include: - Increasing the delay for human detection by the amount you increase this parameter, e.g., a change of 1200ms to 2500ms increases human detection delay by 1300ms. - Cases where a human has two utterances separated by a period of silence (e.g. a "Hello", then 2000ms of silence, and another "Hello") may be interpreted as a machine. Decreasing this value will result in faster human detection. The consequence is that it can lead to increased False Human (detected human, actually machine) detections because a silence gap in a voicemail greeting (not necessarily just in short voicemail scenarios) can be incorrectly interpreted as the end of speech. Check the [Twilio docs](https://www.twilio.com/docs/voice/answering-machine-detection#optional-api-tuning-parameters) for more info. @default 1200
          - `machineDetectionSilenceTimeout` number — The number of milliseconds of initial silence after which an unknown AnsweredBy result will be returned. Default is 5000 milliseconds. Increasing this value will result in waiting for a longer period of initial silence before returning an 'unknown' AMD result. Decreasing this value will result in waiting for a shorter period of initial silence before returning an 'unknown' AMD result. Check the [Twilio docs](https://www.twilio.com/docs/voice/answering-machine-detection#optional-api-tuning-parameters) for more info. @default 5000
        - VapiVoicemailDetectionPlan
          - `beepMaxAwaitSeconds` number — This is the maximum duration from the start of the call that we will wait for a voicemail beep, before speaking our message - If we detect a voicemail beep before this, we will speak the message at that point. - Setting too low a value means that the bot will start speaking its voicemail message too early. If it does so before the actual beep, it will get cut off. You should definitely tune this to your use case. @default 30 @min 0 @max 60
          - `provider` 'vapi', required — This is the provider to use for voicemail detection.
          - `backoffPlan` VoicemailDetectionBackoffPlan
            - `startAtSeconds` number — This is the number of seconds to wait before starting the first retry attempt.
            - `frequencySeconds` number — This is the interval in seconds between retry attempts.
            - `maxRetries` number — This is the maximum number of retry attempts before giving up.
          - `type` 'audio' | 'transcript' — This is the detection type to use for voicemail detection. - 'audio': Uses native audio models (default) - 'transcript': Uses ASR/transcript-based detection @default 'audio' (audio detection)
      - `clientMessages` string[] — These are the messages that will be sent to your Client SDKs. Default is conversation-update,function-call,hang,model-output,speech-update,status-update,transfer-update,transcript,tool-calls,user-interrupted,voice-input,workflow.node.started,assistant.started. You can check the shape of the messages in ClientMessage schema.
      - `serverMessages` string[] — These are the messages that will be sent to your Server URL. Default is conversation-update,end-of-call-report,function-call,hang,speech-update,status-update,tool-calls,transfer-destination-request,handoff-destination-request,user-interrupted,assistant.started. You can check the shape of the messages in ServerMessage schema.
      - `maxDurationSeconds` number — This is the maximum number of seconds that the call will last. When the call reaches this duration, it will be ended. @default 600 (10 minutes)
      - `backgroundSound` union — This is the background sound in the call. Default for phone calls is 'office' and default for web calls is 'off'. You can also provide a custom sound by providing a URL to an audio file.
        - 'off' | 'office'
        - string, uri
      - `modelOutputInMessagesEnabled` boolean — This determines whether the model's output is used in conversation history rather than the transcription of assistant's speech. @default false
      - `transportConfigurations` TransportConfigurationTwilio[] — These are the configurations to be passed to the transport providers of assistant's calls, like Twilio. You can store multiple configurations for different transport providers. For a call, only the configuration matching the call transport provider is used.
        - `provider` 'twilio', required
        - `timeout` number — The integer number of seconds that we should allow the phone to ring before assuming there is no answer. The default is `60` seconds and the maximum is `600` seconds. For some call flows, we will add a 5-second buffer to the timeout value you provide. For this reason, a timeout value of 10 seconds could result in an actual timeout closer to 15 seconds. You can set this to a short time, such as `15` seconds, to hang up before reaching an answering machine or voicemail. @default 60
        - `record` boolean — Whether to record the call. Can be `true` to record the phone call, or `false` to not. The default is `false`. @default false
        - `recordingChannels` 'mono' | 'dual' — The number of channels in the final recording. Can be: `mono` or `dual`. The default is `mono`. `mono` records both legs of the call in a single channel of the recording file. `dual` records each leg to a separate channel of the recording file. The first channel of a dual-channel recording contains the parent call and the second channel contains the child call. @default 'mono'
      - `observabilityPlan` LangfuseObservabilityPlan
        - `provider` 'langfuse', required
        - `promptName` string — The name of a Langfuse prompt to link generations to. This enables tracking which prompt version was used for each generation. https://langfuse.com/docs/prompt-management/features/link-to-traces
        - `promptVersion` number — The version number of the Langfuse prompt to link generations to. Used together with promptName to identify the exact prompt version. https://langfuse.com/docs/prompt-management/features/link-to-traces
        - `traceName` string — Custom name for the Langfuse trace. Supports Liquid templates. Available variables: - {{ call.id }} - Call UUID - {{ call.type }} - 'inboundPhoneCall', 'outboundPhoneCall', 'webCall' - {{ assistant.name }} - Assistant name - {{ assistant.id }} - Assistant ID Example: "{{ assistant.name }} - {{ call.type }}" Defaults to call ID if not provided.
        - `tags` string[], required — This is an array of tags to be added to the Langfuse trace. Tags allow you to categorize and filter traces. https://langfuse.com/docs/tracing-features/tags
        - `metadata` object — This is a JSON object that will be added to the Langfuse trace. Traces can be enriched with metadata to better understand your users, application, and experiments. https://langfuse.com/docs/tracing-features/metadata By default it includes the call metadata, assistant metadata, and assistant overrides.
      - `credentials` union[] — These are dynamic credentials that will be used for the assistant calls. By default, all the credentials are available for use in the call but you can supplement an additional credentials using this. Dynamic credentials override existing credentials.
        - union
          - CreateAnthropicCredentialDTO
            - `provider` 'anthropic', required
            - `apiKey` string, required — This is not returned in the API.
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateAnthropicBedrockCredentialDTO
            - `provider` 'anthropic-bedrock', required
            - `region` 'us-east-1' | 'us-west-2' | 'eu-central-1' | 'eu-west-1' | 'eu-west-3' | 'ap-northeast-1' | 'ap-southeast-2', required — AWS region where Bedrock is configured.
            - `authenticationPlan` union, required — Authentication method - either direct IAM credentials or cross-account role assumption.
              - …
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateAnyscaleCredentialDTO
            - `provider` 'anyscale', required
            - `apiKey` string, required — This is not returned in the API.
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateAssemblyAICredentialDTO
            - `provider` 'assembly-ai', required
            - `apiKey` string, required — This is not returned in the API.
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateAzureCredentialDTO
            - `provider` 'azure', required
            - `service` 'speech' | 'blob_storage', required — This is the service being used in Azure.
            - `region` 'australiaeast' | 'canadaeast' | 'canadacentral' | 'centralus' | 'eastus2' | 'eastus' | 'france' | 'germanywestcentral' | 'india' | 'japaneast' | 'japanwest' | 'northcentralus' | 'norway' | 'polandcentral' | 'southcentralus' | 'spaincentral' | 'swedencentral' | 'switzerland' | 'switzerlandnorth' | 'switzerlandwest' | 'uaenorth' | 'uk' | 'westeurope' | 'westus' | 'westus3' — This is the region of the Azure resource.
            - `apiKey` string — This is not returned in the API.
            - `fallbackIndex` number — This is the order in which this storage provider is tried during upload retries. Lower numbers are tried first in increasing order.
            - `bucketPlan` AzureBlobStorageBucketPlan
              - …
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateAzureOpenAICredentialDTO
            - `provider` 'azure-openai', required
            - `region` 'australiaeast' | 'canadaeast' | 'canadacentral' | 'centralus' | 'eastus2' | 'eastus' | 'france' | 'germanywestcentral' | 'india' | 'japaneast' | 'japanwest' | 'northcentralus' | 'norway' | 'polandcentral' | 'southcentralus' | 'spaincentral' | 'swedencentral' | 'switzerland' | 'switzerlandnorth' | 'switzerlandwest' | 'uaenorth' | 'uk' | 'westeurope' | 'westus' | 'westus3', required
            - `models` string[], required
            - `openAIKey` string, required — This is not returned in the API.
            - `ocpApimSubscriptionKey` string — This is not returned in the API.
            - `openAIEndpoint` string, required
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateByoSipTrunkCredentialDTO
            - `provider` 'byo-sip-trunk' — This can be used to bring your own SIP trunks or to connect to a Carrier.
            - `gateways` SipTrunkGateway[], required — This is the list of SIP trunk's gateways.
              - …
            - `outboundAuthenticationPlan` SipTrunkOutboundAuthenticationPlan
              - …
            - `outboundLeadingPlusEnabled` boolean — This ensures the outbound origination attempts have a leading plus. Defaults to false to match conventional telecom behavior. Usage: - Vonage/Twilio requires leading plus for all outbound calls. Set this to true. @default false
            - `techPrefix` string — This can be used to configure the tech prefix on outbound calls. This is an advanced property.
            - `sipDiversionHeader` string — This can be used to enable the SIP diversion header for authenticating the calling number if the SIP trunk supports it. This is an advanced property.
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateCartesiaCredentialDTO
            - `provider` 'cartesia', required
            - `apiKey` string, required — This is not returned in the API.
            - `apiUrl` string — This can be used to point to an onprem Cartesia instance. Defaults to api.cartesia.ai.
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateCerebrasCredentialDTO
            - `provider` 'cerebras', required
            - `apiKey` string, required — This is not returned in the API.
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateCloudflareCredentialDTO
            - `provider` 'cloudflare', required — Credential provider. Only allowed value is cloudflare
            - `accountId` string — Cloudflare Account Id.
            - `apiKey` string — Cloudflare API Key / Token.
            - `accountEmail` string — Cloudflare Account Email.
            - `fallbackIndex` number — This is the order in which this storage provider is tried during upload retries. Lower numbers are tried first in increasing order.
            - `bucketPlan` CloudflareR2BucketPlan
              - …
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateCustomLLMCredentialDTO
            - `provider` 'custom-llm', required
            - `apiKey` string, required — This is not returned in the API.
            - `authenticationPlan` OAuth2AuthenticationPlan
              - …
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateDeepgramCredentialDTO
            - `provider` 'deepgram', required
            - `apiKey` string, required — This is not returned in the API.
            - `apiUrl` string — This can be used to point to an onprem Deepgram instance. Defaults to api.deepgram.com.
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateDeepInfraCredentialDTO
            - `provider` 'deepinfra', required
            - `apiKey` string, required — This is not returned in the API.
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateDeepSeekCredentialDTO
            - `provider` 'deep-seek', required
            - `apiKey` string, required — This is not returned in the API.
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateElevenLabsCredentialDTO
            - `provider` '11labs', required
            - `apiKey` string, required — This is not returned in the API.
            - `apiUrl` 'https://api.elevenlabs.io' | 'https://api.eu.residency.elevenlabs.io', nullable — ElevenLabs-only API environment for this key: the global endpoint or the EU data residency endpoint. In EU deployments, new credentials must explicitly use the EU data residency endpoint; existing credentials may omit this field on update to retain their saved endpoint. Outside EU deployments, Vapi detects an omitted endpoint automatically and null on update clears and re-detects the endpoint.
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateGcpCredentialDTO
            - `provider` 'gcp', required
            - `fallbackIndex` number — This is the order in which this storage provider is tried during upload retries. Lower numbers are tried first in increasing order.
            - `gcpKey` GcpKey, required
              - …
            - `region` string — This is the region of the GCP resource.
            - `bucketPlan` BucketPlan
              - …
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateGladiaCredentialDTO
            - `provider` 'gladia', required
            - `apiKey` string, required — This is not returned in the API.
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateGoHighLevelCredentialDTO
            - `provider` 'gohighlevel', required
            - `apiKey` string, required — This is not returned in the API.
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateGoogleCredentialDTO
            - `provider` 'google', required — This is the key for Gemini in Google AI Studio. Get it from here: https://aistudio.google.com/app/apikey
            - `apiKey` string, required — This is not returned in the API.
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateGroqCredentialDTO
            - `provider` 'groq', required
            - `apiKey` string, required — This is not returned in the API.
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateHumeCredentialDTO
            - `provider` 'hume', required
            - `apiKey` string, required — This is not returned in the API.
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateInflectionAICredentialDTO
            - `provider` 'inflection-ai', required — This is the api key for Pi in InflectionAI's console. Get it from here: https://developers.inflection.ai/keys, billing will need to be setup
            - `apiKey` string, required — This is not returned in the API.
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateLangfuseCredentialDTO
            - `provider` 'langfuse', required
            - `publicKey` string, required — The public key for Langfuse project. Eg: pk-lf-...
            - `apiKey` string, required — The secret key for Langfuse project. Eg: sk-lf-... .This is not returned in the API.
            - `apiUrl` string, required — The host URL for Langfuse project. Eg: https://cloud.langfuse.com
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateLmntCredentialDTO
            - `provider` 'lmnt', required
            - `apiKey` string, required — This is not returned in the API.
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateMakeCredentialDTO
            - `provider` 'make', required
            - `teamId` string, required — Team ID
            - `region` string, required — Region of your application. For example: eu1, eu2, us1, us2
            - `apiKey` string, required — This is not returned in the API.
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateMistralCredentialDTO
            - `provider` 'mistral', required
            - `apiKey` string, required — This is not returned in the API.
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateNeuphonicCredentialDTO
            - `provider` 'neuphonic', required
            - `apiKey` string, required — This is not returned in the API.
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateOpenAICredentialDTO
            - `provider` 'openai', required
            - `apiKey` string, required — This is not returned in the API.
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateOpenRouterCredentialDTO
            - `provider` 'openrouter', required
            - `apiKey` string, required — This is not returned in the API.
            - `name` string — This is the name of credential. This is just for your reference.
          - CreatePerplexityAICredentialDTO
            - `provider` 'perplexity-ai', required
            - `apiKey` string, required — This is not returned in the API.
            - `name` string — This is the name of credential. This is just for your reference.
          - CreatePlayHTCredentialDTO
            - `provider` 'playht', required
            - `apiKey` string, required — This is not returned in the API.
            - `userId` string, required
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateRimeAICredentialDTO
            - `provider` 'rime-ai', required
            - `apiKey` string, required — This is not returned in the API.
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateRunpodCredentialDTO
            - `provider` 'runpod', required
            - `apiKey` string, required — This is not returned in the API.
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateS3CredentialDTO
            - `provider` 's3', required — Credential provider. Only allowed value is s3
            - `awsAccessKeyId` string, required — AWS access key ID.
            - `awsSecretAccessKey` string, required — AWS access key secret. This is not returned in the API.
            - `region` string, required — AWS region in which the S3 bucket is located.
            - `s3BucketName` string, required — AWS S3 bucket name.
            - `s3PathPrefix` string, required — The path prefix for the uploaded recording. Ex. "recordings/"
            - `fallbackIndex` number — This is the order in which this storage provider is tried during upload retries. Lower numbers are tried first in increasing order.
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateS3CompatibleCredentialDTO
            - `provider` 's3-compatible', required — This is for S3-compatible storage such as MinIO, Garage, Ceph, or Backblaze B2.
            - `bucketPlan` S3CompatibleBucketPlan, required
              - …
            - `fallbackIndex` number — This is the order in which this storage provider is tried during upload retries. Lower numbers are tried first in increasing order.
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateSmallestAICredentialDTO
            - `provider` 'smallest-ai', required
            - `apiKey` string, required — This is not returned in the API.
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateSpeechmaticsCredentialDTO
            - `provider` 'speechmatics', required
            - `apiKey` string, required — This is not returned in the API.
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateSonioxCredentialDTO
            - `provider` 'soniox', required
            - `apiKey` string, required — This is not returned in the API.
            - `apiUrl` string — Custom Soniox WebSocket endpoint (e.g. EU server wss://stt-rt.eu.soniox.com/transcribe-websocket). Defaults to the region-appropriate endpoint when omitted.
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateSupabaseCredentialDTO
            - `provider` 'supabase', required — This is for supabase storage.
            - `fallbackIndex` number — This is the order in which this storage provider is tried during upload retries. Lower numbers are tried first in increasing order.
            - `bucketPlan` SupabaseBucketPlan
              - …
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateTavusCredentialDTO
            - `provider` 'tavus', required
            - `apiKey` string, required — This is not returned in the API.
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateTogetherAICredentialDTO
            - `provider` 'together-ai', required
            - `apiKey` string, required — This is not returned in the API.
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateTwilioCredentialDTO
            - `provider` 'twilio', required
            - `authToken` string — This is not returned in the API.
            - `apiKey` string — This is not returned in the API.
            - `apiSecret` string — This is not returned in the API.
            - `accountSid` string, required
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateVonageCredentialDTO
            - `provider` 'vonage', required
            - `apiSecret` string, required — This is not returned in the API.
            - `apiKey` string, required
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateWebhookCredentialDTO
            - `provider` 'webhook', required
            - `authenticationPlan` union, required — This is the authentication plan. Supports OAuth2 RFC 6749, HMAC signing, and Bearer authentication.
              - …
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateCustomCredentialDTO
            - `provider` 'custom-credential', required
            - `authenticationPlan` union, required — This is the authentication plan. Supports OAuth2 RFC 6749, HMAC signing, and Bearer authentication.
              - …
            - `encryptionPlan` PublicKeyEncryptionPlan
              - …
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateXAiCredentialDTO
            - `provider` 'xai', required — This is the api key for Grok in XAi's console. Get it from here: https://console.x.ai
            - `apiKey` string, required — This is not returned in the API.
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateMicrosoftCredentialDTO
            - `provider` 'microsoft', required
            - `apiKey` string, required — This is not returned in the API.
            - `region` string — Azure region for the Speech resource. Defaults to `eastus` when omitted. MAI-Voice-2 is preview and region-limited.
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateGoogleCalendarOAuth2ClientCredentialDTO
            - `provider` 'google.calendar.oauth2-client', required
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateGoogleCalendarOAuth2AuthorizationCredentialDTO
            - `provider` 'google.calendar.oauth2-authorization', required
            - `authorizationId` string, required — The authorization ID for the OAuth2 authorization
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateGoogleSheetsOAuth2AuthorizationCredentialDTO
            - `provider` 'google.sheets.oauth2-authorization', required
            - `authorizationId` string, required — The authorization ID for the OAuth2 authorization
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateSlackOAuth2AuthorizationCredentialDTO
            - `provider` 'slack.oauth2-authorization', required
            - `authorizationId` string, required — The authorization ID for the OAuth2 authorization
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateGoHighLevelMCPCredentialDTO
            - `provider` 'ghl.oauth2-authorization', required
            - `authenticationSession` Oauth2AuthenticationSession, required
              - …
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateInworldCredentialDTO
            - `provider` 'inworld', required
            - `apiKey` string, required — This is the Inworld Basic (Base64) authentication token. This is not returned in the API.
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateMinimaxCredentialDTO
            - `provider` 'minimax', required
            - `apiKey` string, required — This is not returned in the API.
            - `groupId` string, required — This is the Minimax Group ID.
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateWellSaidCredentialDTO
            - `provider` 'wellsaid', required
            - `apiKey` string, required — This is not returned in the API.
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateEmailCredentialDTO
            - `provider` 'email', required
            - `email` string, required — The recipient email address for alerts
            - `name` string — This is the name of credential. This is just for your reference.
          - CreateSlackWebhookCredentialDTO
            - `provider` 'slack-webhook', required
            - `webhookUrl` string, required — Slack incoming webhook URL. See https://api.slack.com/messaging/webhooks for setup instructions. This is not returned in the API.
            - `name` string — This is the name of credential. This is just for your reference.
      - `hooks` union[] — This is a set of actions that will be performed on certain events.
        - union
          - CallHookCallEnding
            - `on` 'call.ending', required — This is the event that triggers this hook
            - `do` union[], required — This is the set of actions to perform when the hook triggers
              - …
            - `filters` CallHookFilter[] — This is the set of filters that must match for the hook to trigger
              - …
          - CallHookAssistantSpeechInterrupted
            - `on` 'assistant.speech.interrupted', required — This is the event that triggers this hook
            - `do` union[], required — This is the set of actions to perform when the hook triggers
              - …
          - CallHookCustomerSpeechInterrupted
            - `on` 'customer.speech.interrupted', required — This is the event that triggers this hook
            - `do` union[], required — This is the set of actions to perform when the hook triggers
              - …
          - CallHookCustomerSpeechTimeout
            - `on` string, required — Must be either "customer.speech.timeout" or match the pattern "customer.speech.timeout[property=value]"
            - `do` union[], required — This is the set of actions to perform when the hook triggers
              - …
            - `options` CustomerSpeechTimeoutOptions
              - …
            - `name` string — This is the name of the hook, it can be set by the user to identify the hook. If no name is provided, the hook will be auto generated as UUID. @default UUID
          - SessionCreatedHook
            - `on` 'session.created', required — This is the event that triggers this hook
            - `do` ToolCallHookAction[], required — This is the set of actions to perform when the hook triggers.
              - …
            - `name` string — Optional name for this hook instance. If no name is provided, the hook will be auto generated as UUID. @default UUID
      - `name` string — This is the name of the assistant. This is required when you want to transfer between assistants in a call.
      - `voicemailMessage` string — This is the message that the assistant will say if the call is forwarded to voicemail. If unspecified, it will hang up.
      - `endCallMessage` string — This is the message that the assistant will say if it ends the call. If unspecified, it will hang up without saying anything.
      - `endCallPhrases` string[] — This list contains phrases that, if spoken by the assistant, will trigger the call to be hung up. Case insensitive.
      - `compliancePlan` CompliancePlan
        - `hipaaEnabled` boolean — When this is enabled, logs, recordings, and transcriptions will be stored in HIPAA-compliant storage. Defaults to false. Only HIPAA-compliant providers will be available for LLM, Voice, and Transcriber respectively. This setting is only honored if the organization is on an Enterprise subscription or has purchased the HIPAA add-on.
        - `pciEnabled` boolean — When this is enabled, the user will be restricted to use PCI-compliant providers, and no logs or transcripts are stored. At the end of the call, you will receive an end-of-call-report message to store on your server. Defaults to false.
        - `securityFilterPlan` SecurityFilterPlan
          - `enabled` boolean — Whether the security filter is enabled. @default false
          - `filters` SecurityFilterBase[] — Array of security filter types to apply. If array is not empty, only those security filters are run.
          - `mode` 'sanitize' | 'reject' | 'replace' — Mode of operation when a security threat is detected. - 'sanitize': Remove or replace the threatening content - 'reject': Replace the entire transcript with replacement text - 'replace': Replace threatening patterns with replacement text @default 'sanitize'
          - `replacementText` string — Text to use when replacing filtered content. @default '[FILTERED]'
        - `recordingConsentPlan` union
          - RecordingConsentPlanStayOnLine
            - `message` string, required — This is the message asking for consent to record the call. If the type is `stay-on-line`, the message should ask the user to hang up if they do not consent. If the type is `verbal`, the message should ask the user to verbally consent or decline.
            - `voice` union — This is the voice to use for the consent message. If not specified, inherits from the assistant's voice. Use a different voice for the consent message for a better user experience.
              - …
            - `firstMessageMode` 'assistant-speaks-first' | 'assistant-waits-for-user' — This controls whether the consent assistant speaks first or waits for the caller to speak first. Use: - `assistant-speaks-first` (default) to have the consent assistant play the consent message as soon as the call is answered. - `assistant-waits-for-user` to have the consent assistant wait for the caller to speak before playing the consent message. We strongly recommend `assistant-waits-for-user` for outbound calls. Some telephony providers signal "answered" while the line is still ringing, which can cause the consent message to play into a ringing line and be missed by the caller. Waiting for the caller to speak first guarantees they hear the full consent message. Note: when combined with `type: 'stay-on-line'`, silence only counts toward consent after the caller has spoken at least once. @default 'assistant-speaks-first'
            - `type` 'stay-on-line', required — This is the type of recording consent plan. This type assumes consent is granted if the user stays on the line.
            - `waitSeconds` number — Number of seconds to wait before transferring to the assistant if user stays on the call
          - RecordingConsentPlanVerbal
            - `message` string, required — This is the message asking for consent to record the call. If the type is `stay-on-line`, the message should ask the user to hang up if they do not consent. If the type is `verbal`, the message should ask the user to verbally consent or decline.
            - `voice` union — This is the voice to use for the consent message. If not specified, inherits from the assistant's voice. Use a different voice for the consent message for a better user experience.
              - …
            - `firstMessageMode` 'assistant-speaks-first' | 'assistant-waits-for-user' — This controls whether the consent assistant speaks first or waits for the caller to speak first. Use: - `assistant-speaks-first` (default) to have the consent assistant play the consent message as soon as the call is answered. - `assistant-waits-for-user` to have the consent assistant wait for the caller to speak before playing the consent message. We strongly recommend `assistant-waits-for-user` for outbound calls. Some telephony providers signal "answered" while the line is still ringing, which can cause the consent message to play into a ringing line and be missed by the caller. Waiting for the caller to speak first guarantees they hear the full consent message. Note: when combined with `type: 'stay-on-line'`, silence only counts toward consent after the caller has spoken at least once. @default 'assistant-speaks-first'
            - `type` 'verbal', required — This is the type of recording consent plan. This type assumes consent is granted if the user verbally consents or declines.
            - `declineTool` object — Tool to execute if user verbally declines recording consent
            - `declineToolId` string — ID of existing tool to execute if user verbally declines recording consent
      - `metadata` object — This is for metadata you want to store on the assistant.
      - `backgroundSpeechDenoisingPlan` BackgroundSpeechDenoisingPlan
        - `smartDenoisingPlan` SmartDenoisingPlan
          - `enabled` boolean — Whether smart denoising using Krisp is enabled.
        - `fourierDenoisingPlan` FourierDenoisingPlan
          - `enabled` boolean — Whether Fourier denoising is enabled. Note that this is experimental and may not work as expected.
          - `mediaDetectionEnabled` boolean — Whether automatic media detection is enabled. When enabled, the filter will automatically detect consistent background TV/music/radio and switch to more aggressive filtering settings. Only applies when enabled is true.
          - `staticThreshold` number — Static threshold in dB used as fallback when no baseline is established.
          - `baselineOffsetDb` number — How far below the rolling baseline to filter audio, in dB. Lower values (e.g., -10) are more aggressive, higher values (e.g., -20) are more conservative.
          - `windowSizeMs` number — Rolling window size in milliseconds for calculating the audio baseline. Larger windows adapt more slowly but are more stable.
          - `baselinePercentile` number — Percentile to use for baseline calculation (1-99). Higher percentiles (e.g., 85) focus on louder speech, lower percentiles (e.g., 50) include quieter speech.
      - `analysisPlan` AnalysisPlan
        - `minMessagesThreshold` number — The minimum number of messages required to run the analysis plan. If the number of messages is less than this, analysis will be skipped. @default 2
        - `summaryPlan` SummaryPlan
          - `messages` object[] — These are the messages used to generate the summary. @default: ``` [ { "role": "system", "content": "You are an expert note-taker. You will be given a transcript of a call. Summarize the call in 2-3 sentences. DO NOT return anything except the summary." }, { "role": "user", "content": "Here is the transcript:\n\n{{transcript}}\n\n. Here is the ended reason of the call:\n\n{{endedReason}}\n\n" } ]``` You can customize by providing any messages you want. Here are the template variables available: - {{transcript}}: The transcript of the call from `call.artifact.transcript` - {{systemPrompt}}: The system prompt of the call from `assistant.model.messages[type=system].content` - {{messages}}: The messages of the call from `assistant.model.messages` - {{endedReason}}: The ended reason of the call from `call.endedReason`
          - `enabled` boolean — This determines whether a summary is generated and stored in `call.analysis.summary`. Defaults to true. Usage: - If you want to disable the summary, set this to false. @default true
          - `timeoutSeconds` number — This is how long the request is tried before giving up. When request times out, `call.analysis.summary` will be empty. Usage: - To guarantee the summary is generated, set this value high. Note, this will delay the end of call report in cases where model is slow to respond. @default 5 seconds
        - `structuredDataPlan` StructuredDataPlan
          - `messages` object[] — These are the messages used to generate the structured data. @default: ``` [ { "role": "system", "content": "You are an expert data extractor. You will be given a transcript of a call. Extract structured data per the JSON Schema. DO NOT return anything except the structured data.\n\nJson Schema:\\n{{schema}}\n\nOnly respond with the JSON." }, { "role": "user", "content": "Here is the transcript:\n\n{{transcript}}\n\n. Here is the ended reason of the call:\n\n{{endedReason}}\n\n" } ]``` You can customize by providing any messages you want. Here are the template variables available: - {{transcript}}: the transcript of the call from `call.artifact.transcript`- {{systemPrompt}}: the system prompt of the call from `assistant.model.messages[type=system].content`- {{messages}}: the messages of the call from `assistant.model.messages`- {{schema}}: the schema of the structured data from `structuredDataPlan.schema`- {{endedReason}}: the ended reason of the call from `call.endedReason`
          - `enabled` boolean — This determines whether structured data is generated and stored in `call.analysis.structuredData`. Defaults to false. Usage: - If you want to extract structured data, set this to true and provide a `schema`. @default false
          - `schema` JsonSchema
            - `type` 'string' | 'number' | 'integer' | 'boolean' | 'array' | 'object', required — This is the type of output you'd like. `string`, `number`, `integer`, `boolean` are the primitive types and should be obvious. `array` and `object` are more interesting and quite powerful. They allow you to define nested structures. For `array`, you can define the schema of the items in the array using the `items` property. For `object`, you can define the properties of the object using the `properties` property.
            - `items` JsonSchema — recursive
            - `properties` object — This is required if the type is "object". This specifies the properties of the object. This is a map of property names to JsonSchema objects.
            - `description` string — This is the description to help the model understand what it needs to output.
            - `pattern` string — This is the pattern of the string. This is a regex that will be used to validate the data in question. To use a common format, use the `format` property instead. OpenAI documentation: https://platform.openai.com/docs/guides/structured-outputs#supported-properties
            - `format` 'date-time' | 'time' | 'date' | 'duration' | 'email' | 'hostname' | 'ipv4' | 'ipv6' | 'uuid' — This is the format of the string. To pass a regex, use the `pattern` property instead. OpenAI documentation: https://platform.openai.com/docs/guides/structured-outputs?api-mode=chat&type-restrictions=string-restrictions
            - `required` string[] — This is a list of properties that are required. This only makes sense if the type is "object".
            - `enum` string[] — This array specifies the allowed values that can be used to restrict the output of the model.
            - `title` string — This is the title of the schema.
          - `timeoutSeconds` number — This is how long the request is tried before giving up. When request times out, `call.analysis.structuredData` will be empty. Usage: - To guarantee the structured data is generated, set this value high. Note, this will delay the end of call report in cases where model is slow to respond. @default 5 seconds
        - `structuredDataMultiPlan` StructuredDataMultiPlan[] — This is an array of structured data plan catalogs. Each entry includes a `key` and a `plan` for generating the structured data from the call. This outputs to `call.analysis.structuredDataMulti`.
          - `key` string, required — This is the key of the structured data plan in the catalog.
          - `plan` StructuredDataPlan, required
            - `messages` object[] — These are the messages used to generate the structured data. @default: ``` [ { "role": "system", "content": "You are an expert data extractor. You will be given a transcript of a call. Extract structured data per the JSON Schema. DO NOT return anything except the structured data.\n\nJson Schema:\\n{{schema}}\n\nOnly respond with the JSON." }, { "role": "user", "content": "Here is the transcript:\n\n{{transcript}}\n\n. Here is the ended reason of the call:\n\n{{endedReason}}\n\n" } ]``` You can customize by providing any messages you want. Here are the template variables available: - {{transcript}}: the transcript of the call from `call.artifact.transcript`- {{systemPrompt}}: the system prompt of the call from `assistant.model.messages[type=system].content`- {{messages}}: the messages of the call from `assistant.model.messages`- {{schema}}: the schema of the structured data from `structuredDataPlan.schema`- {{endedReason}}: the ended reason of the call from `call.endedReason`
              - …
            - `enabled` boolean — This determines whether structured data is generated and stored in `call.analysis.structuredData`. Defaults to false. Usage: - If you want to extract structured data, set this to true and provide a `schema`. @default false
            - `schema` JsonSchema
              - …
            - `timeoutSeconds` number — This is how long the request is tried before giving up. When request times out, `call.analysis.structuredData` will be empty. Usage: - To guarantee the structured data is generated, set this value high. Note, this will delay the end of call report in cases where model is slow to respond. @default 5 seconds
        - `successEvaluationPlan` SuccessEvaluationPlan
          - `rubric` 'NumericScale' | 'DescriptiveScale' | 'Checklist' | 'Matrix' | 'PercentageScale' | 'LikertScale' | 'AutomaticRubric' | 'PassFail' — This enforces the rubric of the evaluation. The output is stored in `call.analysis.successEvaluation`. Options include: - 'NumericScale': A scale of 1 to 10. - 'DescriptiveScale': A scale of Excellent, Good, Fair, Poor. - 'Checklist': A checklist of criteria and their status. - 'Matrix': A grid that evaluates multiple criteria across different performance levels. - 'PercentageScale': A scale of 0% to 100%. - 'LikertScale': A scale of Strongly Agree, Agree, Neutral, Disagree, Strongly Disagree. - 'AutomaticRubric': Automatically break down evaluation into several criteria, each with its own score. - 'PassFail': A simple 'true' if call passed, 'false' if not. Default is 'PassFail'.
          - `messages` object[] — These are the messages used to generate the success evaluation. @default: ``` [ { "role": "system", "content": "You are an expert call evaluator. You will be given a transcript of a call and the system prompt of the AI participant. Determine if the call was successful based on the objectives inferred from the system prompt. DO NOT return anything except the result.\n\nRubric:\\n{{rubric}}\n\nOnly respond with the result." }, { "role": "user", "content": "Here is the transcript:\n\n{{transcript}}\n\n" }, { "role": "user", "content": "Here was the system prompt of the call:\n\n{{systemPrompt}}\n\n. Here is the ended reason of the call:\n\n{{endedReason}}\n\n" } ]``` You can customize by providing any messages you want. Here are the template variables available: - {{transcript}}: the transcript of the call from `call.artifact.transcript`- {{systemPrompt}}: the system prompt of the call from `assistant.model.messages[type=system].content`- {{messages}}: the messages of the call from `assistant.model.messages`- {{rubric}}: the rubric of the success evaluation from `successEvaluationPlan.rubric`- {{endedReason}}: the ended reason of the call from `call.endedReason`
          - `enabled` boolean — This determines whether a success evaluation is generated and stored in `call.analysis.successEvaluation`. Defaults to true. Usage: - If you want to disable the success evaluation, set this to false. @default true
          - `timeoutSeconds` number — This is how long the request is tried before giving up. When request times out, `call.analysis.successEvaluation` will be empty. Usage: - To guarantee the success evaluation is generated, set this value high. Note, this will delay the end of call report in cases where model is slow to respond. @default 5 seconds
        - `outcomeIds` string[] — This is an array of outcome UUIDs to be calculated during analysis. The outcomes will be calculated and stored in `call.analysis.outcomes`.
      - `artifactPlan` ArtifactPlan
        - `recordingEnabled` boolean — This determines whether assistant's calls are recorded. Defaults to true. Usage: - If you don't want to record the calls, set this to false. - If you want to record the calls when `assistant.hipaaEnabled` (deprecated) or `assistant.compliancePlan.hipaaEnabled` explicity set this to true and make sure to provide S3 or GCP credentials on the Provider Credentials page in the Dashboard. You can find the recording at `call.artifact.recordingUrl` and `call.artifact.stereoRecordingUrl` after the call is ended. @default true
        - `recordingFormat` 'wav;l16' | 'mp3' — This determines the format of the recording. Defaults to `wav;l16`. @default 'wav;l16'
        - `recordingUseCustomStorageEnabled` boolean — This determines whether to use custom storage (S3 or GCP) for call recordings when storage credentials are configured. When set to false, recordings will be stored on Vapi's storage instead of your custom storage, even if you have custom storage credentials configured. Usage: - Set to false if you have custom storage configured but want to store recordings on Vapi's storage for this assistant. - Set to true (or leave unset) to use your custom storage for recordings when available. If your organization has ZDR (zero data retention) or PCI enabled, recordings are never written to Vapi storage. In that case, false means "do not use my custom storage", so nothing is stored at all. @default true
        - `videoRecordingEnabled` boolean — This determines whether the video is recorded during the call. Defaults to false. Only relevant for `webCall` type. You can find the video recording at `call.artifact.videoRecordingUrl` after the call is ended. @default false
        - `fullMessageHistoryEnabled` boolean — This determines whether the artifact contains the full message history, even after handoff context engineering. Defaults to false.
        - `pcapEnabled` boolean — This determines whether the SIP packet capture is enabled. Defaults to true. Only relevant for `phone` type calls where phone number's provider is `vapi` or `byo-phone-number`. You can find the packet capture at `call.artifact.pcapUrl` after the call is ended. @default true
        - `pcapS3PathPrefix` string — This is the path where the SIP packet capture will be uploaded. This is only used if you have provided S3 or GCP credentials on the Provider Credentials page in the Dashboard. If credential.s3PathPrefix or credential.bucketPlan.path is set, this will append to it. Usage: - If you want to upload the packet capture to a specific path, set this to the path. Example: `/my-assistant-captures`. - If you want to upload the packet capture to the root of the bucket, set this to `/`. @default '/'
        - `pcapUseCustomStorageEnabled` boolean — This determines whether to use custom storage (S3 or GCP) for SIP packet captures when storage credentials are configured. When set to false, packet captures will be stored on Vapi's storage instead of your custom storage, even if you have custom storage credentials configured. Usage: - Set to false if you have custom storage configured but want to store packet captures on Vapi's storage for this assistant. - Set to true (or leave unset) to use your custom storage for packet captures when available. If your organization has ZDR (zero data retention) or PCI enabled, packet captures are never written to Vapi storage. In that case, false means "do not use my custom storage", so nothing is stored at all. @default true
        - `loggingEnabled` boolean — This determines whether the call logs are enabled. Defaults to true. @default true
        - `loggingUseCustomStorageEnabled` boolean — This determines whether to use custom storage (S3 or GCP) for call logs when storage credentials are configured. When set to false, logs will be stored on Vapi's storage instead of your custom storage, even if you have custom storage credentials configured. Usage: - Set to false if you have custom storage configured but want to store logs on Vapi's storage for this assistant. - Set to true (or leave unset) to use your custom storage for logs when available. If your organization has ZDR (zero data retention) or PCI enabled, logs are never written to Vapi storage. In that case, false means "do not use my custom storage", so nothing is stored at all. @default true
        - `transcriptPlan` TranscriptPlan
          - `enabled` boolean — This determines whether the transcript is stored in `call.artifact.transcript`. Defaults to true. @default true
          - `assistantName` string — This is the name of the assistant in the transcript. Defaults to 'AI'. Usage: - If you want to change the name of the assistant in the transcript, set this. Example, here is what the transcript would look like with `assistantName` set to 'Buyer': ``` User: Hello, how are you? Buyer: I'm fine. User: Do you want to buy a car? Buyer: No. ``` @default 'AI'
          - `userName` string — This is the name of the user in the transcript. Defaults to 'User'. Usage: - If you want to change the name of the user in the transcript, set this. Example, here is what the transcript would look like with `userName` set to 'Seller': ``` Seller: Hello, how are you? AI: I'm fine. Seller: Do you want to buy a car? AI: No. ``` @default 'User'
        - `recordingPath` string — This is the path where the recording will be uploaded. This is only used if you have provided S3 or GCP credentials on the Provider Credentials page in the Dashboard. If credential.s3PathPrefix or credential.bucketPlan.path is set, this will append to it. Usage: - If you want to upload the recording to a specific path, set this to the path. Example: `/my-assistant-recordings`. - If you want to upload the recording to the root of the bucket, set this to `/`. @default '/'
        - `structuredOutputIds` string[] — This is an array of structured output IDs to be calculated during the call. The outputs will be extracted and stored in `call.artifact.structuredOutputs` after the call is ended.
        - `structuredOutputs` CreateStructuredOutputDTO[] — This is an array of transient structured outputs to be calculated during the call. The outputs will be extracted and stored in `call.artifact.structuredOutputs` after the call is ended. Use this to provide inline structured output configurations instead of referencing existing ones via structuredOutputIds.
          - `type` 'ai' | 'regex' — This is the type of structured output. - 'ai': Uses an LLM to extract structured data from the conversation (default). - 'regex': Uses a regex pattern to extract data from the transcript without an LLM. Defaults to 'ai' if not specified.
          - `regex` string — This is the regex pattern to match against the transcript. Only used when type is 'regex'. Supports both raw patterns (e.g. '\d+') and regex literal format (e.g. '/\d+/gi'). Uses RE2 syntax for safety. The result depends on the schema type: - boolean: true if the pattern matches, false otherwise - string: the first match or first capture group - number/integer: the first match parsed as a number - array: all matches
          - `model` union — This is the model that will be used to extract the structured output. To provide your own custom system and user prompts for structured output extraction, populate the messages array with your system and user messages. You can specify liquid templating in your system and user messages. Between the system or user messages, you must reference either 'transcript' or 'messages' with the `{{}}` syntax to access the conversation history. Between the system or user messages, you must reference a variation of the structured output with the `{{}}` syntax to access the structured output definition. i.e.: `{{structuredOutput}}` `{{structuredOutput.name}}` `{{structuredOutput.description}}` `{{structuredOutput.schema}}` If model is not specified, GPT-4.1 will be used by default for extraction, utilizing default system and user prompts. If messages or required fields are not specified, the default system and user prompts will be used.
            - WorkflowOpenAIModel
              - …
            - WorkflowAnthropicModel
              - …
            - WorkflowAnthropicBedrockModel
              - …
            - WorkflowGoogleModel
              - …
            - WorkflowCustomModel
              - …
          - `compliancePlan` ComplianceOverride
            - `forceStoreOnHipaaEnabled` boolean — Force storage for this output under HIPAA. Only enable if output contains no sensitive data.
          - `conditions` union[], nullable — These are the conditions that gate the execution of this structured output. Every condition must pass for the structured output to run (AND semantics). When omitted or empty, no user-defined conditions gate this output. Send null to clear a previously saved gate.
            - union
              - …
          - `name` string, required — This is the name of the structured output.
          - `schema` JsonSchema, required
            - `type` 'string' | 'number' | 'integer' | 'boolean' | 'array' | 'object', required — This is the type of output you'd like. `string`, `number`, `integer`, `boolean` are the primitive types and should be obvious. `array` and `object` are more interesting and quite powerful. They allow you to define nested structures. For `array`, you can define the schema of the items in the array using the `items` property. For `object`, you can define the properties of the object using the `properties` property.
            - `items` JsonSchema — recursive
            - `properties` object — This is required if the type is "object". This specifies the properties of the object. This is a map of property names to JsonSchema objects.
            - `description` string — This is the description to help the model understand what it needs to output.
            - `pattern` string — This is the pattern of the string. This is a regex that will be used to validate the data in question. To use a common format, use the `format` property instead. OpenAI documentation: https://platform.openai.com/docs/guides/structured-outputs#supported-properties
            - `format` 'date-time' | 'time' | 'date' | 'duration' | 'email' | 'hostname' | 'ipv4' | 'ipv6' | 'uuid' — This is the format of the string. To pass a regex, use the `pattern` property instead. OpenAI documentation: https://platform.openai.com/docs/guides/structured-outputs?api-mode=chat&type-restrictions=string-restrictions
            - `required` string[] — This is a list of properties that are required. This only makes sense if the type is "object".
            - `enum` string[] — This array specifies the allowed values that can be used to restrict the output of the model.
            - `title` string — This is the title of the schema.
          - `description` string — This is the description of what the structured output extracts. Use this to provide context about what data will be extracted and how it will be used.
          - `assistantIds` string[] — These are the assistant IDs that this structured output is linked to. When linked to assistants, this structured output will be available for extraction during those assistant's calls.
          - `workflowIds` string[] — These are the workflow IDs that this structured output is linked to. When linked to workflows, this structured output will be available for extraction during those workflow's execution.
        - `scorecardIds` string[] — This is an array of scorecard IDs that will be evaluated based on the structured outputs extracted during the call. The scorecards will be evaluated and the results will be stored in `call.artifact.scorecards` after the call has ended.
        - `scorecards` CreateScorecardDTO[] — This is the array of scorecards that will be evaluated based on the structured outputs extracted during the call. The scorecards will be evaluated and the results will be stored in `call.artifact.scorecards` after the call has ended.
          - `name` string — This is the name of the scorecard. It is only for user reference and will not be used for any evaluation.
          - `description` string — This is the description of the scorecard. It is only for user reference and will not be used for any evaluation.
          - `metrics` ScorecardMetric[], required — These are the metrics that will be used to evaluate the scorecard. Each metric will have a set of conditions and points that will be used to generate the score.
            - `conditions` union[], required — These are the conditions that will be used to evaluate the scorecard. Each condition will have a comparator, value, and points that will be used to calculate the final score. The points will be added to the overall score if the condition is met. The overall score will be normalized to a 100 point scale to ensure uniformity across different scorecards.
              - …
            - `structuredOutputId` string, required — This is the unique identifier for the structured output that will be used to evaluate the scorecard. The structured output must be of type number or boolean only for now.
          - `assistantIds` string[] — These are the assistant IDs that this scorecard is linked to. When linked to assistants, this scorecard will be available for evaluation during those assistants' calls.
        - `loggingPath` string — This is the path where the call logs will be uploaded. This is only used if you have provided S3 or GCP credentials on the Provider Credentials page in the Dashboard. If credential.s3PathPrefix or credential.bucketPlan.path is set, this will append to it. Usage: - If you want to upload the call logs to a specific path, set this to the path. Example: `/my-assistant-logs`. - If you want to upload the call logs to the root of the bucket, set this to `/`. @default '/'
      - `startSpeakingPlan` StartSpeakingPlan
        - `waitSeconds` number — This is how long assistant waits before speaking. Defaults to 0.4. This is the minimum it will wait but if there is latency is the pipeline, this minimum will be exceeded. This is intended as a stopgap in case the pipeline is moving too fast. Example: - If model generates tokens and voice generates bytes within 100ms, the pipeline still waits 300ms before outputting speech. Usage: - If the customer is taking long pauses, set this to a higher value. - If the assistant is accidentally jumping in too much, set this to a higher value. @default 0.4
        - `smartEndpointingEnabled` union
          - boolean
          - 'livekit'
        - `smartEndpointingPlan` union — This is the plan for smart endpointing. Pick between Vapi smart endpointing, LiveKit, or custom endpointing model (or nothing). We strongly recommend using livekit endpointing when working in English. LiveKit endpointing is not supported in other languages, yet. If this is set, it will override and take precedence over `transcriptionEndpointingPlan`. This plan will still be overridden by any matching `customEndpointingRules`. If this is not set, the system will automatically use the transcriber's built-in endpointing capabilities if available.
          - VapiSmartEndpointingPlan
            - `provider` 'vapi' | 'livekit' | 'custom-endpointing-model', required — This is the provider for the smart endpointing plan.
          - LivekitSmartEndpointingPlan
            - `provider` 'vapi' | 'livekit' | 'custom-endpointing-model', required — This is the provider for the smart endpointing plan.
            - `waitFunction` string — This expression describes how long the bot will wait to start speaking based on the likelihood that the user has reached an endpoint. This is a millisecond valued function. It maps probabilities (real numbers on [0,1]) to milliseconds that the bot should wait before speaking ([0, \infty]). Any negative values that are returned are set to zero (the bot can't start talking in the past). A probability of zero represents very high confidence that the caller has stopped speaking, and would like the bot to speak to them. A probability of one represents very high confidence that the caller is still speaking. Under the hood, this is parsed into a mathjs expression. Whatever you use to write your expression needs to be valid with respect to mathjs @default "20 + 500 * sqrt(x) + 2500 * x^3"
          - CustomEndpointingModelSmartEndpointingPlan
            - `provider` 'vapi' | 'livekit' | 'custom-endpointing-model', required — This is the provider for the smart endpointing plan. Use `custom-endpointing-model` for custom endpointing providers that are not natively supported.
            - `server` Server
              - …
        - `customEndpointingRules` union[] — These are the custom endpointing rules to set an endpointing timeout based on a regex on the customer's speech or the assistant's last message. Usage: - If you have yes/no questions like "are you interested in a loan?", you can set a shorter timeout. - If you have questions where the customer may pause to look up information like "what's my account number?", you can set a longer timeout. - If you want to wait longer while customer is enumerating a list of numbers, you can set a longer timeout. These rules have the highest precedence and will override both `smartEndpointingPlan` and `transcriptionEndpointingPlan` when a rule is matched. The rules are evaluated in order and the first one that matches will be used. Order of precedence for endpointing: 1. customEndpointingRules (if any match) 2. smartEndpointingPlan (if set) 3. transcriptionEndpointingPlan @default []
          - union
            - AssistantCustomEndpointingRule
              - …
            - CustomerCustomEndpointingRule
              - …
            - BothCustomEndpointingRule
              - …
        - `transcriptionEndpointingPlan` TranscriptionEndpointingPlan
          - `onPunctuationSeconds` number — The minimum number of seconds to wait after transcription ending with punctuation before sending a request to the model. Defaults to 0.1. This setting exists because the transcriber punctuates the transcription when it's more confident that customer has completed a thought. @default 0.1
          - `onNoPunctuationSeconds` number — The minimum number of seconds to wait after transcription ending without punctuation before sending a request to the model. Defaults to 1.5. This setting exists to catch the cases where the transcriber was not confident enough to punctuate the transcription, but the customer is done and has been silent for a long time. @default 1.5
          - `onNumberSeconds` number — The minimum number of seconds to wait after transcription ending with a number before sending a request to the model. Defaults to 0.4. This setting exists because the transcriber will sometimes punctuate the transcription ending with a number, even though the customer hasn't uttered the full number. This happens commonly for long numbers when the customer reads the number in chunks. @default 0.5
      - `stopSpeakingPlan` StopSpeakingPlan
        - `numWords` number — This is the number of words that the customer has to say before the assistant will stop talking. Words like "stop", "actually", "no", etc. will always interrupt immediately regardless of this value. Words like "okay", "yeah", "right" will never interrupt. When set to 0, `voiceSeconds` is used in addition to the transcriptions to determine the customer has started speaking. Defaults to 0. @default 0
        - `voiceSeconds` number — This is the seconds customer has to speak before the assistant stops talking. This uses the VAD (Voice Activity Detection) spike to determine if the customer has started speaking. Considerations: - A lower value might be more responsive but could potentially pick up non-speech sounds. - A higher value reduces false positives but might slightly delay the detection of speech onset. This is only used if `numWords` is set to 0. Defaults to 0.2 @default 0.2
        - `backoffSeconds` number — This is the seconds to wait before the assistant will start talking again after being interrupted. Defaults to 1. @default 1
        - `acknowledgementPhrases` string[] — These are the phrases that will never interrupt the assistant, even if numWords threshold is met. These are typically acknowledgement or backchanneling phrases.
        - `interruptionPhrases` string[] — These are the phrases that will always interrupt the assistant immediately, regardless of numWords. These are typically phrases indicating disagreement or desire to stop.
      - `monitorPlan` MonitorPlan
        - `listenEnabled` boolean — This determines whether the assistant's calls allow live listening. Defaults to true. Fetch `call.monitor.listenUrl` to get the live listening URL. @default true
        - `listenAuthenticationEnabled` boolean — This enables authentication on the `call.monitor.listenUrl`. If `listenAuthenticationEnabled` is `true`, the `call.monitor.listenUrl` will require an `Authorization: Bearer <vapi-public-api-key>` header. @default false
        - `controlEnabled` boolean — This determines whether the assistant's calls allow live control. Defaults to true. Fetch `call.monitor.controlUrl` to get the live control URL. To use, send any control message via a POST request to `call.monitor.controlUrl`. Here are the types of controls supported: https://docs.vapi.ai/api-reference/messages/client-inbound-message @default true
        - `controlAuthenticationEnabled` boolean — This enables authentication on the `call.monitor.controlUrl`. If `controlAuthenticationEnabled` is `true`, the `call.monitor.controlUrl` will require an `Authorization: Bearer <vapi-public-api-key>` header. @default false
        - `monitorIds` string[] — This the set of monitor ids that are attached to the assistant. The source of truth for the monitor ids is the assistant_monitor join table. This field can be used for transient assistants and to update assistants with new monitor ids. @default []
      - `credentialIds` string[] — These are the credentials that will be used for the assistant calls. By default, all the credentials are available for use in the call but you can provide a subset using this.
      - `server` Server
        - `timeoutSeconds` number — This is the timeout in seconds for the request. Defaults to 20 seconds. @default 20
        - `credentialId` string — The credential ID for server authentication
        - `staticIpAddressesEnabled` boolean — If enabled, requests will originate from a static set of IPs owned and managed by Vapi. @default false
        - `encryptedPaths` string[] — This is the paths to encrypt in the request body if credentialId and encryptionPlan are defined.
        - `url` string — This is where the request will be sent.
        - `headers` object — These are the headers to include in the request. Each key-value pair represents a header name and its value. Note: Specifying an Authorization header here will override the authorization provided by the `credentialId` (if provided). This is an anti-pattern and should be avoided outside of edge case scenarios.
        - `backoffPlan` BackoffPlan
          - `type` object, required — This is the type of backoff plan to use. Defaults to fixed. @default fixed
          - `maxRetries` number, required — This is the maximum number of retries to attempt if the request fails. Defaults to 0 (no retries). @default 0
          - `baseDelaySeconds` number, required — This is the base delay in seconds. For linear backoff, this is the delay between each retry. For exponential backoff, this is the initial delay.
          - `excludedStatusCodes` object[] — This is the excluded status codes. If the response status code is in this list, the request will not be retried. By default, the request will be retried for any non-2xx status code.
      - `keypadInputPlan` KeypadInputPlan
        - `enabled` boolean — This keeps track of whether the user has enabled keypad input. By default, it is off. @default false
        - `timeoutSeconds` number — This is the time in seconds to wait before processing the input. If the input is not received within this time, the input will be ignored. If set to "off", the input will be processed when the user enters a delimiter or immediately if no delimiter is used. @default 2
        - `delimiters` '#' | '*' | '' — This is the delimiter(s) that will be used to process the input. Can be '#', '*', or an empty array.
    - `assistantOverrides` AssistantOverrides
      - `transcriber` union — These are the options for the assistant's transcriber.
        - AssemblyAITranscriber
          - `provider` 'assembly-ai', required — This is the transcription provider that will be used.
          - `language` 'multi' | 'en' — This is the language that will be set for the transcription.
          - `confidenceThreshold` number — Transcripts below this confidence threshold will be discarded. @default 0.4
          - `formatTurns` boolean — This enables formatting of transcripts. @default true
          - `endOfTurnConfidenceThreshold` number — This is the end of turn confidence threshold. The minimum confidence that the end of turn is detected. Note: Only used if startSpeakingPlan.smartEndpointingPlan is not set. @min 0 @max 1 @default 0.7
          - `minEndOfTurnSilenceWhenConfident` number — This is the minimum end of turn silence when confident in milliseconds. Note: Only used if startSpeakingPlan.smartEndpointingPlan is not set. @default 160
          - `wordFinalizationMaxWaitTime` number
          - `maxTurnSilence` number — This is the maximum turn silence time in milliseconds. Note: Only used if startSpeakingPlan.smartEndpointingPlan is not set. @default 400
          - `vadAssistedEndpointingEnabled` boolean — Use VAD to assist with endpointing decisions from the transcriber. When enabled, transcriber endpointing will be buffered if VAD detects the user is still speaking, preventing premature turn-taking. When disabled, transcriber endpointing will be used immediately regardless of VAD state, allowing for quicker but more aggressive turn-taking. Note: Only used if startSpeakingPlan.smartEndpointingPlan is not set. @default true
          - `mode` 'max_accuracy' | 'min_latency' | 'balanced' — This is the transcription mode used by the `universal-3-5-pro` speech model. Only applies to the `universal-3-5-pro` speech model. @default 'balanced'
          - `prompt` string — This is a prompt that provides additional context to the transcription model. Only applies to the `universal-3-5-pro` speech model.
          - `agentContext` string — This is context about the voice agent that guides the transcription model. Only applies to the `universal-3-5-pro` speech model.
          - `languageCodes` string[] — These are language codes used to steer automatic language detection. Only applies to the `universal-3-5-pro` speech model.
          - `speechModel` 'universal-streaming-english' | 'universal-streaming-multilingual' | 'universal-3-5-pro' — This is the speech model used for the streaming session. Keyterms prompting is supported on universal-streaming-english and universal-3-5-pro. universal-3-5-pro is AssemblyAI's most accurate voice-agent model. @default 'universal-streaming-english'
          - `realtimeUrl` string — The WebSocket URL that the transcriber connects to.
          - `wordBoost` string[] — Add up to 2500 characters of custom vocabulary.
          - `keytermsPrompt` string[] — Keyterms prompting improves recognition accuracy for specific words and phrases. Can include up to 100 keyterms, each up to 50 characters. Costs an additional $0.04/hour on universal-streaming-english and is included at no extra cost on universal-3-5-pro.
          - `endUtteranceSilenceThreshold` number — The duration of the end utterance silence threshold in milliseconds.
          - `disablePartialTranscripts` boolean — Disable partial transcripts. Set to `true` to not receive partial transcripts. Defaults to `false`.
          - `fallbackPlan` FallbackTranscriberPlan
            - `transcribers` union[]
              - …
        - AzureSpeechTranscriber
          - `provider` 'azure', required — This is the transcription provider that will be used.
          - `language` 'af-ZA' | 'am-ET' | 'ar-AE' | 'ar-BH' | 'ar-DZ' | 'ar-EG' | 'ar-IL' | 'ar-IQ' | 'ar-JO' | 'ar-KW' | 'ar-LB' | 'ar-LY' | 'ar-MA' | 'ar-OM' | 'ar-PS' | 'ar-QA' | 'ar-SA' | 'ar-SY' | 'ar-TN' | 'ar-YE' | 'az-AZ' | 'bg-BG' | 'bn-IN' | 'bs-BA' | 'ca-ES' | 'cs-CZ' | 'cy-GB' | 'da-DK' | 'de-AT' | 'de-CH' | 'de-DE' | 'el-GR' | 'en-AU' | 'en-CA' | 'en-GB' | 'en-GH' | 'en-HK' | 'en-IE' | 'en-IN' | 'en-KE' | 'en-NG' | 'en-NZ' | 'en-PH' | 'en-SG' | 'en-TZ' | 'en-US' | 'en-ZA' | 'es-AR' | 'es-BO' | 'es-CL' | 'es-CO' | 'es-CR' | 'es-CU' | 'es-DO' | 'es-EC' | 'es-ES' | 'es-GQ' | 'es-GT' | 'es-HN' | 'es-MX' | 'es-NI' | 'es-PA' | 'es-PE' | 'es-PR' | 'es-PY' | 'es-SV' | 'es-US' | 'es-UY' | 'es-VE' | 'et-EE' | 'eu-ES' | 'fa-IR' | 'fi-FI' | 'fil-PH' | 'fr-BE' | 'fr-CA' | 'fr-CH' | 'fr-FR' | 'ga-IE' | 'gl-ES' | 'gu-IN' | 'he-IL' | 'hi-IN' | 'hr-HR' | 'hu-HU' | 'hy-AM' | 'id-ID' | 'is-IS' | 'it-CH' | 'it-IT' | 'ja-JP' | 'jv-ID' | 'ka-GE' | 'kk-KZ' | 'km-KH' | 'kn-IN' | 'ko-KR' | 'lo-LA' | 'lt-LT' | 'lv-LV' | 'mk-MK' | 'ml-IN' | 'mn-MN' | 'mr-IN' | 'ms-MY' | 'mt-MT' | 'my-MM' | 'nb-NO' | 'ne-NP' | 'nl-BE' | 'nl-NL' | 'pa-IN' | 'pl-PL' | 'ps-AF' | 'pt-BR' | 'pt-PT' | 'ro-RO' | 'ru-RU' | 'si-LK' | 'sk-SK' | 'sl-SI' | 'so-SO' | 'sq-AL' | 'sr-RS' | 'sv-SE' | 'sw-KE' | 'sw-TZ' | 'ta-IN' | 'te-IN' | 'th-TH' | 'tr-TR' | 'uk-UA' | 'ur-IN' | 'uz-UZ' | 'vi-VN' | 'wuu-CN' | 'yue-CN' | 'zh-CN' | 'zh-CN-shandong' | 'zh-CN-sichuan' | 'zh-HK' | 'zh-TW' | 'zu-ZA' — This is the language that will be set for the transcription. The list of languages Azure supports can be found here: https://learn.microsoft.com/en-us/azure/ai-services/speech-service/language-support?tabs=stt
          - `segmentationStrategy` 'Default' | 'Time' | 'Semantic' — Controls how phrase boundaries are detected, enabling either simple time/silence heuristics or more advanced semantic segmentation.
          - `segmentationSilenceTimeoutMs` number — Duration of detected silence after which the service finalizes a phrase. Configure to adjust sensitivity to pauses in speech.
          - `segmentationMaximumTimeMs` number — Maximum duration a segment can reach before being cut off when using time-based segmentation.
          - `fallbackPlan` FallbackTranscriberPlan
            - `transcribers` union[]
              - …
        - CustomTranscriber
          - `provider` 'custom-transcriber', required — This is the transcription provider that will be used. Use `custom-transcriber` for providers that are not natively supported.
          - `server` Server, required
            - `timeoutSeconds` number — This is the timeout in seconds for the request. Defaults to 20 seconds. @default 20
            - `credentialId` string — The credential ID for server authentication
            - `staticIpAddressesEnabled` boolean — If enabled, requests will originate from a static set of IPs owned and managed by Vapi. @default false
            - `encryptedPaths` string[] — This is the paths to encrypt in the request body if credentialId and encryptionPlan are defined.
            - `url` string — This is where the request will be sent.
            - `headers` object — These are the headers to include in the request. Each key-value pair represents a header name and its value. Note: Specifying an Authorization header here will override the authorization provided by the `credentialId` (if provided). This is an anti-pattern and should be avoided outside of edge case scenarios.
            - `backoffPlan` BackoffPlan
              - …
          - `fallbackPlan` FallbackTranscriberPlan
            - `transcribers` union[]
              - …
        - DeepgramTranscriber
          - `provider` 'deepgram', required — This is the transcription provider that will be used.
          - `model` union — This is the Deepgram model that will be used. A list of models can be found here: https://developers.deepgram.com/docs/models-languages-overview
            - 'nova-3' | 'nova-3-general' | 'nova-3-medical' | 'nova-2' | 'nova-2-general' | 'nova-2-meeting' | 'nova-2-phonecall' | 'nova-2-finance' | 'nova-2-conversationalai' | 'nova-2-voicemail' | 'nova-2-video' | 'nova-2-medical' | 'nova-2-drivethru' | 'nova-2-automotive' | 'nova' | 'nova-general' | 'nova-phonecall' | 'nova-medical' | 'enhanced' | 'enhanced-general' | 'enhanced-meeting' | 'enhanced-phonecall' | 'enhanced-finance' | 'base' | 'base-general' | 'base-meeting' | 'base-phonecall' | 'base-finance' | 'base-conversationalai' | 'base-voicemail' | 'base-video' | 'whisper' | 'flux-general-en' | 'flux-general-multi'
            - string
          - `language` 'ar' | 'az' | 'ba' | 'be' | 'bg' | 'bn' | 'br' | 'bs' | 'ca' | 'cs' | 'da' | 'da-DK' | 'de' | 'de-CH' | 'el' | 'en' | 'en-AU' | 'en-CA' | 'en-GB' | 'en-IE' | 'en-IN' | 'en-NZ' | 'en-US' | 'es' | 'es-419' | 'es-LATAM' | 'et' | 'eu' | 'fa' | 'fi' | 'fr' | 'fr-CA' | 'ha' | 'haw' | 'he' | 'hi' | 'hi-Latn' | 'hr' | 'hu' | 'id' | 'is' | 'it' | 'ja' | 'jw' | 'kn' | 'ko' | 'ko-KR' | 'ln' | 'lt' | 'lv' | 'mk' | 'mr' | 'ms' | 'multi' | 'nl' | 'nl-BE' | 'no' | 'pl' | 'pt' | 'pt-BR' | 'pt-PT' | 'ro' | 'ru' | 'sk' | 'sl' | 'sn' | 'so' | 'sr' | 'su' | 'sv' | 'sv-SE' | 'ta' | 'taq' | 'te' | 'th' | 'th-TH' | 'tl' | 'tr' | 'tt' | 'uk' | 'ur' | 'vi' | 'yo' | 'zh' | 'zh-CN' | 'zh-HK' | 'zh-Hans' | 'zh-Hant' | 'zh-TW' — This is the language that will be set for the transcription. The list of languages Deepgram supports can be found here: https://developers.deepgram.com/docs/models-languages-overview
          - `smartFormat` boolean — This will be use smart format option provided by Deepgram. It's default disabled because it can sometimes format numbers as times but it's getting better.
          - `mipOptOut` boolean — If set to true, this will add mip_opt_out=true as a query parameter of all API requests. See https://developers.deepgram.com/docs/the-deepgram-model-improvement-partnership-program#want-to-opt-out This will only be used if you are using your own Deepgram API key. @default false
          - `numerals` boolean — If set to true, this will cause deepgram to convert spoken numbers to literal numerals. For example, "my phone number is nine-seven-two..." would become "my phone number is 972..." @default false
          - `profanityFilter` boolean — If set to true, Deepgram will replace profanity in transcripts with surrounding asterisks, e.g. "f***". @default false
          - `redaction` string[] — Enables redaction of sensitive information from transcripts. Options include: - "pci": Redacts credit card numbers, expiration dates, and CVV. - "pii": Redacts personally identifiable information (names, locations, identifying numbers, etc.). - "phi": Redacts protected health information (medical conditions, drugs, injuries, etc.). - "numbers": Redacts numerical and identifying entities (dates, account numbers, SSNs, etc.). Multiple values can be provided to redact different categories simultaneously. Redacted content is replaced with entity labels like [CREDIT_CARD_1], [SSN_1], etc. See https://developers.deepgram.com/docs/redaction for details.
          - `confidenceThreshold` number — Transcripts below this confidence threshold will be discarded. @default 0.4
          - `eotThreshold` number — End-of-turn confidence required to finish a turn. Only used with Flux models. @default 0.7
          - `eotTimeoutMs` number — A turn will be finished when this much time has passed after speech, regardless of EOT confidence. Only used with Flux models. @default 5000
          - `languages` string[] — Language hints to bias Flux Multilingual (`flux-general-multi`) toward specific languages. Provide BCP-47 language codes (e.g. "en", "es", "fr"). Multiple hints can be given for multilingual or code-switching scenarios. Omit for auto-detection. Only used with `flux-general-multi`.
          - `keywords` string[] — These keywords are passed to the transcription model to help it pick up use-case specific words. Anything that may not be a common word, like your company name, should be added here.
          - `keyterm` string[] — Keyterm Prompting allows you improve Keyword Recall Rate (KRR) for important keyterms or phrases up to 90%.
          - `endpointing` number — This is the timeout after which Deepgram will send transcription on user silence. You can read in-depth documentation here: https://developers.deepgram.com/docs/endpointing. Here are the most important bits: - Defaults to 10. This is recommended for most use cases to optimize for latency. - 10 can cause some missing transcriptions since because of the shorter context. This mostly happens for one-word utterances. For those uses cases, it's recommended to try 300. It will add a bit of latency but the quality and reliability of the experience will be better. - If neither 10 nor 300 work, contact support@vapi.ai and we'll find another solution. @default 10
          - `fallbackPlan` FallbackTranscriberPlan
            - `transcribers` union[]
              - …
        - ElevenLabsTranscriber
          - `provider` '11labs', required — This is the transcription provider that will be used.
          - `model` 'scribe_v1' | 'scribe_v2' | 'scribe_v2_realtime' — This is the model that will be used for the transcription.
          - `language` 'aa' | 'ab' | 'ae' | 'af' | 'ak' | 'am' | 'an' | 'ar' | 'as' | 'av' | 'ay' | 'az' | 'ba' | 'be' | 'bg' | 'bh' | 'bi' | 'bm' | 'bn' | 'bo' | 'br' | 'bs' | 'ca' | 'ce' | 'ch' | 'co' | 'cr' | 'cs' | 'cu' | 'cv' | 'cy' | 'da' | 'de' | 'dv' | 'dz' | 'ee' | 'el' | 'en' | 'eo' | 'es' | 'et' | 'eu' | 'fa' | 'ff' | 'fi' | 'fj' | 'fo' | 'fr' | 'fy' | 'ga' | 'gd' | 'gl' | 'gn' | 'gu' | 'gv' | 'ha' | 'he' | 'hi' | 'ho' | 'hr' | 'ht' | 'hu' | 'hy' | 'hz' | 'ia' | 'id' | 'ie' | 'ig' | 'ii' | 'ik' | 'io' | 'is' | 'it' | 'iu' | 'ja' | 'jv' | 'ka' | 'kg' | 'ki' | 'kj' | 'kk' | 'kl' | 'km' | 'kn' | 'ko' | 'kr' | 'ks' | 'ku' | 'kv' | 'kw' | 'ky' | 'la' | 'lb' | 'lg' | 'li' | 'ln' | 'lo' | 'lt' | 'lu' | 'lv' | 'mg' | 'mh' | 'mi' | 'mk' | 'ml' | 'mn' | 'mr' | 'ms' | 'mt' | 'my' | 'na' | 'nb' | 'nd' | 'ne' | 'ng' | 'nl' | 'nn' | 'no' | 'nr' | 'nv' | 'ny' | 'oc' | 'oj' | 'om' | 'or' | 'os' | 'pa' | 'pi' | 'pl' | 'ps' | 'pt' | 'qu' | 'rm' | 'rn' | 'ro' | 'ru' | 'rw' | 'sa' | 'sc' | 'sd' | 'se' | 'sg' | 'si' | 'sk' | 'sl' | 'sm' | 'sn' | 'so' | 'sq' | 'sr' | 'ss' | 'st' | 'su' | 'sv' | 'sw' | 'ta' | 'te' | 'tg' | 'th' | 'ti' | 'tk' | 'tl' | 'tn' | 'to' | 'tr' | 'ts' | 'tt' | 'tw' | 'ty' | 'ug' | 'uk' | 'ur' | 'uz' | 've' | 'vi' | 'vo' | 'wa' | 'wo' | 'xh' | 'yi' | 'yue' | 'yo' | 'za' | 'zh' | 'zu' — This is the language that will be used for the transcription.
          - `silenceThresholdSeconds` number — This is the number of seconds of silence before VAD commits (0.3-3.0).
          - `confidenceThreshold` number — This is the VAD sensitivity (0.1-0.9, lower indicates more sensitive).
          - `minSpeechDurationMs` number — This is the minimum speech duration for VAD (50-2000ms).
          - `minSilenceDurationMs` number — This is the minimum silence duration for VAD (50-2000ms).
          - `fallbackPlan` FallbackTranscriberPlan
            - `transcribers` union[]
              - …
        - GladiaTranscriber
          - `provider` 'gladia', required — This is the transcription provider that will be used.
          - `model` 'fast' | 'accurate' | 'solaria-1' — This is the Gladia model that will be used. Default is 'fast'
          - `languageBehaviour` 'manual' | 'automatic single language' | 'automatic multiple languages' — Defines how the transcription model detects the audio language. Default value is 'automatic single language'.
          - `language` 'af' | 'sq' | 'am' | 'ar' | 'hy' | 'as' | 'az' | 'ba' | 'eu' | 'be' | 'bn' | 'bs' | 'br' | 'bg' | 'ca' | 'zh' | 'hr' | 'cs' | 'da' | 'nl' | 'en' | 'et' | 'fo' | 'fi' | 'fr' | 'gl' | 'ka' | 'de' | 'el' | 'gu' | 'ht' | 'ha' | 'haw' | 'he' | 'hi' | 'hu' | 'is' | 'id' | 'it' | 'ja' | 'jv' | 'kn' | 'kk' | 'km' | 'ko' | 'lo' | 'la' | 'lv' | 'ln' | 'lt' | 'lb' | 'mk' | 'mg' | 'ms' | 'ml' | 'mt' | 'mi' | 'mr' | 'mn' | 'my' | 'ne' | 'no' | 'nn' | 'oc' | 'ps' | 'fa' | 'pl' | 'pt' | 'pa' | 'ro' | 'ru' | 'sa' | 'sr' | 'sn' | 'sd' | 'si' | 'sk' | 'sl' | 'so' | 'es' | 'su' | 'sw' | 'sv' | 'tl' | 'tg' | 'ta' | 'tt' | 'te' | 'th' | 'bo' | 'tr' | 'tk' | 'uk' | 'ur' | 'uz' | 'vi' | 'cy' | 'yi' | 'yo' — Defines the language to use for the transcription. Required when languageBehaviour is 'manual'.
          - `languages` string[] — Defines the languages to use for the transcription. Required when languageBehaviour is 'manual'.
          - `transcriptionHint` string — Provides a custom vocabulary to the model to improve accuracy of transcribing context specific words, technical terms, names, etc. If empty, this argument is ignored. ⚠️ Warning ⚠️: Please be aware that the transcription_hint field has a character limit of 600. If you provide a transcription_hint longer than 600 characters, it will be automatically truncated to meet this limit.
          - `prosody` boolean — If prosody is true, you will get a transcription that can contain prosodies i.e. (laugh) (giggles) (malefic laugh) (toss) (music)… Default value is false.
          - `audioEnhancer` boolean — If true, audio will be pre-processed to improve accuracy but latency will increase. Default value is false.
          - `confidenceThreshold` number — Transcripts below this confidence threshold will be discarded. @default 0.4
          - `endpointing` number — Endpointing time in seconds - time to wait before considering speech ended
          - `speechThreshold` number — Speech threshold - sensitivity configuration for speech detection (0.0 to 1.0)
          - `customVocabularyEnabled` boolean — Enable custom vocabulary for improved accuracy
          - `customVocabularyConfig` GladiaCustomVocabularyConfigDTO
            - `vocabulary` union[], required — Array of vocabulary items (strings or objects with value, pronunciations, intensity, language)
              - …
            - `defaultIntensity` number — Default intensity for vocabulary items (0.0 to 1.0)
          - `region` 'us-west' | 'eu-west' — Region for processing audio (us-west or eu-west)
          - `receivePartialTranscripts` boolean — Enable partial transcripts for low-latency streaming transcription
          - `fallbackPlan` FallbackTranscriberPlan
            - `transcribers` union[]
              - …
        - GoogleTranscriber
          - `provider` 'google', required — This is the transcription provider that will be used.
          - `model` 'gemini-3.5-flash' | 'gemini-3.1-flash-lite' | 'gemini-3-flash-preview' | 'gemini-2.5-pro' | 'gemini-2.5-flash' | 'gemini-2.5-flash-lite' | 'gemini-2.0-flash-thinking-exp' | 'gemini-2.0-pro-exp-02-05' | 'gemini-2.0-flash' | 'gemini-2.0-flash-lite' | 'gemini-2.0-flash-exp' | 'gemini-2.0-flash-realtime-exp' | 'gemini-1.5-flash' | 'gemini-1.5-flash-002' | 'gemini-1.5-pro' | 'gemini-1.5-pro-002' | 'gemini-1.0-pro' — This is the model that will be used for the transcription.
          - `language` 'Multilingual' | 'Arabic' | 'Bengali' | 'Bulgarian' | 'Chinese' | 'Croatian' | 'Czech' | 'Danish' | 'Dutch' | 'English' | 'Estonian' | 'Finnish' | 'French' | 'German' | 'Greek' | 'Hebrew' | 'Hindi' | 'Hungarian' | 'Indonesian' | 'Italian' | 'Japanese' | 'Korean' | 'Latvian' | 'Lithuanian' | 'Norwegian' | 'Polish' | 'Portuguese' | 'Romanian' | 'Russian' | 'Serbian' | 'Slovak' | 'Slovenian' | 'Spanish' | 'Swahili' | 'Swedish' | 'Thai' | 'Turkish' | 'Ukrainian' | 'Vietnamese' — This is the language that will be set for the transcription.
          - `fallbackPlan` FallbackTranscriberPlan
            - `transcribers` union[]
              - …
        - SpeechmaticsTranscriber
          - `provider` 'speechmatics', required — This is the transcription provider that will be used.
          - `model` 'default' — This is the model that will be used for the transcription.
          - `language` 'auto' | 'ar' | 'ar_en' | 'ba' | 'eu' | 'be' | 'bn' | 'bg' | 'yue' | 'ca' | 'hr' | 'cs' | 'da' | 'nl' | 'en' | 'eo' | 'et' | 'fi' | 'fr' | 'gl' | 'de' | 'el' | 'he' | 'hi' | 'hu' | 'id' | 'ia' | 'ga' | 'it' | 'ja' | 'ko' | 'lv' | 'lt' | 'ms' | 'en_ms' | 'mt' | 'cmn' | 'cmn_en' | 'mr' | 'mn' | 'no' | 'fa' | 'pl' | 'pt' | 'ro' | 'ru' | 'sk' | 'sl' | 'es' | 'en_es' | 'sw' | 'sv' | 'tl' | 'ta' | 'en_ta' | 'th' | 'tr' | 'uk' | 'ur' | 'ug' | 'vi' | 'cy'
          - `operatingPoint` 'standard' | 'enhanced' — This is the operating point for the transcription. Choose between `standard` for faster turnaround with strong accuracy or `enhanced` for highest accuracy when precision is critical. @default 'enhanced'
          - `region` 'eu' | 'us' — This is the region for the Speechmatics API. Choose between EU (Europe) and US (United States) regions for lower latency and data sovereignty compliance. @default 'eu'
          - `enableDiarization` boolean — This enables speaker diarization, which identifies and separates speakers in the transcription. Essential for multi-speaker conversations and conference calls. @default false
          - `maxDelay` number — This sets the maximum delay in milliseconds for partial transcripts. Balances latency and accuracy. @default 3000
          - `customVocabulary` SpeechmaticsCustomVocabularyItem[], required
            - `content` string, required — The word or phrase to add to the custom vocabulary.
            - `soundsLike` string[] — Alternative phonetic representations of how the word might sound. This helps recognition when the word might be pronounced differently.
          - `numeralStyle` 'written' | 'spoken' — This controls how numbers, dates, currencies, and other entities are formatted in the transcription output. @default 'written'
          - `endOfTurnSensitivity` number — This is the sensitivity level for end-of-turn detection, which determines when a speaker has finished talking. Higher values are more sensitive. @default 0.5
          - `removeDisfluencies` boolean — This enables removal of disfluencies (um, uh) from the transcript to create cleaner, more professional output. This is only supported for the English language transcriber. @default false
          - `minimumSpeechDuration` number — This is the minimum duration in seconds for speech segments. Shorter segments will be filtered out. Helps remove noise and improve accuracy. @default 0.0
          - `fallbackPlan` FallbackTranscriberPlan
            - `transcribers` union[]
              - …
        - TalkscriberTranscriber
          - `provider` 'talkscriber', required — This is the transcription provider that will be used.
          - `model` 'whisper' — This is the model that will be used for the transcription.
          - `language` 'en' | 'zh' | 'de' | 'es' | 'ru' | 'ko' | 'fr' | 'ja' | 'pt' | 'tr' | 'pl' | 'ca' | 'nl' | 'ar' | 'sv' | 'it' | 'id' | 'hi' | 'fi' | 'vi' | 'he' | 'uk' | 'el' | 'ms' | 'cs' | 'ro' | 'da' | 'hu' | 'ta' | 'no' | 'th' | 'ur' | 'hr' | 'bg' | 'lt' | 'la' | 'mi' | 'ml' | 'cy' | 'sk' | 'te' | 'fa' | 'lv' | 'bn' | 'sr' | 'az' | 'sl' | 'kn' | 'et' | 'mk' | 'br' | 'eu' | 'is' | 'hy' | 'ne' | 'mn' | 'bs' | 'kk' | 'sq' | 'sw' | 'gl' | 'mr' | 'pa' | 'si' | 'km' | 'sn' | 'yo' | 'so' | 'af' | 'oc' | 'ka' | 'be' | 'tg' | 'sd' | 'gu' | 'am' | 'yi' | 'lo' | 'uz' | 'fo' | 'ht' | 'ps' | 'tk' | 'nn' | 'mt' | 'sa' | 'lb' | 'my' | 'bo' | 'tl' | 'mg' | 'as' | 'tt' | 'haw' | 'ln' | 'ha' | 'ba' | 'jw' | 'su' | 'yue' — This is the language that will be set for the transcription. The list of languages Whisper supports can be found here: https://github.com/openai/whisper/blob/main/whisper/tokenizer.py
          - `fallbackPlan` FallbackTranscriberPlan
            - `transcribers` union[]
              - …
        - OpenAITranscriber
          - `provider` 'openai', required — This is the transcription provider that will be used.
          - `model` 'gpt-4o-transcribe' | 'gpt-4o-mini-transcribe', required — This is the model that will be used for the transcription.
          - `language` 'af' | 'ar' | 'hy' | 'az' | 'be' | 'bs' | 'bg' | 'ca' | 'zh' | 'hr' | 'cs' | 'da' | 'nl' | 'en' | 'et' | 'fi' | 'fr' | 'gl' | 'de' | 'el' | 'he' | 'hi' | 'hu' | 'is' | 'id' | 'it' | 'ja' | 'kn' | 'kk' | 'ko' | 'lv' | 'lt' | 'mk' | 'ms' | 'mr' | 'mi' | 'ne' | 'no' | 'fa' | 'pl' | 'pt' | 'ro' | 'ru' | 'sr' | 'sk' | 'sl' | 'es' | 'sw' | 'sv' | 'tl' | 'ta' | 'th' | 'tr' | 'uk' | 'ur' | 'vi' | 'cy' — This is the language that will be set for the transcription.
          - `fallbackPlan` FallbackTranscriberPlan
            - `transcribers` union[]
              - …
        - CartesiaTranscriber
          - `provider` 'cartesia', required
          - `model` 'ink-whisper' | 'ink-2'
          - `language` 'aa' | 'ab' | 'ae' | 'af' | 'ak' | 'am' | 'an' | 'ar' | 'as' | 'av' | 'ay' | 'az' | 'ba' | 'be' | 'bg' | 'bh' | 'bi' | 'bm' | 'bn' | 'bo' | 'br' | 'bs' | 'ca' | 'ce' | 'ch' | 'co' | 'cr' | 'cs' | 'cu' | 'cv' | 'cy' | 'da' | 'de' | 'dv' | 'dz' | 'ee' | 'el' | 'en' | 'eo' | 'es' | 'et' | 'eu' | 'fa' | 'ff' | 'fi' | 'fj' | 'fo' | 'fr' | 'fy' | 'ga' | 'gd' | 'gl' | 'gn' | 'gu' | 'gv' | 'ha' | 'he' | 'hi' | 'ho' | 'hr' | 'ht' | 'hu' | 'hy' | 'hz' | 'ia' | 'id' | 'ie' | 'ig' | 'ii' | 'ik' | 'io' | 'is' | 'it' | 'iu' | 'ja' | 'jv' | 'ka' | 'kg' | 'ki' | 'kj' | 'kk' | 'kl' | 'km' | 'kn' | 'ko' | 'kr' | 'ks' | 'ku' | 'kv' | 'kw' | 'ky' | 'la' | 'lb' | 'lg' | 'li' | 'ln' | 'lo' | 'lt' | 'lu' | 'lv' | 'mg' | 'mh' | 'mi' | 'mk' | 'ml' | 'mn' | 'mr' | 'ms' | 'mt' | 'my' | 'na' | 'nb' | 'nd' | 'ne' | 'ng' | 'nl' | 'nn' | 'no' | 'nr' | 'nv' | 'ny' | 'oc' | 'oj' | 'om' | 'or' | 'os' | 'pa' | 'pi' | 'pl' | 'ps' | 'pt' | 'qu' | 'rm' | 'rn' | 'ro' | 'ru' | 'rw' | 'sa' | 'sc' | 'sd' | 'se' | 'sg' | 'si' | 'sk' | 'sl' | 'sm' | 'sn' | 'so' | 'sq' | 'sr' | 'ss' | 'st' | 'su' | 'sv' | 'sw' | 'ta' | 'te' | 'tg' | 'th' | 'ti' | 'tk' | 'tl' | 'tn' | 'to' | 'tr' | 'ts' | 'tt' | 'tw' | 'ty' | 'ug' | 'uk' | 'ur' | 'uz' | 've' | 'vi' | 'vo' | 'wa' | 'wo' | 'xh' | 'yi' | 'yue' | 'yo' | 'za' | 'zh' | 'zu'
          - `fallbackPlan` FallbackTranscriberPlan
            - `transcribers` union[]
              - …
        - SonioxTranscriber
          - `provider` 'soniox', required
- … truncated; see the full OpenAPI document linked below

## Changes

> 57 revisions in range; 2 not diffed, 14 could not be searched.

- **2026-03-07** `de31e89ab2f9` — 118 info
  - added the optional property `members/items/assistant/allOf[#/components/schemas/CreateAssistantDTO]/hooks/items/oneOf[subschema #1: CallHookCallEnding]/do/items/oneOf[subschema #1: ToolCallHookAction]/tool/oneOf[subschema #1: ApiRequestTool]/parameters` to the response with the `200` status
  - added the optional property `members/items/assistant/allOf[#/components/schemas/CreateAssistantDTO]/hooks/items/oneOf[subschema #1: CallHookCallEnding]/do/items/oneOf[subschema #1: ToolCallHookAction]/tool/oneOf[subschema #7: FunctionTool]/parameters` to the response with the `200` status
  - added the optional property `members/items/assistant/allOf[#/components/schemas/CreateAssistantDTO]/hooks/items/oneOf[subschema #2: CallHookAssistantSpeechInterrupted]/do/items/oneOf[subschema #2: ToolCallHookAction]/tool/oneOf[subschema #1: ApiRequestTool]/parameters` to the response with the `200` status
  - added the optional property `members/items/assistant/allOf[#/components/schemas/CreateAssistantDTO]/hooks/items/oneOf[subschema #2: CallHookAssistantSpeechInterrupted]/do/items/oneOf[subschema #2: ToolCallHookAction]/tool/oneOf[subschema #7: FunctionTool]/parameters` to the response with the `200` status
  - …114 more
- …earlier changes not shown

[Full history](https://skmtc.dev/vapiai/apis/vapi-api/changes/squad/:id/get.md)

---

[API](https://skmtc.dev/vapiai/apis/vapi-api.md) · [All operations](https://skmtc.dev/vapiai/apis/vapi-api/llms.txt) · [OpenAPI document](https://skmtc-service-production.skmtc.workers.dev/v1/apis/vapiai/vapi-api/revisions/872ff45af53b/schema)
