Clean Audio

Submit a clean audio job

Start an asynchronous job that removes background noise from a speech recording — the same model that powers Clean Audio in the VEED editor.

It is built for speech, in any language. Music and sound effects count as noise and are removed. A video file is accepted: its audio track is cleaned, and the result is audio, not video.

Inputs

  • audio_url — public URL of the recording; any file ffmpeg reads (such as MP3, WAV, M4A/AAC, OGG/Opus, FLAC, MP4, MOV or WebM), up to 30 minutes and 512 MB. Split anything longer into separate jobs
  • strength — optional; how much of the original may remain under speech, between 0 and 1. Lower keeps more room tone behind the voice
  • target_lufs — optional; the output's integrated loudness, between -40 and -8 LUFS
  • normalize_loudness — optional; false keeps the input level instead of normalizing it
  • output_format — optional; flac or wav, both 48 kHz mono 16-bit

What happens next

The job is accepted immediately — you get 202 Accepted with a job_id and status PROCESSING. Processing takes roughly 0.2–0.5× the audio's duration; poll GET /v1/clean-audio/{job_id} until the job is COMPLETED or FAILED.

post/v1/clean-audio

Headers

X-Veed-Store-IO'0' | '1'

Set to 0 to not store this request's and response's bodies. They are then not available in your request logs for debugging.

X-Veed-Media-Expiration-Secondsinteger

Number of seconds before the media URLs returned for this request expire. A value above the maximum is capped rather than rejected.

Request body

$schemastring uri

A URL to the JSON Schema for this object.

audio_urlstring uri required

URL of the recording to clean: any audio or video file, up to 30 minutes and 512 MB. A video's audio track is used; multi-channel audio is mixed down to mono.

normalize_loudnessboolean

Set to false to skip loudness normalization and keep the input level.

output_format'flac' | 'wav'

Container for the 48 kHz mono 16-bit output. FLAC is lossless at about half the size of WAV.

strengthnumber double

How much of the original is allowed to remain under speech: the suppression floor is 1 - strength. Lower keeps more room tone behind the voice; silence between words is always fully cleaned.

target_lufsnumber double

Integrated loudness of the output in LUFS (ITU-R BS.1770); true peak is capped at -1.1 dBTP. Ignored when normalize_loudness is false. A null reads as omitted: set normalize_loudness to false to skip normalization.

Example request

{
  "$schema": "https://api.veed.io/schemas/CleanAudioInput.json",
  "audio_url": "https://static-assets.veed.io/api-examples/clean-audio-input.m4a",
  "output_format": "flac"
}

Response

Accepted

$schemastring uri

A URL to the JSON Schema for this object.

Example response

{
  "$schema": "https://api.veed.io/schemas/ResourceCleanAudioJob.json",
  "data": {
    "result": {
      "audio": {
        "content_type": "video/mp4"
      }
    }
  }
}

Changes

Changed in 2 of the 4 revisions of this API.11