↓
Rivya AI Docs

Rivya Audio Uploads and Recording Requirements

Use recordings with supported Rivya video and reference workflows. Check audio formats, duration, companion files, and why recording cleanup is currently paused.

Last reviewed on September 9, 2026

Choose the result you need before uploading a recording. An audio file can be an input to a video model; it does not have to go through Audio Studio. A file picker or a model description does not establish that the selected operation is currently runnable. Check model availability first.

Choose the Right Use for the Recording

What you needWhere to startWhat the file does
A portrait animated with recorded speechKling AI Avatar Standard, Kling AI Avatar Pro, or OmniHuman 1.5One image and one audio file produce a video; this is not a cleaned audio export
An existing video matched to supplied audioVolcengine Video Lip SyncOne video and one audio file are inputs to a new video result
Audio as a reference for video generationSeedance 2.0 or Seedance 2.0 Fast, in reference modeThe recording guides a video request; it is not a promise to repair or reproduce the source exactly
New speech from a script, dialogue, music, or soundsAudio Studio and audio workflowsStart with text and the selected model's controls; these uses do not require uploading the old recording
Noise cleanup or speech isolation from an existing recordingSee the current audio-cleanup explanationElevenLabs Audio Isolation is paused; its documented input is not an available cleanup service

Text-to-speech, dialogue, and Suno Sounds generate new audio. They do not repair an uploaded recording. The current chat attachment picker accepts supported images, not audio files; do not upload a recording there expecting transcription or a listening review. These routes also do not establish a general voice-cloning service.

Match the File to the Model

Use the selected form and reference-file requirements for the full input contract. The following are specific examples, not universal audio limits:

Model or modeAudio requirementsOther inputs and limits
Seedance 2.0 / 2.0 Fast reference modeMP3 or WAV, up to 15 MB per file; 2–15 seconds per fileUp to 3 audio references, with a combined audio duration of no more than 15 seconds; check separate image and video requirements
Kling AI Avatar Standard / ProMP3, M4A, WAV, AAC, or OGG; Rivya's current upload limit is 95 MB; duration must be greater than 0 and no more than 300 secondsExactly one JPEG or PNG image, up to 10 MB, and one audio file
OmniHuman 1.5MP3, M4A, WAV, AAC, or OGG, up to 10 MB; duration must be greater than 0 and strictly less than 60 secondsExactly one JPEG, PNG, or WebP image up to 10 MB, and one audio file
Volcengine Video Lip SyncMP3, M4A, WAV, AAC, or OGG, up to 10 MB, with a readable positive durationExactly one supported video and one audio file; video properties and Lite/Basic settings have their own requirements

For example, two 8-second audio references exceed Seedance 2.0's 15-second combined limit even though each file passes its individual duration limit. Choose shorter relevant excerpts and re-export them before uploading. Do not apply that limit to a different Seedance version or another model without checking its own form.

Open the exported copy locally. Renaming a file extension does not convert its format, and a locally playable file can still fail a model's size or duration checks. Use browser and file troubleshooting if playback or upload confirmation fails. For OmniHuman, also check the model page's prompt-language and length requirements; they are separate from the recording's language.

Prepare a Copy You Can Safely Send

Keep the original on your device. In a separate copy, trim unrelated conversation and private material, verify that the beginning and end of the required speech remain intact, and listen for clipped words, distortion, or accidental silence. A smaller excerpt should still contain the complete passage needed for the chosen task.

Confirm your right to use the recording, the speaker's voice, and any accompanying image or video for the intended output. Follow the safe-upload guide before uploading. Uploading sends the file beyond your browser before you generate; removing the form attachment is not confirmation of deletion from other systems. Use only the material the request needs.

Upload and Send One Request

  1. Sign in and open the chosen model and operation. Save unsent text, model settings, and source files locally if sign-in or a page change is needed; see login and draft protection.

  2. Select the prepared audio and any required image or video. In Seedance 2.0 reference mode, use the audio-reference area. Wait for upload completion, then verify the displayed files, roles, counts, and durations where shown.

  3. Review the selected mode, relevant settings, and estimated credits before submitting once. Upload completion does not start generation. A mode or model change can reset settings or remove incompatible references, so recheck before sending.

  4. Check the existing task in History if the response or connection is unclear. Do not immediately repeat the request. Follow queue and waiting guidance for an in-progress task and failed tasks and credit returns for a recorded failure.

Preparing or uploading the file does not guarantee acceptance, successful generation, or a particular quality. A second generation request can have a separate credit cost.

Review the Result for Its Actual Purpose

For avatar or lip-sync video, watch the whole result while listening: check missing words, mouth timing, visible facial changes, unintended content, and the start and end of the speech. For audio-guided video, check whether the result uses the intended rhythm or sound direction; do not assume it preserves every detail of the recording.

Compare with the original before replacing or sharing anything. Save the usable result according to output downloads, and retain the original separately. An expired result link or a preview that will not play is not proof that the generation failed or qualifies for a credit return. Check the actual file first.

Resolve Rejections Without Repeating the Same Mistake

A missing audio field can mean the selected operation does not accept recordings. A paused cleanup model cannot be made available by paying for a plan or uploading a different file. Choose an appropriate available operation only if its output meets your real need.

If a file is rejected, read the message and check its actual format, size, duration, companion files, and the mode's count or combined-duration limits. Fix the relevant issue in a separate copy and inspect it locally before another upload. If a task was already submitted, check its record before starting a replacement.

If you still need help, use Contact. Include the model, mode, file format, size, duration, visible error, approximate time, and task identifier if one exists. Upload problems can occur before a task identifier is created. Do not send passwords, API keys, private access links, or the sensitive recording itself in the initial message.