Use Volcengine Video Lip Sync when the source footage already has the right person, framing, and motion, but the mouth movement needs to follow a different voice recording.
Exactly one source video and one target audio trackLite mode for front-facing single-person footage
Basic mode for more complex single-person scenes
Input
Up to 2 reference files
Output
AI Video Generator
Credits
per_input_second · default=8 · ⌈total⌉
Best for
Redubbing one presenter clip with a replacement voice track
Check the input, output, credits, and example result, then try this model directly on the page. Move into Studio when you need saved history, assets, or longer iteration.
Example output
Volcengine Video Lip Sync
Best for
Redubbing one presenter clip with a replacement voice trackLocalizing existing single-speaker videosCorrecting dialogue while keeping approved footageFront-facing social, tutorial, or spokesperson clips
Supports
Video + Audio to Video
Online trial
Try Volcengine Video Lip Sync
Upload one supported source video and one vocal track, then choose Lite or Basic processing and only the controls available for that mode.
Results
Volcengine Video Lip Sync
VolcengineVideo + Audio to Video
Use Volcengine Video Lip Sync when the source footage already has the right person, framing, and motion, but the mouth movement needs to follow a different voice recording.
Volcengine Video Lip Sync
No result yet. Submit a prompt to start a run.
Video + Audio to Video
Decision fit
When this model is the right choice
Decision fit
Fit signals
Exactly one source video and one target audio track
Lite mode for front-facing single-person footage
Basic mode for more complex single-person scenes
Optional vocal separation and mode-specific alignment controls
Credits and output duration follow the verified audio track
Best-fit tasks
Use it when the task looks like this
Redubbing one presenter clip with a replacement voice trackLocalizing existing single-speaker videosCorrecting dialogue while keeping approved footageFront-facing social, tutorial, or spokesperson clips
Quiet facts
Inputs, output, and credits to confirm
Provider
Volcengine
Category
Video
Capabilities
Video + Audio to Video
Credit model
per_input_second · default=8 · ⌈total⌉
Input path
Up to 2 reference files
Prompt setup
Prompt-guided
FAQ
Volcengine Video Lip Sync FAQ
What is Volcengine Video Lip Sync best used for?
Use it when the source video is already the visual result you want to keep and only the spoken performance needs to follow a different vocal track. It is a video-and-audio lip-sync workflow, not a prompt-led video generator or portrait-animation model.
What files can I upload on Rivya?
Upload exactly one MP4 or MOV video up to 95 MB and one MP3, M4A, WAV, AAC, or OGG audio file up to 10 MB. The video must be within the supported 360p to 1080p geometry, 24 to 60 fps, and 1 to 30 Mbps range, and both files must have a positive verified duration.
Should I choose Lite or Basic mode?
Choose Lite for a single person facing the camera when faster processing and audio-to-video alignment controls are useful. Choose Basic for a more complex single-person scene, especially when scene segmentation and speaker identification are needed.
How do the Lite-only controls work?
In Lite mode, audio alignment can loop the video when the target audio is longer. Reverse looping is available only when alignment is on, and the template start value chooses where the source sequence begins. Basic mode does not use those controls.
How are credits calculated?
Rivya uses a rate of 8 credits per verified audio second, and the output duration follows that audio. The decimal duration remains part of the full calculation, and only the reservation total is rounded up to a whole credit. When valid actual usage no higher than the reservation is reported, Rivya settles to it; missing, invalid, or higher-than-reserved usage remains in reconciliation instead of triggering a hidden extra charge.
What can I do on this lip-sync page?
Pair one existing single-person video with one target vocal track so the visible speech follows the new audio while the source footage remains the visual base.
What should I check before uploading?
Use a supported MP4 or MOV file with a clearly visible speaker, then prepare a clean vocal track. Check that the video is no larger than 95 MB and that its resolution, frame rate, and bitrate fall within the listed limits.
Which mode should I start with?
Start with Lite for straightforward front-facing footage. Use Basic when the scene is more complex or when scene segmentation and speaker identification are useful.
Do I need to sign in, and how do credits work?
The details are public, but uploads and generation require sign-in. Pricing is 8 credits per verified audio second, the result follows the audio duration, and only the final calculated total is rounded up.
When should I open the full Studio?
Stay here for one source-video and voice-track test. Open Studio for saved dubs, repeated languages or takes, comparison work, or a longer localization workflow.
Compare alternatives
Other models to consider next
Video
Seedance 2.0
ByteDance's full Seedance 2.0 video model with explicit support for prompt-only generation, frame-driven animation, and multimodal reference generation. Rivya keeps the documented role split explicit so frame inputs and multimodal references stay mutually exclusive instead of collapsing into one ambiguous upload bucket.
Why consider it
Consider it when your next task needs Prompt, frames, or image/video/audio references.
InputPrompt, frames, or image/video/audio references
ByteDance's faster Seedance 2.0 video model with full scene routing for prompt-only generation, frame-driven image animation, and multimodal reference video generation. Rivya keeps the documented scene split explicit so first/last-frame inputs do not collide with reference image, video, and audio roles.
Why consider it
Consider it when your next task needs Prompt, frames, or image/video/audio references.
InputPrompt, frames, or image/video/audio references
Best forFast ad previs from prompts or storyboard frames
Video
Seedance 1.5 Pro
A Seedance video model for text-to-video and image-to-video with native audio-visual sync. It supports 480p–1080p, 4–12s clips, 6 aspect ratios, dynamic or fixed lens control, optional audio generation, and lip-sync.
Why consider it
Consider it when your next task needs Prompt + optional images.