Use OmniHuman 1.5 when a person, pet, or illustrated character should perform from a finished voice track while the source portrait remains the visual anchor. Subject identification, subject detection, and mask-based selection are not available on this page.
Exactly one portrait image and one audio track keep the workflow focusedWorks with people, pets, and illustrated or anime-style subjects
Check the input, output, credits, and example result, then try this model directly on the page. Move into Studio when you need saved history, assets, or longer iteration.
Example output
OmniHuman 1.5
Best for
Presenter videos driven by a finished voice trackShort character performances from one approved portraitPet or illustrated-character speaking clipsNarration-led social videos and explainers
Supports
Image + Audio to Video
Online trial
Try OmniHuman 1.5
Upload one portrait and one audio clip under 60 seconds, then choose the output resolution, processing mode, and optional seed.
Results
OmniHuman 1.5
ByteDanceImage + Audio to Video
Use OmniHuman 1.5 when a person, pet, or illustrated character should perform from a finished voice track while the source portrait remains the visual anchor. Subject identification, subject detection, and mask-based selection are not available on this page.
OmniHuman 1.5
No result yet. Submit a prompt to start a run.
Image + Audio to Video
Decision fit
When this model is the right choice
Decision fit
Fit signals
Exactly one portrait image and one audio track keep the workflow focused
Works with people, pets, and illustrated or anime-style subjects
Selectable 720p and 1080p output
Optional faster processing mode
Credits follow the verified audio duration
Best-fit tasks
Use it when the task looks like this
Presenter videos driven by a finished voice trackShort character performances from one approved portraitPet or illustrated-character speaking clipsNarration-led social videos and explainers
Quiet facts
Inputs, output, and credits to confirm
Provider
ByteDance
Category
Video
Capabilities
Image + Audio to Video
Credit model
per_input_second · default=27 · ⌈total⌉
Input path
1 image + 1 audio
Prompt setup
Up to 300 chars
FAQ
OmniHuman 1.5 FAQ
What is OmniHuman 1.5 best used for?
OmniHuman 1.5 is designed for a focused image-and-audio workflow: one portrait supplies the subject, and one voice track drives the resulting performance. It fits presenter clips, character narration, and short talking-subject videos better than a general prompt-to-video task.
Which portrait and audio files can I upload?
Upload exactly one JPEG, PNG, or WebP image up to 10 MB and one MP3, M4A, WAV, AAC, or OGG audio file up to 10 MB. The verified audio duration must be greater than zero and strictly less than 60 seconds.
How are OmniHuman 1.5 credits calculated?
Rivya uses a rate of 27 credits per verified audio second. The decimal duration remains part of the full calculation, and only the reservation total is rounded up to a whole credit. When valid actual usage no higher than the reservation is reported, Rivya settles to it; missing, invalid, or higher-than-reserved usage remains in reconciliation instead of triggering a hidden extra charge.
What does fast mode change?
Fast mode prioritizes shorter processing time with a possible quality tradeoff. Leave it off for a quality-first run, and turn it on when turnaround matters more for an early test.
Can I add a prompt to the portrait and audio?
Yes, but it is optional and limited to 300 characters on Rivya. The current prompt path accepts Chinese, English, Japanese, Korean, Spanish, or Indonesian text; the portrait and audio remain the required inputs.
What can I create on this OmniHuman 1.5 page?
Create an audio-driven character video from exactly one portrait image and one audio clip. The subject can be a person, pet, or illustrated character, and output can be 720p or 1080p.
What should I prepare before generating?
Choose a clear portrait that already represents the subject you want to keep, then prepare the final voice take. The audio must be shorter than 60 seconds, and both files must pass upload checks before generation.
Which settings matter most?
Choose 720p or 1080p, decide whether to use the faster processing option, and leave the seed at -1 for a random result unless you need a more repeatable comparison.
Do I need to sign in, and how do credits work?
The details are public, but uploading and generating require sign-in. OmniHuman 1.5 costs 27 credits per verified audio second, with only the final calculated total rounded up.
When should I open the full Studio?
Stay here to test one portrait and one voice track. Open Studio when you need saved takes, several characters, repeated voice versions, or a longer production workflow.
Compare alternatives
Other models to consider next
Video
Seedance 2.0
ByteDance's full Seedance 2.0 video model with explicit support for prompt-only generation, frame-driven animation, and multimodal reference generation. Rivya keeps the documented role split explicit so frame inputs and multimodal references stay mutually exclusive instead of collapsing into one ambiguous upload bucket.
Why consider it
Consider it when your next task needs Prompt, frames, or image/video/audio references.
InputPrompt, frames, or image/video/audio references
ByteDance's faster Seedance 2.0 video model with full scene routing for prompt-only generation, frame-driven image animation, and multimodal reference video generation. Rivya keeps the documented scene split explicit so first/last-frame inputs do not collide with reference image, video, and audio roles.
Why consider it
Consider it when your next task needs Prompt, frames, or image/video/audio references.
InputPrompt, frames, or image/video/audio references
Best forFast ad previs from prompts or storyboard frames
Video
Seedance 1.5 Pro
A Seedance video model for text-to-video and image-to-video with native audio-visual sync. It supports 480p–1080p, 4–12s clips, 6 aspect ratios, dynamic or fixed lens control, optional audio generation, and lip-sync.
Why consider it
Consider it when your next task needs Prompt + optional images.