Create Wan 3.0 video from text, guide frames, or verified multimodal references, then choose Standard or Prime, resolution, aspect ratio, duration, generated audio, and seed.
Standard and Prime share one scene and parameter contract480P, 720P, and 1080P output with six aspect-ratio choicesWhole-number output durations from 2 to 30 seconds or intelligent duration
Check the input, output, credits, and example result, then try this model directly on the page. Move into Studio when you need saved history, assets, or longer iteration.
Example output
Wan 3.0 Video
Best for
Storyboards and product motion studies that need explicit output lengthFirst-to-last-frame transformations with controlled start and finishReference-led campaign clips using visual, motion, and short sound cuesLonger text-led shots up to 30 seconds
Supports
Text to VideoImage to VideoReference to VideoVideo to Video
Online trial
Create with Wan 3.0 Video
Choose the exclusive scene first, then set tier, resolution, aspect ratio, duration, audio, seed, and any verified guide or reference media.
Results
Wan 3.0 Video
AlibabaText to Video
Create Wan 3.0 video from text, guide frames, or verified multimodal references, then choose Standard or Prime, resolution, aspect ratio, duration, generated audio, and seed.
Wan 3.0 Video
No result yet. Submit a prompt to start a run.
Text to VideoImage to VideoReference to VideoVideo to Video
Decision fit
When this model is the right choice
Decision fit
Fit signals
Standard and Prime share one scene and parameter contract
480P, 720P, and 1080P output with six aspect-ratio choices
Whole-number output durations from 2 to 30 seconds or intelligent duration
First/last-frame guidance plus signed image, video, and audio references
Exact version-and-resolution reservation with fail-closed actual-use settlement
Best-fit tasks
Use it when the task looks like this
Storyboards and product motion studies that need explicit output lengthFirst-to-last-frame transformations with controlled start and finishReference-led campaign clips using visual, motion, and short sound cuesLonger text-led shots up to 30 seconds
Quiet facts
Inputs, output, and credits to confirm
Provider
Alibaba
Category
Video
Capabilities
Text to Video · Image to Video · Reference to Video · Video to Video
How do Wan 3.0 Video's three generation scenes differ?
Text accepts no uploaded media. Frames requires a first image and optionally accepts one last image. Reference accepts a signed multimodal bundle of images, videos, and audio, but audio cannot be the only media. The scenes are mutually exclusive. A prompt is required in text mode and optional in frames and reference modes, with a trimmed maximum of 20,000 characters.
How much does Wan 3.0 Standard cost?
Standard: 8 / 16 / 32; Prime: 12.2 / 25.2 / 50.4 per billed second (480P / 720P / 1080P). The reservation is the variant-and-resolution rate multiplied by verified input-video seconds plus reserved output seconds, rounded up once. Intelligent duration reserves 30 output seconds and is unavailable with reference video.
How much does Wan 3.0 Prime cost?
Standard: 8 / 16 / 32; Prime: 12.2 / 25.2 / 50.4 per billed second (480P / 720P / 1080P). The reservation is the variant-and-resolution rate multiplied by verified input-video seconds plus reserved output seconds, rounded up once. Intelligent duration reserves 30 output seconds and is unavailable with reference video.
What are the current Wan 3.0 reference limits on Rivya?
Reference accepts up to 10 images, 5 MP4 or MOV videos, and 5 MP3 or WAV audio clips. Each video or audio clip must run from 1 to 15 seconds, and video and audio each have a separate 15-second aggregate limit. Audio cannot be submitted by itself.
Which image and video uploads can Wan 3.0 use?
Images may be JPEG, PNG without transparency, WebP, or BMP, up to 20 MB each, with width and height from 240 to 8,000 pixels and an aspect ratio no greater than 8. Videos must be MP4 or MOV, up to 95 MB each, from 240 to 4,096 pixels per side, and also remain within an 8:1 ratio boundary.
How does intelligent duration affect reservation and settlement?
The -1 choice lets the model select output length, so Rivya reserves the 30-second ceiling. Intelligent duration cannot be combined with reference video; with video input, choose an explicit duration so verified input plus output remains within 30 seconds. Valid actual usage at or below the reservation settles and refunds the difference; missing, invalid, or higher usage keeps the reservation and enters reconciliation without a hidden extra debit.
What can I configure on this Wan 3.0 Video page?
Choose text, frames, or multimodal references; select Standard or Prime; set 480P, 720P, or 1080P; choose adaptive or a fixed aspect ratio; use 2-30 seconds or intelligent duration; and control generated audio and seed.
Can I use a first and last frame together?
Yes. The frames scene requires one first image and accepts one optional last image. It rejects video and audio so the framing contract remains unambiguous.
Can a reference request contain only audio?
No. Audio can accompany image or video references, but it cannot be the only uploaded media.
Does reference video change the duration rule?
Yes. Reference video requires an explicit output duration, and verified input-video seconds plus requested output seconds must not exceed 30.
Why must references be uploaded through Rivya?
Rivya binds the user, model, URL, true MIME type, byte size, geometry, and media duration in signed metadata before task creation and credit reservation. Arbitrary external links and file shortcuts remain unavailable.
When should I continue in Studio?
Use this page to validate one scene, tier, and reference hierarchy. Continue in Studio when the project needs saved iterations, several reference bundles, or Standard-versus-Prime comparisons.
Compare alternatives
Other models to consider next
Video
Seedance 2.0
ByteDance's full Seedance 2.0 video model with explicit support for prompt-only generation, frame-driven animation, and multimodal reference generation. Rivya keeps the documented role split explicit so frame inputs and multimodal references stay mutually exclusive instead of collapsing into one ambiguous upload bucket.
Why consider it
Consider it when your next task needs Prompt, frames, or image/video/audio references.
InputPrompt, frames, or image/video/audio references
ByteDance's faster Seedance 2.0 video model with full scene routing for prompt-only generation, frame-driven image animation, and multimodal reference video generation. Rivya keeps the documented scene split explicit so first/last-frame inputs do not collide with reference image, video, and audio roles.
Why consider it
Consider it when your next task needs Prompt, frames, or image/video/audio references.
InputPrompt, frames, or image/video/audio references
Best forFast ad previs from prompts or storyboard frames
Video
Seedance 1.5 Pro
A Seedance video model for text-to-video and image-to-video with native audio-visual sync. It supports 480p–1080p, 4–12s clips, 6 aspect ratios, dynamic or fixed lens control, optional audio generation, and lip-sync.
Why consider it
Consider it when your next task needs Prompt + optional images.
Teams comparing Standard and Prime on the same reference bundle
Up to 20,000 chars
Developer access
Partly available via API
Wan 3.0 Video has callable Public API modes, but reference-media modes still need Files API upload support. Check the model reference before submitting.