Create a 4-to-15-second MiniMax H3 video from text, one or two guide frames, or a verified multimodal bundle, with a dedicated 2K tier and explicit reference pricing.
A dedicated 2K tier alongside 768P outputSeparate text, frame-guided, and multimodal reference pathsReference bundles with up to nine images, three videos, and three audio files
Check the input, output, credits, and example result, then try this model directly on the page. Move into Studio when you need saved history, assets, or longer iteration.
Example output
MiniMax H3
Best for
2K short-form video concepts built from a written shot briefFirst-frame or first-and-last-frame animation with one clear subjectMotion references that combine visual direction with supporting audioReference-heavy scenes where image-count pricing needs to stay visible
Supports
Text to VideoImage to VideoReference to Video
Online trial
Create with MiniMax H3
Select text, frame guidance, or multimodal reference input, then choose 768P or 2K and a 4-to-15-second duration.
Results
MiniMax H3
MiniMaxText to Video
Create a 4-to-15-second MiniMax H3 video from text, one or two guide frames, or a verified multimodal bundle, with a dedicated 2K tier and explicit reference pricing.
MiniMax H3
No result yet. Submit a prompt to start a run.
Text to VideoImage to VideoReference to Video
Decision fit
When this model is the right choice
Decision fit
Fit signals
A dedicated 2K tier alongside 768P output
Separate text, frame-guided, and multimodal reference paths
Reference bundles with up to nine images, three videos, and three audio files
Whole-number output durations from 4 to 15 seconds
A transparent formula for input-video time and images beyond the first five
Best-fit tasks
Use it when the task looks like this
2K short-form video concepts built from a written shot briefFirst-frame or first-and-last-frame animation with one clear subjectMotion references that combine visual direction with supporting audioReference-heavy scenes where image-count pricing needs to stay visible
Quiet facts
Inputs, output, and credits to confirm
Provider
MiniMax
Category
Video
Capabilities
Text to Video · Image to Video · Reference to Video
Use MiniMax H3 for 4-to-15-second video work that needs either prompt-only framing, one or two guide frames, or a multimodal reference bundle. Its distinguishing controls are the 768P and 2K tiers plus explicit accounting for input-video duration and more than five images.
How do the MiniMax H3 scenes differ?
All three scenes use a non-empty prompt of up to 7,000 characters. Text accepts no uploads and requires a specific non-adaptive aspect ratio. Frames accepts exactly one image as either the starting or ending frame, or two ordered images as start and end, with no video or audio. Reference requires at least one upload; audio cannot be used by itself and must accompany an image or video.
What media can I use in a MiniMax H3 reference request?
Rivya accepts up to nine JPEG, PNG, or WebP images, three MP4 or MOV videos, and three MP3 or WAV audio files, with no more than 15 files overall. Images are limited to 30 MB each. Video is limited to 50 MB and 2 to 15 seconds per file, with 15 video seconds in total; audio is limited to 15 MB and 2 to 15 seconds per file, with 15 audio seconds in total.
Which image and video geometry does MiniMax H3 accept?
Images and reference videos must each have width and height from 256 to 5,760 pixels and an aspect ratio from 0.4 to 2.5. Reference videos must also run at 23.976 to 60 frames per second.
How are MiniMax H3 credits calculated?
The rate is 16 credits at 768P or 26 at 2K, multiplied by requested output seconds plus verified input-video seconds. The first five images add no image charge; images six through nine add 8 credits each. Input audio does not add to this formula. Rivya rounds the combined total up once.
Which aspect ratios are available in MiniMax H3?
Text-to-video requires 21:9, 16:9, 4:3, 1:1, 3:4, or 9:16 and does not accept adaptive. Reference-to-video can also use adaptive. The frame-guided path derives its framing from the supplied image or image pair rather than sending a separate aspect-ratio value.
How is the final MiniMax H3 charge settled?
Rivya reserves the calculated formula before generation, preserving all rate terms and rounding only the full total. When valid actual usage at or below the reservation is reported, Rivya settles to it. Missing, invalid, or higher-than-reserved usage remains in reconciliation instead of triggering a hidden extra charge.
What can I set up on this MiniMax H3 page?
Choose text, frame guidance, or a multimodal reference scene; set 768P or 2K; choose a whole-number duration from 4 to 15 seconds; and use an aspect ratio only where that scene accepts it.
Can I upload audio by itself?
No. In the reference scene, audio must be accompanied by at least one image or video. The text and frames scenes do not accept audio.
Can I use an end frame without a starting frame?
Yes. With one image, choose whether it is the starting or ending frame. With two images, order them as start and then end.
When do images add to the MiniMax H3 price?
The first five images are included in the base time formula. Each additional image from six through nine adds 8 credits to the combined estimate before the total is rounded up once.
Which MiniMax H3 resolution should I start with?
Use 768P at 16 credits per billed second while testing motion and reference roles. Choose 2K at 26 credits per billed second once the scene direction is ready for the higher tier.
When should I continue in Studio?
Use this page to validate one scene, duration, and reference hierarchy. Move to Studio for saved 768P-versus-2K comparisons, repeated reference revisions, or a longer project thread.
Compare alternatives
Other models to consider next
Video
Seedance 2.0
ByteDance's full Seedance 2.0 video model with explicit support for prompt-only generation, frame-driven animation, and multimodal reference generation. Rivya keeps the documented role split explicit so frame inputs and multimodal references stay mutually exclusive instead of collapsing into one ambiguous upload bucket.
Why consider it
Consider it when your next task needs Prompt, frames, or image/video/audio references.
InputPrompt, frames, or image/video/audio references
ByteDance's faster Seedance 2.0 video model with full scene routing for prompt-only generation, frame-driven image animation, and multimodal reference video generation. Rivya keeps the documented scene split explicit so first/last-frame inputs do not collide with reference image, video, and audio roles.
Why consider it
Consider it when your next task needs Prompt, frames, or image/video/audio references.
InputPrompt, frames, or image/video/audio references
Best forFast ad previs from prompts or storyboard frames
Video
Seedance 1.5 Pro
A Seedance video model for text-to-video and image-to-video with native audio-visual sync. It supports 480p–1080p, 4–12s clips, 6 aspect ratios, dynamic or fixed lens control, optional audio generation, and lip-sync.
Why consider it
Consider it when your next task needs Prompt + optional images.