Use HappyHorse 1.1 when the same video idea needs to move between text, a first-frame image, and a set of visual references without leaving one focused workflow.
Three distinct starting points: text only, one first-frame image, or 2 to 9 reference imagesSelectable 720p and 1080p output
Check the input, output, credits, and example result, then try this model directly on the page. Move into Studio when you need saved history, assets, or longer iteration.
Example output
HappyHorse 1.1
Best for
Comparing a written scene with image-led versions of the same conceptAnimating one approved key visual into a short clipGuiding a video with several character, product, wardrobe, or style referencesCreating landscape, portrait, square, or extra-wide campaign variations
Supports
Text to VideoImage to VideoReference to Video
Online trial
Try HappyHorse 1.1
Start with text, one first-frame image, or 2 to 9 visual references, then choose the output resolution, duration, and supported frame shape.
Results
HappyHorse 1.1
HappyHorseText to Video
Use HappyHorse 1.1 when the same video idea needs to move between text, a first-frame image, and a set of visual references without leaving one focused workflow.
HappyHorse 1.1
No result yet. Submit a prompt to start a run.
Text to VideoImage to VideoReference to Video
Decision fit
When this model is the right choice
Decision fit
Fit signals
Three distinct starting points: text only, one first-frame image, or 2 to 9 reference images
Selectable 720p and 1080p output
Integer output durations from 3 to 15 seconds
Nine frame shapes for prompt-only and multi-reference runs
Per-second pricing stays explicit before the task is reserved
Best-fit tasks
Use it when the task looks like this
Comparing a written scene with image-led versions of the same conceptAnimating one approved key visual into a short clipGuiding a video with several character, product, wardrobe, or style referencesCreating landscape, portrait, square, or extra-wide campaign variations
Quiet facts
Inputs, output, and credits to confirm
Provider
HappyHorse
Category
Video
Capabilities
Text to Video · Image to Video · Reference to Video
HappyHorse 1.1 is useful when a short video concept may begin as text, one first-frame image, or a collection of reference images. It keeps those three workflows on one model page while preserving separate input rules for each mode.
How does HappyHorse 1.1 choose the video workflow?
Leave the reference area empty for text-to-video, upload exactly one image for first-frame image-to-video, or upload 2 to 9 images for reference-to-video. Video and audio inputs are not accepted on this page.
What image limits apply?
Use JPEG, PNG, or WebP images no larger than 20 MB each. A single first-frame image must have both sides at least 300 pixels and an aspect ratio between 1:2.5 and 2.5:1. Multi-image references must have a shortest side of at least 400 pixels. Prompts allow up to 2,500 characters when they contain Chinese and up to 5,000 otherwise.
How are HappyHorse 1.1 credits calculated?
The rate is 22.5 credits per requested output second at 720p and 29 credits per requested output second at 1080p. Rivya keeps the decimal rate through the full duration calculation and rounds only the reservation total up to a whole credit. When valid actual usage no higher than the reservation is reported, Rivya settles to it; missing, invalid, or higher-than-reserved usage remains in reconciliation instead of triggering a hidden extra charge.
When does the aspect-ratio control apply?
Choose an aspect ratio for prompt-only and multi-image reference runs. A single-image run follows the first frame instead, so the separate aspect-ratio control is not used in that mode.
What can I create on this HappyHorse 1.1 page?
Create a short video from a prompt, animate one first-frame image, or guide a new clip with 2 to 9 reference images. All three paths support 720p or 1080p output and a 3 to 15 second duration.
How many images should I upload?
Upload no images for text-to-video, exactly one image for first-frame animation, or 2 to 9 images for multi-reference generation. The number of images determines which workflow is used.
Which settings should I choose first?
Start with the input mode, then choose resolution and duration. For prompt-only or multi-reference work, also select the frame shape that matches the intended placement. Prompts allow up to 2,500 characters when they contain Chinese and up to 5,000 otherwise.
Do I need to sign in, and how do credits work?
The details are public, but uploading media and generating require sign-in. Pricing is 22.5 credits per output second at 720p or 29 at 1080p, with only the final total rounded up.
When should I open the full Studio?
Stay here for a focused first clip. Open Studio when you need saved versions, repeated reference combinations, or a longer production thread around the same video idea.
Compare alternatives
Other models to consider next
Video
Seedance 2.0
ByteDance's full Seedance 2.0 video model with explicit support for prompt-only generation, frame-driven animation, and multimodal reference generation. Rivya keeps the documented role split explicit so frame inputs and multimodal references stay mutually exclusive instead of collapsing into one ambiguous upload bucket.
Why consider it
Consider it when your next task needs Prompt, frames, or image/video/audio references.
InputPrompt, frames, or image/video/audio references
ByteDance's faster Seedance 2.0 video model with full scene routing for prompt-only generation, frame-driven image animation, and multimodal reference video generation. Rivya keeps the documented scene split explicit so first/last-frame inputs do not collide with reference image, video, and audio roles.
Why consider it
Consider it when your next task needs Prompt, frames, or image/video/audio references.
InputPrompt, frames, or image/video/audio references
Best forFast ad previs from prompts or storyboard frames
Video
Seedance 1.5 Pro
A Seedance video model for text-to-video and image-to-video with native audio-visual sync. It supports 480p–1080p, 4–12s clips, 6 aspect ratios, dynamic or fixed lens control, optional audio generation, and lip-sync.
Why consider it
Consider it when the credit note fits your next run: per_output_second · 480p_without_audio=1.75 · 480p_with_audio=3.5 · 720p_without_audio=3.5 · 720p_with_audio=7 · 1080p_without_audio=7.5 · 1080p_with_audio=15 · ⌈total⌉.
Teams that want output duration and resolution to remain explicit before generation
Developer access
Partly available via API
HappyHorse 1.1 has callable Public API modes, but reference-media modes still need Files API upload support. Check the model reference before submitting.