Open image-to-video world model for high-quality physical motion, scene dynamics, and stylized transformations from a first-frame image.
Added Jun 2, 2026
Starting Price
From $0.050 per video
Final price depends on the selected settings and is shown before generation.
Model Type
image-to-video
Settings
Generation controls available for this model.
Output Format
Default Duration
7
7 duration options
Aspect Ratio
Default
16:9
Options (3)
Landscape (16:9), Portrait (9:16), Square (1:1)
Output aspect ratio
Duration
Default
7
Options (7)
1 second, 2 seconds, 3 seconds, 4 seconds +3 more
Length of the generated video in seconds
Prompt Expansion
Default
Yes
Use the model reasoner to expand the prompt from the first frame
Resolution
Default
480p
Options (3)
256p, 480p, 720p
Output video resolution tier
Safety Checker
Default
Yes
Run prompt and input-image safety checks
Benchmarks
Benchmarks
Human preference benchmarks sourced from Artificial Analysis.
Image to Video
#30 / 76
ELO
1253.0
Appearances
3,029
95% CI
-11/11
Release Date 2026-05 · Matched as Cosmos3-Super-Image2Video
Artificial Analysis APIExamples
Loading examples…
Related video models
Compare Cosmos 3 Super Image-to-Video with similar models from the same provider or model family.
MiniMax H3 Max Lip Sync
minimax/h3-max/lip-sync/image-to-videoAnimate a portrait or character image from supplied audio with transcription-guided lip sync. Audio must be at least 5 seconds; longer inputs are clipped to the first 15 seconds. Supports talking and singing clips from 480p through 2K.
Wan 3.0 Prime Video Edit
alibaba/wan-3.0-prime/video-editAccelerated Wan 3.0 editing for an existing video with optional image or audio references. Uses the first 15 seconds of the source and supports 2–15 second output at 480p, 720p, or 1080p.
Wan 3.0 Prime Video Extend
alibaba/wan-3.0-prime/video-extendAccelerated video extension that appends 2–30 seconds with optional target last-frame guidance. Source audio is preserved, and clips longer than 120 seconds use their final 120 seconds as context.
Wan 3.0 Video Edit
alibaba/wan-3.0/video-editEdit an existing video with natural-language instructions and optional image or audio references. Uses the first 15 seconds of the source and supports 2–15 second output at 480p, 720p, or 1080p.
Wan 3.0 Video Extend
alibaba/wan-3.0/video-extendExtend an existing video by 2–30 seconds with optional target last-frame guidance. Source audio is preserved, and clips longer than 120 seconds use their final 120 seconds as context.
Seedance 2.5 Talking Avatar
bytedance/seedance-2.5/talking-avatarTurn a portrait and voice recording into an expressive talking-avatar video with natural lip sync, facial motion, and optional direction for gestures or presentation style. Supports up to 120 seconds at 480p or 720p.