Expressive image-to-video avatar generation with natural facial performance, realistic body motion, accurate A/V sync, and optional driving audio for lip-sync mode.
Added Apr 3, 2026
Starting Price
From $0.250 per video
Final price depends on the selected settings and is shown before generation.
Model Type
image-to-video
Settings
Generation controls available for this model.
Output Format
Default Duration
5
30 duration options
Driving Audio URL
Default
N/A
Optional audio URL for lip-sync mode. If omitted, audio is generated from the prompt.
Duration
Default
5
Options (30)
1 second, 2 seconds, 3 seconds, 4 seconds +26 more
Length of the generated video in seconds.
Guidance Scale
Default
5
Optional classifier-free guidance scale (0-20).
Inference Steps
Default
8
Optional denoising steps (1-50).
Resolution
Default
256p
Options (4)
256p, 540p, 720p, 1080p
Output resolution.
Safety Checker
Default
Yes
Run prompt and image safety checks before generation.
Seed
Optional seed for reproducible output.
Benchmarks
Benchmarks
No benchmark data is available yet for this model.
Examples
Loading examples…
Related video models
Compare DaVinci MagiHuman with similar models from the same provider or model family.
MiniMax H3 Max Lip Sync
minimax/h3-max/lip-sync/image-to-videoAnimate a portrait or character image from supplied audio with transcription-guided lip sync. Audio must be at least 5 seconds; longer inputs are clipped to the first 15 seconds. Supports talking and singing clips from 480p through 2K.
FLUX.3 Edit Video
blackforestlabs/flux-3/edit-videoEdit existing footage with natural-language instructions. Change objects, weather, or the look of a scene while preserving motion, timing, and framing. Requires an MP4 under 15 seconds and 50 MB; output is 720p.
MiniMax H3 Max Multi Angle
minimax/h3-max/multi-angle/image-to-videoAnimate a starting image with precise camera control: orbit around the subject, move closer, pull back, or rise above the scene. Choose a camera movement or supply custom keyframes. Supports 5–15 seconds at 480p, 768p, or refined 1080p.
LTX-2.5 Fast
lightricks/ltx-2.5/fastSpeed-optimized audiovisual generation from text, an image, or a 2-20 second audio clip. Creates synchronized video and audio in one pass, with output up to 4K and optional start/end-frame control.
LTX-2.5 Pro
lightricks/ltx-2.5/proHigh-fidelity audiovisual generation from text, an image, or a 2-20 second audio clip. Creates polished synchronized video and audio in one pass, with 720p/1080p output and optional start/end-frame control.
FLUX.3
flux-3Generate up to 20-second videos with native audio from a prompt, a start image, start/end frames, multiple keyframes, or a source clip. FLUX.3 chooses the matching workflow automatically from what you attach.