Speed-optimized audiovisual generation from text, an image, or a 2-20 second audio clip. Creates synchronized video and audio in one pass, with output up to 4K and optional start/end-frame control.
Added Aug 11, 2026
Starting Price
From $0.200 per video
Final price depends on the selected settings and is shown before generation.
Model Type
both
Settings
Generation controls available for this model.
Output Format
Default Duration
6
8 duration options
Aspect Ratio
Default
auto
Options (3)
Auto, Landscape (16:9), Portrait (9:16)
Auto follows an input image; text-only generation defaults to 16:9.
Camera Motion
Default
N/A
Options (9)
Automatic, Dolly In, Dolly Out, Dolly Left +5 more
Optional camera movement.
Duration
Default
6
Options (8)
6 seconds, 8 seconds, 10 seconds, 12 seconds +4 more
Clip length for text-to-video and image-to-video. Audio-to-video follows the input audio length.
Frames Per Second
Default
25
Options (4)
24 FPS, 25 FPS, 48 FPS, 50 FPS
Output frame rate for text-to-video and image-to-video.
Generate Audio
Default
Yes
Generate synchronized audio with text-to-video or image-to-video.
Guidance Scale
Default
5
Prompt adherence for audio-to-video (1-50).
Resolution
Default
1080p
Options (4)
720p, 1080p, 1440p, 4K
Output video resolution. Audio-to-video uses 1080p.
Benchmarks
Benchmarks
Human preference benchmarks sourced from Artificial Analysis.
Text to Video
#21 / 31
ELO
941.0
Appearances
2,935
95% CI
-13/13
Image to Video
#43 / 76
ELO
1214.0
Appearances
2,350
95% CI
-12/12
Release Date 2026-08 · Matched as LTX-2.5 Fast
Artificial Analysis APIExamples
Loading examples…
Related video models
Compare LTX-2.5 Fast with similar models from the same provider or model family.
LTX-2.5 Pro
lightricks/ltx-2.5/proHigh-fidelity audiovisual generation from text, an image, or a 2-20 second audio clip. Creates polished synchronized video and audio in one pass, with 720p/1080p output and optional start/end-frame control.
LTX-2.3 Quality
ltx-2.3-qualityLTX-2.3 Quality routes text, image, audio, reference video, extend-video, video-to-HDR, and optional LoRA inputs to the appropriate generation mode. Supports output sizes such as landscape, portrait, square, and auto with native synchronized audio.
MiniMax H3 Max Lip Sync
minimax/h3-max/lip-sync/image-to-videoAnimate a portrait or character image from supplied audio with transcription-guided lip sync. Audio must be at least 5 seconds; longer inputs are clipped to the first 15 seconds. Supports talking and singing clips from 480p through 2K.
FLUX.3 Edit Video
blackforestlabs/flux-3/edit-videoEdit existing footage with natural-language instructions. Change objects, weather, or the look of a scene while preserving motion, timing, and framing. Requires an MP4 under 15 seconds and 50 MB; output is 720p.
MiniMax H3 Max Multi Angle
minimax/h3-max/multi-angle/image-to-videoAnimate a starting image with precise camera control: orbit around the subject, move closer, pull back, or rise above the scene. Choose a camera movement or supply custom keyframes. Supports 5–15 seconds at 480p, 768p, or refined 1080p.
FLUX.3
flux-3Generate up to 20-second videos with native audio from a prompt, a start image, start/end frames, multiple keyframes, or a source clip. FLUX.3 chooses the matching workflow automatically from what you attach.