Vidu Q2 is a high-end text-to-image model with cinematic lighting and clean composition. Supports up to 4K resolution and flexible aspect ratios. Upload reference images to guide generation with subject/composition consistency.
Added Dec 2, 2025
Approx. Price
$0.030 per image
Model Type
both
Settings
Generation controls available for this model.
Images Per Run
Up to 1
Output images
Input Images
Up to 7
Reference/edit images accepted • Route max 30 MB
Output Sizes
Aspect Ratio
Default
1:1
Options (9)
Auto (match references), Square, 16:9 Widescreen, 9:16 Vertical +5 more
Canvas shape for generation. Use "auto" to match reference images.
Number of Images
Default
1
Resolution
Default
1080p
Options (3)
1080p (Fast preview (1920x1080)), 2K (Higher detail (2560x1440)), 4K (Maximum sharpness (3840x2160))
Seed
Default
-1
Control reproducibility (-1 for random).
Benchmarks
Benchmarks
Human preference benchmarks sourced from Artificial Analysis.
Text to Image
#88 / 166
ELO
920.0
Appearances
6,451
95% CI
-7/7
Image Editing
#57 / 80
ELO
966.0
Appearances
7,152
95% CI
-7/7
Release Date 2025-11 · Matched as Vidu Q2
Artificial Analysis APIExamples
Loading examples…
Related image models
Compare Vidu Q2 with similar models from the same provider or model family.
Vidu Q2 Reference
vidu-q2-referenceVidu Q2 Reference-to-Image generates images based on 1-7 reference images with customizable prompts. Ideal for keeping product, character, or actor identity consistent across shots.
Qwen Image 2.1 Edit
wavespeed-ai/qwen-image-2.1/editEdit from up to ten reference images with natural-language instructions, strong subject preservation, flexible framing, and native output up to 2K.
Qwen Image 2.1 Edit LoRA
wavespeed-ai/qwen-image-2.1/edit-loraEdit from up to ten references while applying up to three custom LoRAs for precise identity, product, or style control at up to 2K.
Qwen Image 2.1
wavespeed-ai/qwen-image-2.1/text-to-imageCreate high-quality images with strong prompt following, multilingual text rendering, flexible framing, and native output up to 2K.
Qwen Image 2.1 LoRA
wavespeed-ai/qwen-image-2.1/text-to-image-loraGenerate Qwen Image 2.1 artwork with up to three custom character, product, or style LoRAs and native output up to 2K.
Bria Product Holding
bria/product-holdingPlace one to three referenced products naturally into a person's hands while preserving the subject and scene. Built for e-commerce, lifestyle advertising, influencer content, and product campaigns.