
Vidu Q2 · Image Generation / Referencia a Imagen
Next-gen image-to-video model with "emotional acting" for nuanced facial expressions. Supports 2–8s clips at 1080p with cinematic camera control. Features multi-modal reference (2 videos + 4 images) enabling precise replication of expressions, actions, textures, and effects. Includes video editing: element add/delete/replace, style transfer, aspect adjustment. Lightning mode outputs 5s in 20s. Supports up to 7 consistent subjects, 3x faster than Q1
Upload Images
JPG, JPEG, PNG (Max 10MB)
Image Playground Ready
Describa lo que desea generar en el panel de prompt y ajuste las opciones a la izquierda.
Vidu Q2Reference-to-Image at 1080p–4K
Vidu Q2 generates stills from text or references at 1080p, 2K, or 4K with nine aspect ratios and optional seed control.

At a Glance
Vidu Q2 Capabilities
Reference-aware stills with multi-tier fidelity.
Reference Conditioning
Feed reference images so identity and style stay locked across generated stills.

1080p / 2K / 4K Tiers
Draft at 1080p, ship at 2K, and promote print or OOH heroes to 4K.

Nine Aspect Ratios
16:9, 9:16, 1:1, 3:4, 4:3, 21:9, 2:3, 3:2, and auto cover social, print, and cinematic placements.

Seed Reproducibility
Pin a seed to regenerate the same creative baseline after brief revisions.

How It Works
From brief to reference-aware still.
Write the Brief
Describe subject, style, and use-case.
Optional References
Add reference images to lock identity and style.
Pick Fidelity & Ratio
1080p/2K/4K; choose from nine ratios or auto.
Pin Seed if Needed
Lock a seed for reproducible campaign frames.
Vidu Q2 Domains
Where reference stills and multi-tier fidelity matter.
Character Campaigns
Lock talent identity across still variants.
Social Carousels
Multi-ratio frames with consistent style.
Print Masters
4K output for large-format placements.
Mood Boards
Reference-conditioned exploration boards.
Prompt Tips
Better Vidu Q2 stills on the first pass.
Reference images work best when each has a single clear job.
Draft at 1080p and promote final masters to 4K.
Reuse a seed so small brief edits do not re-roll the entire composition.
Vidu Q2 Quickstart
Reference-aware still generation.
curl -X POST "https://api.powertokens.ai/v1/images/generations" \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type": "application/json" \
-d '{
"model": "viduq2",
"prompt": "Editorial close-up portrait with sculptural hat, photorealistic studio light.",
"size": "2048",
"response_format": "url",
"watermark": false,
"aspect_ratio": "16:9"
}'Technical Specifications
Confirmed parameters and runtime execution protocols.
Detalles de precios
La facturación real de este modelo se calcula dinámicamente en función de los parámetros específicos de su solicitud API. A continuación se muestran las combinaciones específicas y sus precios correspondientes:
| Modalidad | Créditos | Precio (USD) |
|---|---|---|
| Text to Image/1080P | 29/ Image | $0.029 |
| Text to Image/2K | 38/ Image | $0.038 |
| Text to Image/4K | 48/ Image | $0.048 |
| Reference to Image/1080P/1-3 Images | 38/ Image | $0.038 |
| Reference to Image/1080P/4-7 Images | 48/ Image | $0.048 |
| Reference to Image/2K/1-3 Images | 57/ Image | $0.057 |
| Reference to Image/2K/4-7 Images | 76/ Image | $0.076 |
| Reference to Image/4K/1-3 Images | 95/ Image | $0.095 |
| Reference to Image/4K/4-7 Images | 143/ Image | $0.143 |
Recommended Related Models
Explore complementary video and multimodal models with your unified API key.


Seedream 5.0
seedream-5-0-260128
Supports text, single-image and multi-image inputs, and enables the generation of image sets


Wan 2.7 Image Pro
wan2.7-image-pro
Wan2.7–image-pro,supports text to image, text/image to sequential images, image editing, multi-image reference generation, and interactive editing. Delivers enhanced performance in text rendering, subject consistency, and complex instruction following.


Kling Image O1
kling-image-o1
Advanced image generation model from the Kling O1 family. Supports text-to-image and image-to-image workflows with deep semantic understanding. Features character/subject consistency across generations, multi-reference input (up to 7 images), and intelligent aspect ratio adaptation. Professional-grade output suitable for commercial creative pipelines


Kling V2
kling-v2
V2.0 foundational image model supporting text-to-image and image-to-image generation. Enables multi-type reference control—character, face, subject, scene, and style. Built for role consistency and thematic control across generated visuals. Multiple aspect ratio support
Frequently Asked Questions
Everything you need to know before integrating this model.
Start Building with Vidu Q2 Today
Create an account in seconds to receive 100 free credits and start generating immediately. No credit card or upfront contract required.