Get 100 free credits on sign up to explore and build your AI applicationsClaim Free
viduq3-pro

Vidu Q3 Pro / Text to Video

Commercial
ID: viduq3-pro

Flagship video model supporting text-to-video, image-to-video, and start/end-frame workflows. Generates up to 16-second clips with native audio sync and storyboard capabilities. Available in 540p–1080p resolutions with premium motion dynamics

Model Type:
Price$0.024/ Second(24 Credits)
Input Prompt
0 / 5000
5

Upload Wm Url

JPG, JPEG, PNG (Max 10MB)

Video Generation

Video Playground Ready

Enter prompts in the left parameter panel, configure options, and click Generate.

Vidu Q3 Pro Hero
Vidu Q3 Pro • Audio-Ready Cinema

Vidu Q3 ProT2V / I2V / Dual-Frame with Native Audio

Vidu Q3 Pro generates video at 540p–1080p with optional native audio, audio type selection, start-and-end frames, and watermark position control.

540p / 720p / 1080p
Optional Native Audio
Start & End Frames
Audio Type Select
ByteDance Seed Foundation Architecture
Commercial License & Enterprise SLA

At a Glance

1080p
Max Resolution
Also 540p / 720p
Audio
Native Optional
3 audio types
3
Modes
T2V · I2V · Dual
Seed
Reproducible
Optional pin
Capability Highlights

Vidu Q3 Pro Architecture

Audio-aware multimodal video for production.

Optional Native Audio

Enable audio generation with type selection — all, speech only, or sound effects only.

audiospeech_onlysound_effect_only
Native audio

Three Conditioning Modes

Text-to-video, image-to-video, and start-and-end-frame cover most production needs.

T2VI2VDual Frame

540p to 1080p

Draft at 540p, ship social at 720p, and promote heroes to 1080p.

540p720p1080p

Watermark Position Control

Optional watermark with selectable corner positions (1–4) and custom watermark URL.

wm_positionwm_url
How It Works

How It Works

From brief to audio-ready clip.

01

Pick a Mode

Text, image, or start-and-end frame conditioning.

02

Set Resolution

540p draft, 720p social, or 1080p hero.

03

Enable Audio

Choose all, speech only, or sound effects only.

04

Watermark & Deliver

Optionally place a watermark and pull the finished clip.

Vidu Q3 Pro Domains

Where sound and picture ship together.

Content

Social with Sound

Native audio for Reels and TikTok without post Foley.

Audio
Commerce

Product Film

I2V packshots with optional speech or SFX.

I2V
Film

Transition Shots

Start-and-end frames for match cuts.

Dual Frame
Platform

Branded Watermarks

Corner position and custom watermark URL.

wm_position
Best Practices

Prompt Tips

Cleaner Vidu Q3 Pro output.

Mention acoustic mood

If audio is on, name environmental cues so sound matches picture.

Audio type is a product choice

Use speech_only for VO-led cuts and sound_effect_only for ambient plates.

Draft at 540p

Explore motion cheaply, then promote to 1080p for finals.

Developer Quickstart

Vidu Q3 Pro Quickstart

Audio-ready multimodal video.

curl -X POST "https://api.powertokens.ai/v1/videos" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "viduq3-pro",
  "prompt": "Cinematic tracking shot through a neon cyberpunk city in heavy rain.",
  "seconds": "5",
  "size": "1080p",
  "ratio": "16:9",
  "resolution": "1080p",
  "duration": 5,
  "aspect_ratio": "16:9",
  "audio": true
}'

Technical Specifications

Confirmed parameters and runtime execution protocols.

Provider & Model ID
Vidu • viduq3-pro
Modes
Text-to-video, Image-to-video, Start & end frame
Resolutions
540p, 720p, 1080p (default 540p)
Duration
Configurable via duration slider (default 5)
Aspect Ratios
16:9, 9:16, 3:4, 4:3, 1:1 (default 16:9 for T2V)
Audio
Optional; all / speech_only / sound_effect_only
Seed
Optional reproducible seed
Watermark
Optional with wm_position 1–4 and wm_url
Billing
See pricing matrix
API Endpoint
POST /v1/videos

Pricing Details

The actual billing for this model is dynamically calculated based on the specific parameters passed in your API request. Below are the specific combinations and their corresponding pricing:

Text to Video540PPeak Shifting
Credits24/ Second
Price (USD)$0.024
Text to Video540P
Credits43/ Second
Price (USD)$0.043
Text to Video720PPeak Shifting
Credits48/ Second
Price (USD)$0.048
Text to Video720P
Credits95/ Second
Price (USD)$0.095
Text to Video1080PPeak Shifting
Credits57/ Second
Price (USD)$0.057
Text to Video1080P
Credits114/ Second
Price (USD)$0.114
Image to Video540PPeak Shifting
Credits24/ Second
Price (USD)$0.024
Image to Video540P
Credits43/ Second
Price (USD)$0.043
Image to Video720PPeak Shifting
Credits48/ Second
Price (USD)$0.048
Image to Video720P
Credits95/ Second
Price (USD)$0.095
Image to Video1080PPeak Shifting
Credits57/ Second
Price (USD)$0.057
Image to Video1080P
Credits114/ Second
Price (USD)$0.114
Start&End Frame to Video540PPeak Shifting
Credits24/ Second
Price (USD)$0.024
Start&End Frame to Video540P
Credits43/ Second
Price (USD)$0.043
Start&End Frame to Video720PPeak Shifting
Credits48/ Second
Price (USD)$0.048
Start&End Frame to Video720P
Credits95/ Second
Price (USD)$0.095
Start&End Frame to Video1080PPeak Shifting
Credits57/ Second
Price (USD)$0.057
Start&End Frame to Video1080P
Credits114/ Second
Price (USD)$0.114
Ecosystem Models

Recommended Related Models

Explore complementary video and multimodal models with your unified API key.

Browse All Models
Seedance 2.5
videoCommercial
Seedance 2.5

Seedance 2.5

dreamina-seedance-2-5-260628

ByteDance's latest flagship video generation model, built for longer-form storytelling and production-ready output. Generates up to 30 seconds of continuous, cinematic video with native audio sync in a single pass . Accepts up to 50 multimodal references (images, videos, audio, character sheets, storyboards) for precise scene, character, and motion consistency . Features localized region editing to fix specific areas without full regeneration……

Text to VideoImage to Video
Wan 3.0
videoCommercial
Wan 3.0

Wan 3.0

wan3.0-video

Wan3.0-Video is an all-in-one video generation model unified support for multiple creative capabilities, including reference, editing, replication, and driving. It generates videos up to 30 seconds with omni-modal reference, and can parse files, web pages and complex images. With production-grade character consistency and lifelike visuals and sound, it delivers an immersive audiovisual experience.

Text to VideoImage to Video
Seedance 2.0 Mini
videoCommercial
Seedance 2.0 Mini

Seedance 2.0 Mini

dreamina-seedance-2-0-mini-260615

Lightweight, cost-efficient video model from ByteDance, optimized for speed and high-volume content creation. Supports text-to-video, image-to-video, and reference-based generation with up to 12 references (6 images, 3 audio, 3 video). Delivers faster generation and lower credit consumption than Seedance 2.0, with strong motion quality and character consistency. Ideal for social media content, product videos, AI short dramas, and rapid creative iteration

Text to VideoImage to Video
Seedance 2.0
videoCommercial
Seedance 2.0

Seedance 2.0

dreamina-seedance-2-0-260128

Generate videos from reference images, videos, and audio; edit videos; extend videos; generate videos from start and end frames

Text to VideoImage to Video

Frequently Asked Questions

Everything you need to know before integrating this model.

540p, 720p, and 1080p. Default is 540p.

Start Building with Vidu Q3 Pro Today

Create an account in seconds to receive 100 free credits and start generating immediately. No credit card or upfront contract required.