신규 가입 시 100 무료 크레딧 증정, 지금 바로 AI 애플리케이션을 탐색하고 구축하세요무료 받기
kling-video-o1

Kling Video O1 / 비디오 생성

Commercial
ID: kling-video-o1

World's first unified multimodal video model built on MVL (Multi-modal Visual Language) architecture. Accepts multimodal inputs—text, images, videos, and elements—for all-in-one creation and editing. Supports reference-based generation, start/end frame interpolation, video in/outpainting, stylization, and multi-subject consistency. Generates 3–10s clips with up to 7 reference images

가격$0.084/ Second(84 크레딧)
Input Prompt
0 / 2500

Upload Image List

JPG, JPEG, PNG (Max 10MB)

Upload Video List

MP4, MOV (Max 50MB)

비디오 생성

Video Playground Ready

왼쪽 매개변수 패널에서 프롬프트를 입력하고 옵션을 설정한 후 Generate를 클릭하세요.

Kling Video O1 Hero Visual
Kling Video Omni O1 • Multimodal Story Video

Kling Video O1Omni Video with Multi-Shot Story Control

Kling Video O1 unifies text, image, and video references into cinematic multi-shot clips — with optional sound, Pro quality mode, and 3–10 second durations.

Multi-Shot Custom or Auto Storyboarding
Image + Video Reference Conditioning
Optional Synchronized Sound
Pro Mode & 1:1 / 16:9 / 9:16
ByteDance Seed Foundation Architecture
Commercial License & Enterprise SLA
Omni
Multimodal Input
Text + image + video
3–10s
Clip Duration
Default 5 seconds
Multi
Shot Storyboard
Custom or auto
Pro
Default Quality
Standard available
Capability Highlights

Omni Storytelling Architecture

Blend prompts, stills, and motion references into one coherent multi-shot narrative.

Multi-Shot Story Control

Write custom shot lists or let the model storyboard automatically. Each shot stays coherent across cuts and camera changes.

Multi-ShotCustom / AutoStory Arc
Multi-shot storyboard
Multimodal conditioning

Multimodal Reference Input

Condition on still images and existing video clips together so subject identity, lighting, and motion style carry through the generation.

Image ReferenceVideo ReferenceIdentity Lock

Optional Synchronized Sound

Enable native sound for finished social cuts, or keep it off when you will score the clip in post.

Sound On/OffSocial ReadyPost Flexibility
Native Sound
Pro Fidelity

Pro Mode Fidelity

Default Pro mode prioritizes temporal coherence and material detail for hero placements; Standard is available for drafts.

Pro DefaultTemporal CoherenceHero Clips
How It Works

How It Works

A production-ready path from brief to finished asset.

01

Assemble References

Collect stills and optional motion references that define subject identity and style.

02

Outline the Story Beats

Write custom shot prompts or let auto mode structure the narrative arc.

03

Set Canvas & Sound

Choose 1:1 / 16:9 / 9:16, duration, Pro mode, and whether synchronized sound is needed.

04

Generate & Chain

Pull the finished multi-shot clip and chain further shots with consistent references.

Cinematic Omni Domains

Where multimodal storytelling replaces multi-tool pipelines.

Marketing

Brand Story Shorts

Multi-shot narrative ads from stills plus a brief.

Multi-Shot
Content

Character Episodes

Keep persona consistent across sequential scenes.

Identity
Commerce

Product Demo Films

Animate packshots with optional synchronized sound.

Sound On
Growth

Social Series

9:16 vertical multi-beat clips for Reels and TikTok.

9:16
Best Practices

Prompt & Usage Tips

Practical guidance for reliable first-pass results.

Describe each shot’s action

In multi-shot mode, write motion per shot instead of one long paragraph.

Lock identity with stills

A clear front-lit subject still beats text-only identity descriptions.

Sound only when useful

Leave sound off for plates you will score externally; enable it for finished social cuts.

Developer Quickstart

Omni Video Quickstart

Generate multi-shot clips with references and optional sound.

curl -X POST "https://api.powertokens.ai/v1/videos" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "kling-video-o1",
  "prompt": "Cinematic tracking shot through a neon cyberpunk city in heavy rain.",
  "seconds": "5",
  "size": "1080p",
  "ratio": "16:9",
  "mode": "pro",
  "aspect_ratio": "16:9",
  "duration": "5",
  "sound": "on"
}'

Technical Specifications

Confirmed parameters and runtime execution protocols.

Provider & Model ID
Kuaishou (Kling) • kling-video-o1
Modes
Omni video generation (text / image / video references)
Multi-Shot
Custom shot lists or auto storyboarding
Duration
3 to 10 seconds (default 5)
Aspect Ratios
1:1, 16:9, 9:16 (default 1:1)
Quality Mode
Standard or Pro (default Pro)
Sound
Optional synchronized audio (off by default)
References
Image list and video list supported
Watermark
Optional — off by default
API Endpoint
POST /v1/videos

가격 상세

이 모델의 실제 요금은 API 요청에서 전달된 특정 매개변수를 기반으로 동적으로 계산됩니다. 아래는 구체적인 조합과 해당 가격입니다:

720PWithout Video
크레딧84/ Second
가격 (USD)$0.084
720PWith Video
크레딧126/ Second
가격 (USD)$0.126
1080PWithout Video
크레딧112/ Second
가격 (USD)$0.112
1080PWith Video
크레딧168/ Second
가격 (USD)$0.168
Same Channel

Models from the Same Channel

Explore complementary models and alternative versions from the same provider channel.

Kling 3.0 Omni
image,videoCommercial
Kling 3.0 Omni

Kling 3.0 Omni

kling-v3-omni

Flagship unified multimodal model integrating text-to-video, image-to-video, and reference-based generation. Supports up to 15-second cinematic clips with native synchronized audio (dialogue, SFX, BGM). Enables multi-shot control (up to 6 shots) and consistent subject/character preservation across scenes. Pro mode outputs 1080p with enhanced motion realism

Text to ImageImage to Image
kling v3
image,videoCommercial
kling v3

kling v3

kling-v3

Next-generation video generation model offering Standard and Pro tiers. Generates 3–15 second clips at up to 1080p resolution from text or image inputs. Features first-frame and last-frame control for precise scene composition. Supports 16:9, 9:16, and 1:1 aspect ratios. Native audio generation available as an optional feature

Text to ImageImage to Image
Kling 2.5 Turbo
videoCommercial
Kling 2.5 Turbo

Kling 2.5 Turbo

kling-v2-5-turbo

Cost-optimized turbo variant of the V2.5 series, reducing generation costs by nearly 30%. Delivers high-speed video generation at 1080p with fewer refinement passes. Ideal for high-volume creative pipelines and rapid social content iteration. Supports text-to-video and image-to-video workflows

Text to VideoImage to Video
Kling 2.1 Master
videoCommercial
Kling 2.1 Master

Kling 2.1 Master

kling-v2-1-master

Premium Master edition of the V2.1 series. Generates 5-second and 10-second videos at 1080p from text or image inputs. Offers superior motion dynamics, stronger prompt adherence, and enhanced subject consistency. Features first-frame and last-frame control for guided composition

Text to VideoImage to Video
Ecosystem Models

Recommended Related Models

Explore complementary video and multimodal models with your unified API key.

Browse All Models
Seedance 2.5
videoCommercial
Seedance 2.5

Seedance 2.5

dreamina-seedance-2-5-260628

🔥 LIMITED-TIME OFFER: 1080P at 20% OFF! 🔥 ByteDance's latest flagship video generation model, built for longer-form storytelling and production-ready output. Generates up to 30 seconds of continuous, cinematic video with native audio sync in a single pass . Accepts up to 50 multimodal references (images, videos, audio, character sheets, storyboards) for precise scene, character, and motion consistency . Features localized region editing to fix specific areas without full regeneration……

Text to VideoImage to Video
Wan 3.0
videoCommercial
Wan 3.0

Wan 3.0

wan3.0-video

Wan3.0-Video is an all-in-one video generation model unified support for multiple creative capabilities, including reference, editing, replication, and driving. It generates videos up to 30 seconds with omni-modal reference, and can parse files, web pages and complex images. With production-grade character consistency and lifelike visuals and sound, it delivers an immersive audiovisual experience.

Text to VideoImage to Video
Seedance 2.0 Mini
videoCommercial
Seedance 2.0 Mini

Seedance 2.0 Mini

dreamina-seedance-2-0-mini-260615

Lightweight, cost-efficient video model from ByteDance, optimized for speed and high-volume content creation. Supports text-to-video, image-to-video, and reference-based generation with up to 12 references (6 images, 3 audio, 3 video). Delivers faster generation and lower credit consumption than Seedance 2.0, with strong motion quality and character consistency. Ideal for social media content, product videos, AI short dramas, and rapid creative iteration

Text to VideoImage to Video
Seedance 2.0
videoCommercial
Seedance 2.0

Seedance 2.0

dreamina-seedance-2-0-260128

Generate videos from reference images, videos, and audio; edit videos; extend videos; generate videos from start and end frames

Text to VideoImage to Video

Frequently Asked Questions

Everything you need to know before integrating this model.

It accepts text prompts plus image and video references in one request, so you can lock subject identity and motion style together.

Start Building with Kling Video O1 Today

Create an account in seconds to receive 100 free credits and start generating immediately. No credit card or upfront contract required.