세계 최고 수준의 AI 모델을 매우 합리적인 가격으로 이용하십시오.

🔥 LIMITED-TIME OFFER: 1080P at 20% OFF! 🔥 ByteDance's latest flagship video generation model, built for longer-form storytelling and production-ready output. Generates up to 30 seconds of continuous, cinematic video with native audio sync in a single pass . Accepts up to 50 multimodal references (images, videos, audio, character sheets, storyboards) for precise scene, character, and motion consistency . Features localized region editing to fix specific areas without full regeneration……
참고: "비디오 입력 포함" 모델의 단가가 더 저렴하다는 점을 확인하셨을 수 있습니다. 이는 과금 방식이 다르기 때문입니다:
비디오 입력 없음: 총 비용 = 단가 × 출력 시간
비디오 입력 포함: 총 비용 = 단가 × (입력 + 출력) 시간
| 모달리티 | 크레딧 / 단위 | 가격 (USD) | ≈초당 가격 (USD) |
|---|---|---|---|
480P/Without Video videoBytePlus | 10,700/ 1M Tokens | $10.700 | $0.105 |
480P/With Video videoBytePlus | 6,400/ 1M Tokens | $6.400 | $0.125 |
720P/Without Video videoBytePlus | 10,700/ 1M Tokens | $10.700 | $0.232 |
720P/With Video videoBytePlus | 6,400/ 1M Tokens | $6.400 | $0.277 |
1080P/Without Video videoBytePlus | 9,360/ 1M Tokens | $9.360 | $0.455 |
1080P/With Video videoBytePlus | 5,600/ 1M Tokens | $5.600 | $0.545 |

Wan3.0-Video is an all-in-one video generation model unified support for multiple creative capabilities, including reference, editing, replication, and driving. It generates videos up to 30 seconds with omni-modal reference, and can parse files, web pages and complex images. With production-grade character consistency and lifelike visuals and sound, it delivers an immersive audiovisual experience.
| 모달리티 | 크레딧 / 단위 | 가격 (USD) |
|---|---|---|
480P videoAlibaba | 50/ Second | $0.050 |
720P videoAlibaba | 100/ Second | $0.100 |
1080P videoAlibaba | 200/ Second | $0.200 |

Lightweight, cost-efficient video model from ByteDance, optimized for speed and high-volume content creation. Supports text-to-video, image-to-video, and reference-based generation with up to 12 references (6 images, 3 audio, 3 video). Delivers faster generation and lower credit consumption than Seedance 2.0, with strong motion quality and character consistency. Ideal for social media content, product videos, AI short dramas, and rapid creative iteration
참고: "비디오 입력 포함" 모델의 단가가 더 저렴하다는 점을 확인하셨을 수 있습니다. 이는 과금 방식이 다르기 때문입니다:
비디오 입력 없음: 총 비용 = 단가 × 출력 시간
비디오 입력 포함: 총 비용 = 단가 × (입력 + 출력) 시간
| 모달리티 | 크레딧 / 단위 | 가격 (USD) | ≈초당 가격 (USD) |
|---|---|---|---|
Without Video videoBytePlus | 3,500/ 1M Tokens | $3.500 | $0.076 |
With Video videoBytePlus | 2,100/ 1M Tokens | $2.100 | $0.091 |

Zhipu AI's flagship text model optimized for complex software engineering and long-horizon Agent tasks. Features a 1M-token context window with mandatory reasoning (3 levels: low/high/max). Coding capability improved 50% over GLM-5.2 on Z.ai Code Bench; scores SOTA on Terminal Bench 3.0. Excels in cybersecurity tasks (cyber vulnerability discovery)
| 모달리티 | 입력 크레딧 | 출력 크레딧 | 입력 가격 | 출력 가격 | 암시적 캐시 적중 | 명시적 캐시 적중 | 캐시 생성 |
|---|---|---|---|---|---|---|---|
Standard chatZ.ai | 1,400/ 1M Tokens | 4,400/ 1M Tokens | $1.400 | $4.400 | $0.260 260/ 1M Tokens | $0.260 260/ 1M Tokens | -- |

Generate videos from reference images, videos, and audio; edit videos; extend videos; generate videos from start and end frames
참고: "비디오 입력 포함" 모델의 단가가 더 저렴하다는 점을 확인하셨을 수 있습니다. 이는 과금 방식이 다르기 때문입니다:
비디오 입력 없음: 총 비용 = 단가 × 출력 시간
비디오 입력 포함: 총 비용 = 단가 × (입력 + 출력) 시간
| 모달리티 | 크레딧 / 단위 | 가격 (USD) | ≈초당 가격 (USD) |
|---|---|---|---|
480P/Without Video videoBytePlus | 7,000/ 1M Tokens | $7.000 | $0.069 |
480P/With Video videoBytePlus | 4,300/ 1M Tokens | $4.300 | $0.084 |
720P/Without Video videoBytePlus | 7,000/ 1M Tokens | $7.000 | $0.152 |
720P/With Video videoBytePlus | 4,300/ 1M Tokens | $4.300 | $0.186 |
1080P/Without Video videoBytePlus | 7,700/ 1M Tokens | $7.700 | $0.375 |
1080P/With Video videoBytePlus | 4,700/ 1M Tokens | $4.700 | $0.457 |
4K/Without Video videoBytePlus | 4,000/ 1M Tokens | $4.000 | $0.778 |
4K/With Video videoBytePlus | 2,400/ 1M Tokens | $2.400 | $0.934 |

Wan3.0-Video-Prime is the high-speed video generation model of Wan3.0, with capabilities aligned to the standard version of Wan3.0-Video. It supports four-modal all-in-one reference and generates videos of up to 30 seconds, delivering an immersive audiovisual experience with significantly faster end-to-end generation.
| 모달리티 | 크레딧 / 단위 | 가격 (USD) |
|---|---|---|
480P videoAlibaba | 68/ Second | $0.068 |
720P videoAlibaba | 140/ Second | $0.140 |
1080P videoAlibaba | 280/ Second | $0.280 |

Generate videos with reference to images/videos/audio, edit videos, extend videos, generate videos from first and last frames
참고: "비디오 입력 포함" 모델의 단가가 더 저렴하다는 점을 확인하셨을 수 있습니다. 이는 과금 방식이 다르기 때문입니다:
비디오 입력 없음: 총 비용 = 단가 × 출력 시간
비디오 입력 포함: 총 비용 = 단가 × (입력 + 출력) 시간
| 모달리티 | 크레딧 / 단위 | 가격 (USD) | ≈초당 가격 (USD) |
|---|---|---|---|
Without Video videoBytePlus | 5,600/ 1M Tokens | $5.600 | $0.121 |
With Video videoBytePlus | 3,300/ 1M Tokens | $3.300 | $0.143 |

Flagship unified multimodal model integrating text-to-video, image-to-video, and reference-based generation. Supports up to 15-second cinematic clips with native synchronized audio (dialogue, SFX, BGM). Enables multi-shot control (up to 6 shots) and consistent subject/character preservation across scenes. Pro mode outputs 1080p with enhanced motion realism
| 모달리티 | 크레딧 / 단위 | 가격 (USD) |
|---|---|---|
1K imagekling | 28/ Image | $0.028 |
2K imagekling | 28/ Image | $0.028 |
4K imagekling | 56/ Image | $0.056 |
720P/Without Audio/Without Video videokling | 84/ Second | $0.084 |
720P/With Audio/Without Video videokling | 112/ Second | $0.112 |
720P/Without Audio/With Video videokling | 126/ Second | $0.126 |
1080P/Without Audio/Without Video videokling | 112/ Second | $0.112 |
1080P/With Audio/Without Video videokling | 140/ Second | $0.140 |
1080P/Without Audio/With Video videokling | 168/ Second | $0.168 |
4K/Without Audio/Without Video videokling | 420/ Second | $0.420 |
4K/With Audio/Without Video videokling | 420/ Second | $0.420 |
4K/Without Audio/With Video videokling | 420/ Second | $0.420 |

Next-generation video generation model offering Standard and Pro tiers. Generates 3–15 second clips at up to 1080p resolution from text or image inputs. Features first-frame and last-frame control for precise scene composition. Supports 16:9, 9:16, and 1:1 aspect ratios. Native audio generation available as an optional feature
| 모달리티 | 크레딧 / 단위 | 가격 (USD) |
|---|---|---|
Standard imagekling | 28/ Image | $0.028 |
Motion Control/720P videokling | 126/ Second | $0.126 |
Motion Control/1080P videokling | 168/ Second | $0.168 |
720P/Without Audio videokling | 84/ Second | $0.084 |
720P/With Audio/Unspecified Voice videokling | 126/ Second | $0.126 |
720P/With Audio/Specified Voice videokling | 154/ Second | $0.154 |
1080P/Without Audio videokling | 112/ Second | $0.112 |
1080P/With Audio/Unspecified Voice videokling | 168/ Second | $0.168 |
1080P/With Audio/Specified Voice videokling | 196/ Second | $0.196 |
4K/Without Audio videokling | 420/ Second | $0.420 |
4K/With Audio/Unspecified Voice videokling | 420/ Second | $0.420 |

World's first unified multimodal video model built on MVL (Multi-modal Visual Language) architecture. Accepts multimodal inputs—text, images, videos, and elements—for all-in-one creation and editing. Supports reference-based generation, start/end frame interpolation, video in/outpainting, stylization, and multi-subject consistency. Generates 3–10s clips with up to 7 reference images
| 모달리티 | 크레딧 / 단위 | 가격 (USD) |
|---|---|---|
720P/Without Video videokling | 84/ Second | $0.084 |
720P/With Video videokling | 126/ Second | $0.126 |
1080P/Without Video videokling | 112/ Second | $0.112 |
1080P/With Video videokling | 168/ Second | $0.168 |

Alibaba's cost-efficient multimodal reasoning model. Supports text, image, and video inputs with text output. Features a native 1M-token context window for long documents, codebases, and agentic workflows. Excels in coding assistance, desktop interaction, chart analysis, and long-video understanding. Compatible with OpenAI/Anthropic protocols for seamless integration
| 모달리티 | 입력 크레딧 | 출력 크레딧 | 입력 가격 | 출력 가격 | 암시적 캐시 적중 | 명시적 캐시 적중 | 캐시 생성 |
|---|---|---|---|---|---|---|---|
Standard chatAlibaba | 150/ 1M Tokens | 470/ 1M Tokens | $0.150 | $0.470 | $0.016 16/ 1M Tokens | $0.016 16/ 1M Tokens | $0.200 200/ 1M Tokens |

Flagship text model purpose-built for long-horizon agentic workflows. Features a 1M context window supporting project-level engineering in a single session. Excels at autonomous coding: can complete development, testing, and multi-platform deployment from a single prompt. Top open-weight model per Artificial Analysis; #1 globally on Code Arena. MIT-licensed and Day-0 optimized for domestic AI chips
| 모달리티 | 입력 크레딧 | 출력 크레딧 | 입력 가격 | 출력 가격 | 암시적 캐시 적중 | 명시적 캐시 적중 | 캐시 생성 |
|---|---|---|---|---|---|---|---|
Standard chatZ.ai | 1,400/ 1M Tokens | 4,400/ 1M Tokens | $1.400 | $4.400 | $0.260 260/ 1M Tokens | $0.260 260/ 1M Tokens | -- |

Open-weight general-purpose multimodal video model. Unifies text, image, video, and audio understanding in a single context window, generating up to 2K resolution, 15-second clips with native stereo audio at 24fps. Supports multimodal reference inputs: up to 9 images, 3 videos, and 3 audio clips (12 total references) per generation. Features first-frame, last-frame, and full reference modes with conversational editing capabilities.
참고: 각 요청의 처음 5개 입력 이미지는 무료입니다. 이후 입력 이미지는 표에 표시된 입력 가격에 따라 청구됩니다.
요금 청구 규칙: 입력 비디오와 출력 비디오 모두 비디오 초당 요금이 청구되며, 청구 시간 = 입력 비디오 시간 + 출력 비디오 시간입니다.
| 모달리티 | 입력 | 출력 |
|---|---|---|
768P videoMiniMax | $0.040 40/ Image | $0.080 80/ Second |
2K videoMiniMax | $0.040 40/ Image | $0.130 130/ Second |

Compared with the September 23, 2025 version, the newly upgraded Qwen-3 Max seamlessly integrates thinking and non-thinking modes, bringing an all-round obvious performance boost. Its thinking mode supports web search, web content extraction and code interpreter. It can conduct in-depth logical reasoning and call external tools to solve intricate problems more precisely
| 모달리티 | 입력 크레딧 | 출력 크레딧 | 입력 가격 | 출력 가격 | 암시적 캐시 적중 | 명시적 캐시 적중 | 캐시 생성 |
|---|---|---|---|---|---|---|---|
0 - 32K chatAlibaba | 1,140/ 1M Tokens | 5,700/ 1M Tokens | $1.140 | $5.700 | $0.228 228/ 1M Tokens | $0.114 114/ 1M Tokens | $1.425 1,425/ 1M Tokens |
32K - 128K chatAlibaba | 2,280/ 1M Tokens | 11,400/ 1M Tokens | $2.280 | $11.400 | $0.456 456/ 1M Tokens | $0.228 228/ 1M Tokens | $2.850 2,850/ 1M Tokens |
128K - 256K chatAlibaba | 2,850/ 1M Tokens | 14,250/ 1M Tokens | $2.850 | $14.250 | $0.570 570/ 1M Tokens | $0.285 285/ 1M Tokens | $3.563 3,563/ 1M Tokens |

GLM-5 is Z.ai’s flagship open-source foundation model engineered for complex systems design and long-horizon agent workflows. Built for expert developers, it delivers production-grade performance on large-scale programming tasks, rivaling leading closed-source models. With advanced agentic planning, deep backend reasoning, and iterative self-correction, GLM-5 moves beyond code generation to full-system construction and autonomous execution.
| 모달리티 | 입력 크레딧 | 출력 크레딧 | 입력 가격 | 출력 가격 | 암시적 캐시 적중 | 명시적 캐시 적중 | 캐시 생성 |
|---|---|---|---|---|---|---|---|
Standard chatZ.ai | 970/ 1M Tokens | 3,104/ 1M Tokens | $0.970 | $3.104 | $0.194 194/ 1M Tokens | $0.194 194/ 1M Tokens | -- |

Supports text, single-image and multi-image inputs, and enables the generation of image sets
| 모달리티 | 크레딧 / 단위 | 가격 (USD) |
|---|---|---|
Standard imageBytePlus | 34/ Image | $0.034 |

Flagship multimodal foundation model supporting text, image, and video inputs with text output. Features a 1M-token context window via MiniMax Sparse Attention (MSA), cutting per-token compute to ~1/20 of previous gen at full context. Excels at long-horizon agentic work, coding, and tool use. Native multimodal training from step zero ensures deep semantic alignment. Scores 59.0% on SWE-Bench Pro and 83.5 on BrowseComp, surpassing Opus 4.7
| 모달리티 | 입력 크레딧 | 출력 크레딧 | 입력 가격 | 출력 가격 | 암시적 캐시 적중 | 명시적 캐시 적중 | 캐시 생성 |
|---|---|---|---|---|---|---|---|
0 - 512K chatMiniMax | 600/ 1M Tokens | 2,400/ 1M Tokens | $0.600 | $2.400 | $0.120 120/ 1M Tokens | $0.120 120/ 1M Tokens | -- |
Over 512K chatMiniMax | 1,200/ 1M Tokens | 4,800/ 1M Tokens | $1.200 | $4.800 | $0.240 240/ 1M Tokens | $0.240 240/ 1M Tokens | -- |

Powered by Qwen3, this is a powerful Coding Agent that excels in tool calling and environment interaction to achieve autonomous programming. It combines outstanding coding proficiency with versatile general-purpose abilities.
| 모달리티 | 입력 크레딧 | 출력 크레딧 | 입력 가격 | 출력 가격 | 암시적 캐시 적중 | 명시적 캐시 적중 | 캐시 생성 |
|---|---|---|---|---|---|---|---|
0 - 32K chatAlibaba | 1,000/ 1M Tokens | 5,000/ 1M Tokens | $1.000 | $5.000 | $0.200 200/ 1M Tokens | $0.100 100/ 1M Tokens | $1.250 1,250/ 1M Tokens |
32K - 128K chatAlibaba | 1,800/ 1M Tokens | 9,000/ 1M Tokens | $1.800 | $9.000 | $0.360 360/ 1M Tokens | $0.180 180/ 1M Tokens | $2.250 2,250/ 1M Tokens |
128K - 256K chatAlibaba | 3,000/ 1M Tokens | 15,000/ 1M Tokens | $3.000 | $15.000 | $0.600 600/ 1M Tokens | $0.300 300/ 1M Tokens | $3.750 3,750/ 1M Tokens |
256K - 1M chatAlibaba | 6,000/ 1M Tokens | 60,000/ 1M Tokens | $6.000 | $60.000 | $1.200 1,200/ 1M Tokens | $0.600 600/ 1M Tokens | $7.500 7,500/ 1M Tokens |

The Qwen3.6 native vision-language Plus series models demonstrate exceptional performance on par with the current state-of-the-art models, with a significant improvement in overall results compared to the 3.5 series. The models have been markedly enhanced in code-related capabilities such as agentic coding, front-end programming, and Vibe coding, as well as in multi-modal general object recognition, OCR, and object localization.
| 모달리티 | 입력 크레딧 | 출력 크레딧 | 입력 가격 | 출력 가격 | 암시적 캐시 적중 | 명시적 캐시 적중 | 캐시 생성 |
|---|---|---|---|---|---|---|---|
0 - 256K chatAlibaba | 475/ 1M Tokens | 2,850/ 1M Tokens | $0.475 | $2.850 | -- | $0.048 48/ 1M Tokens | $0.594 594/ 1M Tokens |
256K - 1M chatAlibaba | 1,900/ 1M Tokens | 5,700/ 1M Tokens | $1.900 | $5.700 | -- | $0.190 190/ 1M Tokens | $2.375 2,375/ 1M Tokens |

Supports generating video with audio from text and images, and supports first and last frames.
| 모달리티 | 크레딧 / 단위 | 가격 (USD) | ≈초당 가격 (USD) |
|---|---|---|---|
Without Audio videoBytePlus | 1,200/ 1M Tokens | $1.200 | $0.026 |
With Audio videoBytePlus | 2,280/ 1M Tokens | $2.280 | $0.050 |
* 예상 평균 동영상 = 720p에서 5초 또는 1024x1024 이미지를 기준으로 합니다. 실제 출력 비용은 특정 모델 매개변수, 해상도 배율 및 프롬프트 복잡도에 따라 달라질 수 있습니다.