世界クラスの AI モデルを、驚くほど手頃な価格で。

🔥 LIMITED-TIME OFFER: 1080P at 20% OFF! 🔥 ByteDance's latest flagship video generation model, built for longer-form storytelling and production-ready output. Generates up to 30 seconds of continuous, cinematic video with native audio sync in a single pass . Accepts up to 50 multimodal references (images, videos, audio, character sheets, storyboards) for precise scene, character, and motion consistency . Features localized region editing to fix specific areas without full regeneration……
注: 「動画入力あり」のモデルの単価が低くなっていることにお気づきかもしれません。これは課金方法が異なるためです:
動画入力なし:合計費用 = 単価 × 出力時間
動画入力あり:合計費用 = 単価 × (入力 + 出力) 時間
| モダリティ | クレジット / 単位 | 料金 (USD) | ≈秒あたりの価格 (USD) |
|---|---|---|---|
480P/Without Video videoBytePlus | 10,700/ 1M Tokens | $10.700 | $0.105 |
480P/With Video videoBytePlus | 6,400/ 1M Tokens | $6.400 | $0.125 |
720P/Without Video videoBytePlus | 10,700/ 1M Tokens | $10.700 | $0.232 |
720P/With Video videoBytePlus | 6,400/ 1M Tokens | $6.400 | $0.277 |
1080P/Without Video videoBytePlus | 9,360/ 1M Tokens | $9.360 | $0.455 |
1080P/With Video videoBytePlus | 5,600/ 1M Tokens | $5.600 | $0.545 |

Wan3.0-Video is an all-in-one video generation model unified support for multiple creative capabilities, including reference, editing, replication, and driving. It generates videos up to 30 seconds with omni-modal reference, and can parse files, web pages and complex images. With production-grade character consistency and lifelike visuals and sound, it delivers an immersive audiovisual experience.
| モダリティ | クレジット / 単位 | 料金 (USD) |
|---|---|---|
480P videoAlibaba | 50/ Second | $0.050 |
720P videoAlibaba | 100/ Second | $0.100 |
1080P videoAlibaba | 200/ Second | $0.200 |

Lightweight, cost-efficient video model from ByteDance, optimized for speed and high-volume content creation. Supports text-to-video, image-to-video, and reference-based generation with up to 12 references (6 images, 3 audio, 3 video). Delivers faster generation and lower credit consumption than Seedance 2.0, with strong motion quality and character consistency. Ideal for social media content, product videos, AI short dramas, and rapid creative iteration
注: 「動画入力あり」のモデルの単価が低くなっていることにお気づきかもしれません。これは課金方法が異なるためです:
動画入力なし:合計費用 = 単価 × 出力時間
動画入力あり:合計費用 = 単価 × (入力 + 出力) 時間
| モダリティ | クレジット / 単位 | 料金 (USD) | ≈秒あたりの価格 (USD) |
|---|---|---|---|
Without Video videoBytePlus | 3,500/ 1M Tokens | $3.500 | $0.076 |
With Video videoBytePlus | 2,100/ 1M Tokens | $2.100 | $0.091 |

Zhipu AI's flagship text model optimized for complex software engineering and long-horizon Agent tasks. Features a 1M-token context window with mandatory reasoning (3 levels: low/high/max). Coding capability improved 50% over GLM-5.2 on Z.ai Code Bench; scores SOTA on Terminal Bench 3.0. Excels in cybersecurity tasks (cyber vulnerability discovery)
| モダリティ | 入力クレジット | 出力クレジット | 入力価格 | 出力価格 | 暗黙的キャッシュヒット | 明示的キャッシュヒット | キャッシュ作成 |
|---|---|---|---|---|---|---|---|
Standard chatZ.ai | 1,400/ 1M Tokens | 4,400/ 1M Tokens | $1.400 | $4.400 | $0.260 260/ 1M Tokens | $0.260 260/ 1M Tokens | -- |

Generate videos from reference images, videos, and audio; edit videos; extend videos; generate videos from start and end frames
注: 「動画入力あり」のモデルの単価が低くなっていることにお気づきかもしれません。これは課金方法が異なるためです:
動画入力なし:合計費用 = 単価 × 出力時間
動画入力あり:合計費用 = 単価 × (入力 + 出力) 時間
| モダリティ | クレジット / 単位 | 料金 (USD) | ≈秒あたりの価格 (USD) |
|---|---|---|---|
480P/Without Video videoBytePlus | 7,000/ 1M Tokens | $7.000 | $0.069 |
480P/With Video videoBytePlus | 4,300/ 1M Tokens | $4.300 | $0.084 |
720P/Without Video videoBytePlus | 7,000/ 1M Tokens | $7.000 | $0.152 |
720P/With Video videoBytePlus | 4,300/ 1M Tokens | $4.300 | $0.186 |
1080P/Without Video videoBytePlus | 7,700/ 1M Tokens | $7.700 | $0.375 |
1080P/With Video videoBytePlus | 4,700/ 1M Tokens | $4.700 | $0.457 |
4K/Without Video videoBytePlus | 4,000/ 1M Tokens | $4.000 | $0.778 |
4K/With Video videoBytePlus | 2,400/ 1M Tokens | $2.400 | $0.934 |

Wan3.0-Video-Prime is the high-speed video generation model of Wan3.0, with capabilities aligned to the standard version of Wan3.0-Video. It supports four-modal all-in-one reference and generates videos of up to 30 seconds, delivering an immersive audiovisual experience with significantly faster end-to-end generation.
| モダリティ | クレジット / 単位 | 料金 (USD) |
|---|---|---|
480P videoAlibaba | 68/ Second | $0.068 |
720P videoAlibaba | 140/ Second | $0.140 |
1080P videoAlibaba | 280/ Second | $0.280 |

Generate videos with reference to images/videos/audio, edit videos, extend videos, generate videos from first and last frames
注: 「動画入力あり」のモデルの単価が低くなっていることにお気づきかもしれません。これは課金方法が異なるためです:
動画入力なし:合計費用 = 単価 × 出力時間
動画入力あり:合計費用 = 単価 × (入力 + 出力) 時間
| モダリティ | クレジット / 単位 | 料金 (USD) | ≈秒あたりの価格 (USD) |
|---|---|---|---|
Without Video videoBytePlus | 5,600/ 1M Tokens | $5.600 | $0.121 |
With Video videoBytePlus | 3,300/ 1M Tokens | $3.300 | $0.143 |

Flagship unified multimodal model integrating text-to-video, image-to-video, and reference-based generation. Supports up to 15-second cinematic clips with native synchronized audio (dialogue, SFX, BGM). Enables multi-shot control (up to 6 shots) and consistent subject/character preservation across scenes. Pro mode outputs 1080p with enhanced motion realism
| モダリティ | クレジット / 単位 | 料金 (USD) |
|---|---|---|
1K imagekling | 28/ Image | $0.028 |
2K imagekling | 28/ Image | $0.028 |
4K imagekling | 56/ Image | $0.056 |
720P/Without Audio/Without Video videokling | 84/ Second | $0.084 |
720P/With Audio/Without Video videokling | 112/ Second | $0.112 |
720P/Without Audio/With Video videokling | 126/ Second | $0.126 |
1080P/Without Audio/Without Video videokling | 112/ Second | $0.112 |
1080P/With Audio/Without Video videokling | 140/ Second | $0.140 |
1080P/Without Audio/With Video videokling | 168/ Second | $0.168 |
4K/Without Audio/Without Video videokling | 420/ Second | $0.420 |
4K/With Audio/Without Video videokling | 420/ Second | $0.420 |
4K/Without Audio/With Video videokling | 420/ Second | $0.420 |

Next-generation video generation model offering Standard and Pro tiers. Generates 3–15 second clips at up to 1080p resolution from text or image inputs. Features first-frame and last-frame control for precise scene composition. Supports 16:9, 9:16, and 1:1 aspect ratios. Native audio generation available as an optional feature
| モダリティ | クレジット / 単位 | 料金 (USD) |
|---|---|---|
Standard imagekling | 28/ Image | $0.028 |
Motion Control/720P videokling | 126/ Second | $0.126 |
Motion Control/1080P videokling | 168/ Second | $0.168 |
720P/Without Audio videokling | 84/ Second | $0.084 |
720P/With Audio/Unspecified Voice videokling | 126/ Second | $0.126 |
720P/With Audio/Specified Voice videokling | 154/ Second | $0.154 |
1080P/Without Audio videokling | 112/ Second | $0.112 |
1080P/With Audio/Unspecified Voice videokling | 168/ Second | $0.168 |
1080P/With Audio/Specified Voice videokling | 196/ Second | $0.196 |
4K/Without Audio videokling | 420/ Second | $0.420 |
4K/With Audio/Unspecified Voice videokling | 420/ Second | $0.420 |

World's first unified multimodal video model built on MVL (Multi-modal Visual Language) architecture. Accepts multimodal inputs—text, images, videos, and elements—for all-in-one creation and editing. Supports reference-based generation, start/end frame interpolation, video in/outpainting, stylization, and multi-subject consistency. Generates 3–10s clips with up to 7 reference images
| モダリティ | クレジット / 単位 | 料金 (USD) |
|---|---|---|
720P/Without Video videokling | 84/ Second | $0.084 |
720P/With Video videokling | 126/ Second | $0.126 |
1080P/Without Video videokling | 112/ Second | $0.112 |
1080P/With Video videokling | 168/ Second | $0.168 |

Alibaba's cost-efficient multimodal reasoning model. Supports text, image, and video inputs with text output. Features a native 1M-token context window for long documents, codebases, and agentic workflows. Excels in coding assistance, desktop interaction, chart analysis, and long-video understanding. Compatible with OpenAI/Anthropic protocols for seamless integration
| モダリティ | 入力クレジット | 出力クレジット | 入力価格 | 出力価格 | 暗黙的キャッシュヒット | 明示的キャッシュヒット | キャッシュ作成 |
|---|---|---|---|---|---|---|---|
Standard chatAlibaba | 150/ 1M Tokens | 470/ 1M Tokens | $0.150 | $0.470 | $0.016 16/ 1M Tokens | $0.016 16/ 1M Tokens | $0.200 200/ 1M Tokens |

Flagship text model purpose-built for long-horizon agentic workflows. Features a 1M context window supporting project-level engineering in a single session. Excels at autonomous coding: can complete development, testing, and multi-platform deployment from a single prompt. Top open-weight model per Artificial Analysis; #1 globally on Code Arena. MIT-licensed and Day-0 optimized for domestic AI chips
| モダリティ | 入力クレジット | 出力クレジット | 入力価格 | 出力価格 | 暗黙的キャッシュヒット | 明示的キャッシュヒット | キャッシュ作成 |
|---|---|---|---|---|---|---|---|
Standard chatZ.ai | 1,400/ 1M Tokens | 4,400/ 1M Tokens | $1.400 | $4.400 | $0.260 260/ 1M Tokens | $0.260 260/ 1M Tokens | -- |

Open-weight general-purpose multimodal video model. Unifies text, image, video, and audio understanding in a single context window, generating up to 2K resolution, 15-second clips with native stereo audio at 24fps. Supports multimodal reference inputs: up to 9 images, 3 videos, and 3 audio clips (12 total references) per generation. Features first-frame, last-frame, and full reference modes with conversational editing capabilities.
注: 各リクエストの最初の5枚の入力画像は無料です。それ以降の入力画像は表に表示されている入力価格に基づいて課金されます。
課金ルール:入力動画と出力動画の両方が課金対象となり、動画の秒数単位で計算されます。課金対象時間 = 入力動画の長さ + 出力動画の長さ。
| モダリティ | 入力 | 出力 |
|---|---|---|
768P videoMiniMax | $0.040 40/ Image | $0.080 80/ Second |
2K videoMiniMax | $0.040 40/ Image | $0.130 130/ Second |

Compared with the September 23, 2025 version, the newly upgraded Qwen-3 Max seamlessly integrates thinking and non-thinking modes, bringing an all-round obvious performance boost. Its thinking mode supports web search, web content extraction and code interpreter. It can conduct in-depth logical reasoning and call external tools to solve intricate problems more precisely
| モダリティ | 入力クレジット | 出力クレジット | 入力価格 | 出力価格 | 暗黙的キャッシュヒット | 明示的キャッシュヒット | キャッシュ作成 |
|---|---|---|---|---|---|---|---|
0 - 32K chatAlibaba | 1,140/ 1M Tokens | 5,700/ 1M Tokens | $1.140 | $5.700 | $0.228 228/ 1M Tokens | $0.114 114/ 1M Tokens | $1.425 1,425/ 1M Tokens |
32K - 128K chatAlibaba | 2,280/ 1M Tokens | 11,400/ 1M Tokens | $2.280 | $11.400 | $0.456 456/ 1M Tokens | $0.228 228/ 1M Tokens | $2.850 2,850/ 1M Tokens |
128K - 256K chatAlibaba | 2,850/ 1M Tokens | 14,250/ 1M Tokens | $2.850 | $14.250 | $0.570 570/ 1M Tokens | $0.285 285/ 1M Tokens | $3.563 3,563/ 1M Tokens |

GLM-5 is Z.ai’s flagship open-source foundation model engineered for complex systems design and long-horizon agent workflows. Built for expert developers, it delivers production-grade performance on large-scale programming tasks, rivaling leading closed-source models. With advanced agentic planning, deep backend reasoning, and iterative self-correction, GLM-5 moves beyond code generation to full-system construction and autonomous execution.
| モダリティ | 入力クレジット | 出力クレジット | 入力価格 | 出力価格 | 暗黙的キャッシュヒット | 明示的キャッシュヒット | キャッシュ作成 |
|---|---|---|---|---|---|---|---|
Standard chatZ.ai | 970/ 1M Tokens | 3,104/ 1M Tokens | $0.970 | $3.104 | $0.194 194/ 1M Tokens | $0.194 194/ 1M Tokens | -- |

Supports text, single-image and multi-image inputs, and enables the generation of image sets
| モダリティ | クレジット / 単位 | 料金 (USD) |
|---|---|---|
Standard imageBytePlus | 34/ Image | $0.034 |

Flagship multimodal foundation model supporting text, image, and video inputs with text output. Features a 1M-token context window via MiniMax Sparse Attention (MSA), cutting per-token compute to ~1/20 of previous gen at full context. Excels at long-horizon agentic work, coding, and tool use. Native multimodal training from step zero ensures deep semantic alignment. Scores 59.0% on SWE-Bench Pro and 83.5 on BrowseComp, surpassing Opus 4.7
| モダリティ | 入力クレジット | 出力クレジット | 入力価格 | 出力価格 | 暗黙的キャッシュヒット | 明示的キャッシュヒット | キャッシュ作成 |
|---|---|---|---|---|---|---|---|
0 - 512K chatMiniMax | 600/ 1M Tokens | 2,400/ 1M Tokens | $0.600 | $2.400 | $0.120 120/ 1M Tokens | $0.120 120/ 1M Tokens | -- |
Over 512K chatMiniMax | 1,200/ 1M Tokens | 4,800/ 1M Tokens | $1.200 | $4.800 | $0.240 240/ 1M Tokens | $0.240 240/ 1M Tokens | -- |

Powered by Qwen3, this is a powerful Coding Agent that excels in tool calling and environment interaction to achieve autonomous programming. It combines outstanding coding proficiency with versatile general-purpose abilities.
| モダリティ | 入力クレジット | 出力クレジット | 入力価格 | 出力価格 | 暗黙的キャッシュヒット | 明示的キャッシュヒット | キャッシュ作成 |
|---|---|---|---|---|---|---|---|
0 - 32K chatAlibaba | 1,000/ 1M Tokens | 5,000/ 1M Tokens | $1.000 | $5.000 | $0.200 200/ 1M Tokens | $0.100 100/ 1M Tokens | $1.250 1,250/ 1M Tokens |
32K - 128K chatAlibaba | 1,800/ 1M Tokens | 9,000/ 1M Tokens | $1.800 | $9.000 | $0.360 360/ 1M Tokens | $0.180 180/ 1M Tokens | $2.250 2,250/ 1M Tokens |
128K - 256K chatAlibaba | 3,000/ 1M Tokens | 15,000/ 1M Tokens | $3.000 | $15.000 | $0.600 600/ 1M Tokens | $0.300 300/ 1M Tokens | $3.750 3,750/ 1M Tokens |
256K - 1M chatAlibaba | 6,000/ 1M Tokens | 60,000/ 1M Tokens | $6.000 | $60.000 | $1.200 1,200/ 1M Tokens | $0.600 600/ 1M Tokens | $7.500 7,500/ 1M Tokens |

The Qwen3.6 native vision-language Plus series models demonstrate exceptional performance on par with the current state-of-the-art models, with a significant improvement in overall results compared to the 3.5 series. The models have been markedly enhanced in code-related capabilities such as agentic coding, front-end programming, and Vibe coding, as well as in multi-modal general object recognition, OCR, and object localization.
| モダリティ | 入力クレジット | 出力クレジット | 入力価格 | 出力価格 | 暗黙的キャッシュヒット | 明示的キャッシュヒット | キャッシュ作成 |
|---|---|---|---|---|---|---|---|
0 - 256K chatAlibaba | 475/ 1M Tokens | 2,850/ 1M Tokens | $0.475 | $2.850 | -- | $0.048 48/ 1M Tokens | $0.594 594/ 1M Tokens |
256K - 1M chatAlibaba | 1,900/ 1M Tokens | 5,700/ 1M Tokens | $1.900 | $5.700 | -- | $0.190 190/ 1M Tokens | $2.375 2,375/ 1M Tokens |

Supports generating video with audio from text and images, and supports first and last frames.
| モダリティ | クレジット / 単位 | 料金 (USD) | ≈秒あたりの価格 (USD) |
|---|---|---|---|
Without Audio videoBytePlus | 1,200/ 1M Tokens | $1.200 | $0.026 |
With Audio videoBytePlus | 2,280/ 1M Tokens | $2.280 | $0.050 |
* 推定平均ビデオ = 720p で 5 秒、または画像 1024x1024 に基づいています。実際の出力コストは、特定のモデルパラメータ、解像度倍率、プロンプトの複雑さによって異なる場合があります。