新規登録で 100 クレジットを無料プレゼント、今すぐAIアプリを探索・構築無料受取
MiniMax-H3

MiniMax H3 / テキストから動画

Commercial
ID: MiniMax-H3

Open-weight general-purpose multimodal video model. Unifies text, image, video, and audio understanding in a single context window, generating up to 2K resolution, 15-second clips with native stereo audio at 24fps. Supports multimodal reference inputs: up to 9 images, 3 videos, and 3 audio clips (12 total references) per generation. Features first-frame, last-frame, and full reference modes with conversational editing capabilities.

モデルタイプ:
入力$0.04/ Image(40 クレジット)
出力$0.08/ Second(80 クレジット)
Input Prompt
0
動画生成

Video Playground Ready

左側のパラメータパネルでプロンプトを入力し、設定を行って Generate をクリックしてください。

MiniMax H3 Hero Visual
MiniMax H3 • Multimodal Reference Video

MiniMax H35-Mode Cinematic Video with 2K Output

MiniMax H3 is a multimodal video foundation model with five generation modes — text, image, end-frame, start-and-end-frame, and reference conditioning — plus native 2K resolution and up to 15-second clips.

5 Generation Modes
768P / 2K Native Resolution
4–15 Second Clips
Reference Conditioning + 6 Ratios
ByteDance Seed Foundation Architecture
Commercial License & Enterprise SLA
5
Generation Modes
Text → Reference
2K
Max Resolution
Also 768P
4–15s
Clip Duration
1-second steps
6
Aspect Ratios
Plus adaptive
Capability Highlights

Multimodal Video Architecture

One model for every conditioning path — from pure text to multi-reference cinematography.

Five Conditioning Modes

Text-to-video, image-to-video, end-frame, start-and-end-frame, and multimodal reference-to-video — switch modes without changing providers.

T2VI2VEnd FrameReference
Five modes montage
2K cinema still

Native 2K Resolution

Ship 768P for drafts and social, or promote hero placements to 2K without a separate upscaler pass.

768P2KHero Ready

Start & End Frame Bridges

Land exactly on a target end frame for match cuts, product morphs, and seamless scene transitions.

Start FrameEnd FrameMatch Cut
Frame bridge concept
Multi ratio canvas

Cinematic Ratio Suite

Text-to-video supports 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16; reference mode also allows adaptive framing.

6 Ratios21:9 Ultra-WideAdaptive
How It Works

How It Works

From brief to finished multimodal clip.

01

Pick a Mode

Start from text, a still, an end frame, both frames, or multimodal references.

02

Set Canvas & Length

Choose 768P or 2K, aspect ratio, and 4–15 second duration for the placement.

03

Condition with References

In reference mode, feed stills or motion cues that lock identity and style.

04

Deliver & Chain

Pull the finished clip and bridge into the next shot with end-frame continuity.

H3 Production Domains

Where multimodal control replaces stitched pipelines.

Film

Brand Film Previs

Reference-conditioned previs before full live-action shoots.

Reference
Commerce

Product Morph Ads

End-frame and dual-frame transitions for packshot reveals.

Frame Bridge
Marketing

Ultra-Wide Storytelling

Native 21:9 cinematic boards for hero brand placements.

21:9
Content

Social Vertical Cuts

9:16 clips up to 15 seconds for Reels and TikTok.

9:16
Best Practices

Prompt & Usage Tips

Get cleaner first-pass H3 generations.

Name the mode intent early

If you need a bridge, mention the end composition in the prompt so the model aims for that landing frame.

2K only for heroes

Draft at 768P, promote only final hero placements to 2K to control spend.

Reference with one clear subject

Reference mode is strongest when each image has a single clear job — identity, style, or scene.

Developer Quickstart

H3 Multimodal Quickstart

Submit text or reference-driven clips with native 2K support.

curl -X POST "https://api.powertokens.ai/v1/videos" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "MiniMax-H3",
  "prompt": "Cinematic tracking shot through a neon cyberpunk city in heavy rain.",
  "seconds": "5",
  "size": "1080p",
  "ratio": "16:9",
  "resolution": "2K",
  "duration": 8
}'

Technical Specifications

Confirmed parameters and runtime execution protocols.

Provider & Model ID
MiniMax • MiniMax-H3
Modes
Text-to-video, Image-to-video, End-frame, Start & end frame, Reference-to-video
Resolutions
768P (default) or 2K
Duration
4 to 15 seconds (default 4)
Aspect Ratios (T2V)
21:9, 16:9, 4:3, 1:1, 3:4, 9:16 (default 21:9)
Aspect Ratios (I2V)
Adaptive framing
Reference Mode Ratios
adaptive + 21:9 / 16:9 / 4:3 / 1:1 / 3:4 / 9:16
Billing Unit
Per second of generated video
Tags
Video to Video supported
API Endpoint
POST /v1/videos

料金詳細

このモデルの実際の課金は、API リクエストで渡される特定のパラメータに基づいて動的に計算されます。以下は具体的な組み合わせとそれに対応する料金です:

: 各リクエストの最初の5枚の入力画像は無料です。それ以降の入力画像は表に表示されている入力価格に基づいて課金されます。

課金ルール:入力動画と出力動画の両方が課金対象となり、動画の秒数単位で計算されます。課金対象時間 = 入力動画の長さ + 出力動画の長さ。

768P
入力
$0.040(40 / Image)
出力
$0.080(80 / Second)
2K
入力
$0.040(40 / Image)
出力
$0.130(130 / Second)
Ecosystem Models

Recommended Related Models

Explore complementary video and multimodal models with your unified API key.

Browse All Models
Seedance 2.5
videoCommercial
Seedance 2.5

Seedance 2.5

dreamina-seedance-2-5-260628

🔥 LIMITED-TIME OFFER: 1080P at 20% OFF! 🔥 ByteDance's latest flagship video generation model, built for longer-form storytelling and production-ready output. Generates up to 30 seconds of continuous, cinematic video with native audio sync in a single pass . Accepts up to 50 multimodal references (images, videos, audio, character sheets, storyboards) for precise scene, character, and motion consistency . Features localized region editing to fix specific areas without full regeneration……

Text to VideoImage to Video
Wan 3.0
videoCommercial
Wan 3.0

Wan 3.0

wan3.0-video

Wan3.0-Video is an all-in-one video generation model unified support for multiple creative capabilities, including reference, editing, replication, and driving. It generates videos up to 30 seconds with omni-modal reference, and can parse files, web pages and complex images. With production-grade character consistency and lifelike visuals and sound, it delivers an immersive audiovisual experience.

Text to VideoImage to Video
Seedance 2.0 Mini
videoCommercial
Seedance 2.0 Mini

Seedance 2.0 Mini

dreamina-seedance-2-0-mini-260615

Lightweight, cost-efficient video model from ByteDance, optimized for speed and high-volume content creation. Supports text-to-video, image-to-video, and reference-based generation with up to 12 references (6 images, 3 audio, 3 video). Delivers faster generation and lower credit consumption than Seedance 2.0, with strong motion quality and character consistency. Ideal for social media content, product videos, AI short dramas, and rapid creative iteration

Text to VideoImage to Video
Seedance 2.0
videoCommercial
Seedance 2.0

Seedance 2.0

dreamina-seedance-2-0-260128

Generate videos from reference images, videos, and audio; edit videos; extend videos; generate videos from start and end frames

Text to VideoImage to Video

Frequently Asked Questions

Everything you need to know before integrating this model.

It supports five conditioning paths in one model: text-to-video, image-to-video, end-frame, start-and-end-frame, and reference-to-video.

Start Building with MiniMax H3 Today

Create an account in seconds to receive 100 free credits and start generating immediately. No credit card or upfront contract required.