신규 가입 시 100 무료 크레딧 증정, 지금 바로 AI 애플리케이션을 탐색하고 구축하세요무료 받기
MiniMax-H3

MiniMax H3 / 텍스트로 비디오

Commercial
ID: MiniMax-H3

Open-weight general-purpose multimodal video model. Unifies text, image, video, and audio understanding in a single context window, generating up to 2K resolution, 15-second clips with native stereo audio at 24fps. Supports multimodal reference inputs: up to 9 images, 3 videos, and 3 audio clips (12 total references) per generation. Features first-frame, last-frame, and full reference modes with conversational editing capabilities.

모델 유형:
입력$0.04/ Image(40 크레딧)
출력$0.08/ Second(80 크레딧)
Input Prompt
0
비디오 생성

Video Playground Ready

왼쪽 매개변수 패널에서 프롬프트를 입력하고 옵션을 설정한 후 Generate를 클릭하세요.

MiniMax H3 Hero Visual
MiniMax H3 • Multimodal Reference Video

MiniMax H35-Mode Cinematic Video with 2K Output

MiniMax H3 is a multimodal video foundation model with five generation modes — text, image, end-frame, start-and-end-frame, and reference conditioning — plus native 2K resolution and up to 15-second clips.

5 Generation Modes
768P / 2K Native Resolution
4–15 Second Clips
Reference Conditioning + 6 Ratios
ByteDance Seed Foundation Architecture
Commercial License & Enterprise SLA
5
Generation Modes
Text → Reference
2K
Max Resolution
Also 768P
4–15s
Clip Duration
1-second steps
6
Aspect Ratios
Plus adaptive
Capability Highlights

Multimodal Video Architecture

One model for every conditioning path — from pure text to multi-reference cinematography.

Five Conditioning Modes

Text-to-video, image-to-video, end-frame, start-and-end-frame, and multimodal reference-to-video — switch modes without changing providers.

T2VI2VEnd FrameReference
Five modes montage
2K cinema still

Native 2K Resolution

Ship 768P for drafts and social, or promote hero placements to 2K without a separate upscaler pass.

768P2KHero Ready

Start & End Frame Bridges

Land exactly on a target end frame for match cuts, product morphs, and seamless scene transitions.

Start FrameEnd FrameMatch Cut
Frame bridge concept
Multi ratio canvas

Cinematic Ratio Suite

Text-to-video supports 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16; reference mode also allows adaptive framing.

6 Ratios21:9 Ultra-WideAdaptive
How It Works

How It Works

From brief to finished multimodal clip.

01

Pick a Mode

Start from text, a still, an end frame, both frames, or multimodal references.

02

Set Canvas & Length

Choose 768P or 2K, aspect ratio, and 4–15 second duration for the placement.

03

Condition with References

In reference mode, feed stills or motion cues that lock identity and style.

04

Deliver & Chain

Pull the finished clip and bridge into the next shot with end-frame continuity.

H3 Production Domains

Where multimodal control replaces stitched pipelines.

Film

Brand Film Previs

Reference-conditioned previs before full live-action shoots.

Reference
Commerce

Product Morph Ads

End-frame and dual-frame transitions for packshot reveals.

Frame Bridge
Marketing

Ultra-Wide Storytelling

Native 21:9 cinematic boards for hero brand placements.

21:9
Content

Social Vertical Cuts

9:16 clips up to 15 seconds for Reels and TikTok.

9:16
Best Practices

Prompt & Usage Tips

Get cleaner first-pass H3 generations.

Name the mode intent early

If you need a bridge, mention the end composition in the prompt so the model aims for that landing frame.

2K only for heroes

Draft at 768P, promote only final hero placements to 2K to control spend.

Reference with one clear subject

Reference mode is strongest when each image has a single clear job — identity, style, or scene.

Developer Quickstart

H3 Multimodal Quickstart

Submit text or reference-driven clips with native 2K support.

curl -X POST "https://api.powertokens.ai/v1/videos" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "MiniMax-H3",
  "prompt": "Cinematic tracking shot through a neon cyberpunk city in heavy rain.",
  "seconds": "5",
  "size": "1080p",
  "ratio": "16:9",
  "resolution": "2K",
  "duration": 8
}'

Technical Specifications

Confirmed parameters and runtime execution protocols.

Provider & Model ID
MiniMax • MiniMax-H3
Modes
Text-to-video, Image-to-video, End-frame, Start & end frame, Reference-to-video
Resolutions
768P (default) or 2K
Duration
4 to 15 seconds (default 4)
Aspect Ratios (T2V)
21:9, 16:9, 4:3, 1:1, 3:4, 9:16 (default 21:9)
Aspect Ratios (I2V)
Adaptive framing
Reference Mode Ratios
adaptive + 21:9 / 16:9 / 4:3 / 1:1 / 3:4 / 9:16
Billing Unit
Per second of generated video
Tags
Video to Video supported
API Endpoint
POST /v1/videos

가격 상세

이 모델의 실제 요금은 API 요청에서 전달된 특정 매개변수를 기반으로 동적으로 계산됩니다. 아래는 구체적인 조합과 해당 가격입니다:

참고: 각 요청의 처음 5개 입력 이미지는 무료입니다. 이후 입력 이미지는 표에 표시된 입력 가격에 따라 청구됩니다.

요금 청구 규칙: 입력 비디오와 출력 비디오 모두 비디오 초당 요금이 청구되며, 청구 시간 = 입력 비디오 시간 + 출력 비디오 시간입니다.

768P
입력
$0.040(40 / Image)
출력
$0.080(80 / Second)
2K
입력
$0.040(40 / Image)
출력
$0.130(130 / Second)
Ecosystem Models

Recommended Related Models

Explore complementary video and multimodal models with your unified API key.

Browse All Models
Seedance 2.5
videoCommercial
Seedance 2.5

Seedance 2.5

dreamina-seedance-2-5-260628

🔥 LIMITED-TIME OFFER: 1080P at 20% OFF! 🔥 ByteDance's latest flagship video generation model, built for longer-form storytelling and production-ready output. Generates up to 30 seconds of continuous, cinematic video with native audio sync in a single pass . Accepts up to 50 multimodal references (images, videos, audio, character sheets, storyboards) for precise scene, character, and motion consistency . Features localized region editing to fix specific areas without full regeneration……

Text to VideoImage to Video
Wan 3.0
videoCommercial
Wan 3.0

Wan 3.0

wan3.0-video

Wan3.0-Video is an all-in-one video generation model unified support for multiple creative capabilities, including reference, editing, replication, and driving. It generates videos up to 30 seconds with omni-modal reference, and can parse files, web pages and complex images. With production-grade character consistency and lifelike visuals and sound, it delivers an immersive audiovisual experience.

Text to VideoImage to Video
Seedance 2.0 Mini
videoCommercial
Seedance 2.0 Mini

Seedance 2.0 Mini

dreamina-seedance-2-0-mini-260615

Lightweight, cost-efficient video model from ByteDance, optimized for speed and high-volume content creation. Supports text-to-video, image-to-video, and reference-based generation with up to 12 references (6 images, 3 audio, 3 video). Delivers faster generation and lower credit consumption than Seedance 2.0, with strong motion quality and character consistency. Ideal for social media content, product videos, AI short dramas, and rapid creative iteration

Text to VideoImage to Video
Seedance 2.0
videoCommercial
Seedance 2.0

Seedance 2.0

dreamina-seedance-2-0-260128

Generate videos from reference images, videos, and audio; edit videos; extend videos; generate videos from start and end frames

Text to VideoImage to Video

Frequently Asked Questions

Everything you need to know before integrating this model.

It supports five conditioning paths in one model: text-to-video, image-to-video, end-frame, start-and-end-frame, and reference-to-video.

Start Building with MiniMax H3 Today

Create an account in seconds to receive 100 free credits and start generating immediately. No credit card or upfront contract required.