新用户免费领取 100 积分,即刻探索与构建您的 AI 应用免费领取
MiniMax-M3

MiniMax M3 / 对话

Commercial
ID: MiniMax-M3

Flagship multimodal foundation model supporting text, image, and video inputs with text output. Features a 1M-token context window via MiniMax Sparse Attention (MSA), cutting per-token compute to ~1/20 of previous gen at full context. Excels at long-horizon agentic work, coding, and tool use. Native multimodal training from step zero ensures deep semantic alignment. Scores 59.0% on SWE-Bench Pro and 83.5 on BrowseComp, surpassing Opus 4.7

输入$0.6/ 1M Tokens(600 积分)
输出$2.4/ 1M Tokens(2400 积分)
0.7
0.95

Press Enter to add, Backspace to remove

Chat 对话

Playground 对话

与 AI 模型开启对话,您可以询问任何问题。

0

AI 生成结果的准确性可能有所不同。

MiniMax M3 • Production Chat

MiniMax M3Reliable Streaming Chat Completions

MiniMax M3 is a production chat model with streaming defaults, temperature and top-p sampling, stop sequences, and configurable max output tokens.

Streaming Default On
Temperature / Top-p
Stop Sequences
Max Output Tokens
ByteDance Seed Foundation Architecture
Commercial License & Enterprise SLA
MiniMax M3 Hero

At a Glance

M3
MiniMax Chat
Production tier
Stream
Default On
SSE tokens
0.7
Default Temp
Balanced tone
Stop
Sequences
Bounded output
Capability Highlights

M3 Chat Capabilities

Straightforward chat completions for product surfaces.

01

Streaming by Default

stream defaults to true so tokens arrive as they are generated for chat UIs.

stream: trueSSE
Streaming Default
02

Sampling Controls

Temperature default 0.7 and top-p default 0.95 for tone and diversity control.

temperature: 0.7top_p: 0.95
Sampling Controls
03

Stop Sequences

Bound generation so parsers and UIs never see trailing chatter.

stopBounded Output
Stop Sequences
04

Max Output Tokens

Cap completion length for cost control and predictable latency.

max_tokensCost Control
Token Caps
How It Works

How It Works

Simple chat completions path.

01

POST Messages

Send system and user turns in chat completions schema.

02

Stream Tokens

Consume SSE for immediate UI paint.

03

Bound Output

Use stop sequences and max tokens for parser-safe replies.

04

Tune Sampling

Adjust temperature and top-p for brand voice.

M3 Product Domains

Where dependable chat is the product.

CX

Customer Chat

Streaming support replies with stop-bounded answers.

Chat
Content

Content Assist

Drafts, rewrites, and summaries.

Rewrite
Ops

Ops Automation

Ticket drafting and internal Q&A.

Ops
Search

Search Answers

Grounded replies over retrieved context.

RAG
Best Practices

Prompt Tips

Keep M3 responses reliable.

System message sets the contract

Put role, constraints, and output format in the system message.

Prefer streaming UX

Keep stream on for chat surfaces; disable only for batch jobs.

Stop for schemas

Bound generation when you will parse JSON or field lists.

Developer Quickstart

MiniMax M3 Quickstart

Streaming production chat.

curl -X POST "https://api.powertokens.ai/v1/chat/completions" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "MiniMax-M3",
    "messages": [
      {"role": "user", "content": "Draft a concise out-of-office reply for a client email."}
    ],
    "stream": true,
    "temperature": 0.7
  }'

Technical Specifications

Confirmed parameters and runtime execution protocols.

Provider & Model ID
MiniMax • MiniMax-M3
Streaming
Streaming enabled by default
Sampling
Temperature (default 0.7), top_p (default 0.95)
Output Limits
max_tokens optional
Stop Sequences
stop array supported
Protocol
Chat completions
Billing
Token-based
API Endpoint
POST /v1/chat/completions
Vendor
MiniMax
Family
M3

价格详情

此模型的实际计费根据您在 API 请求中传入的具体参数动态计算。以下是具体的组合及其对应的价格:

0 - 512K
输入价格
$0.600(600 / 1M Tokens)
输出价格
$2.400(2,400 / 1M Tokens)
隐式缓存命中
$0.120(120/ 1M Tokens)
显式缓存命中
$0.120(120/ 1M Tokens)
创建缓存
--
Over 512K
输入价格
$1.200(1,200 / 1M Tokens)
输出价格
$4.800(4,800 / 1M Tokens)
隐式缓存命中
$0.240(240/ 1M Tokens)
显式缓存命中
$0.240(240/ 1M Tokens)
创建缓存
--
Same Channel

Models from the Same Channel

Explore complementary models and alternative versions from the same provider channel.

MiniMax M2.7
chatCommercial
MiniMax M2.7

MiniMax M2.7

MiniMax-M2.7

MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent collaboration, enabling it to plan, execute, and refine complex tasks across dynamic environments.

Chat
MiniMax M2.7 Highspeed
chatCommercial
MiniMax M2.7 Highspeed

MiniMax M2.7 Highspeed

MiniMax-M2.7-highspeed

A high-speed variant of MiniMax’s flagship M2.7 LLM, delivering 100 TPS—~60% faster than standard M2.7—with identical top-tier quality. It features a 204,800-token context window, near-Opus-level SWE performance, and recursive self-improvement. Optimized for low-latency coding, office tasks, and agent workflows, it balances extreme speed, reliability, and cost-efficiency for high-throughput real-world applications。

Chat
MiniMax M2.5
chatCommercial
MiniMax M2.5

MiniMax M2.5

MiniMax-M2.5

MiniMax-M2.5 is a SOTA large language model designed for real-world productivity. Trained in a diverse range of complex real-world digital working environments, M2.5 builds upon the coding expertise of M2.1 to extend into general office work, reaching fluency in generating and operating Word, Excel, and Powerpoint files, context switching between diverse software environments, and working across different agent and human teams

Chat
MiniMax M2.5 Highspeed
chatCommercial
MiniMax M2.5 Highspeed

MiniMax M2.5 Highspeed

MiniMax-M2.5-highspeed

MiniMax-M2.5-highspeed is an ultra-fast inference version of MiniMax M2.5. It maintains the full intelligent capability of the standard M2.5 model, featuring ultra-low latency, high throughput and outstanding cost performance. Built on advanced MoE architecture, it delivers rapid response speed while ensuring stable output quality. Optimized for high-concurrency business scenarios such as real-time dialogue, content generation and API service calls, it perfectly meets enterprise-level demands fo

Chat
Ecosystem Models

Recommended Related Models

Explore complementary video and multimodal models with your unified API key.

Browse All Models
GLM-5.3
chatCommercial
GLM-5.3

GLM-5.3

glm-5.3

Zhipu AI's flagship text model optimized for complex software engineering and long-horizon Agent tasks. Features a 1M-token context window with mandatory reasoning (3 levels: low/high/max). Coding capability improved 50% over GLM-5.2 on Z.ai Code Bench; scores SOTA on Terminal Bench 3.0. Excels in cybersecurity tasks (cyber vulnerability discovery)

ChatText Generation
Qwen3.8 Flash
chatCommercial
Qwen3.8 Flash

Qwen3.8 Flash

qwen3.8-flash

Alibaba's cost-efficient multimodal reasoning model. Supports text, image, and video inputs with text output. Features a native 1M-token context window for long documents, codebases, and agentic workflows. Excels in coding assistance, desktop interaction, chart analysis, and long-video understanding. Compatible with OpenAI/Anthropic protocols for seamless integration

ChatText Generation
GLM-5.2
chatCommercial
GLM-5.2

GLM-5.2

glm-5.2

Flagship text model purpose-built for long-horizon agentic workflows. Features a 1M context window supporting project-level engineering in a single session. Excels at autonomous coding: can complete development, testing, and multi-platform deployment from a single prompt. Top open-weight model per Artificial Analysis; #1 globally on Code Arena. MIT-licensed and Day-0 optimized for domestic AI chips

ChatText Generation
Qwen3 Max
chatCommercial
Qwen3 Max

Qwen3 Max

qwen3-max

Compared with the September 23, 2025 version, the newly upgraded Qwen-3 Max seamlessly integrates thinking and non-thinking modes, bringing an all-round obvious performance boost. Its thinking mode supports web search, web content extraction and code interpreter. It can conduct in-depth logical reasoning and call external tools to solve intricate problems more precisely

Chat

Frequently Asked Questions

Everything you need to know before integrating this model.

Yes. stream defaults to true.

Start Building with MiniMax M3 Today

Create an account in seconds to receive 100 free credits and start generating immediately. No credit card or upfront contract required.