
DeepSeek V4 Pro / 对话
Large-scale MoE model with 1.6T total parameters and 49B activated, supporting a 1M-token context window . Designed for advanced reasoning, coding, and long-horizon agent workflows . Top performance on GPQA Diamond (88.8%) and Terminal-Bench Hard (46.2%)
Press Enter to add, Backspace to remove
Playground 对话
与 AI 模型开启对话,您可以询问任何问题。
AI 生成结果的准确性可能有所不同。
DeepSeek V4 ProDeep Reasoning with Logprobs
DeepSeek V4 Pro offers thinking toggles, high or max reasoning effort, logprobs evaluation signals, and production sampling for demanding analytical workloads.

At a Glance
V4 Pro Reasoning Stack
Analytical depth with evaluation hooks.
High or Max Effort
Dial reasoning between high and max for multi-step proofs and hard analysis.

Thinking Toggle
Enable or disable deep thinking per request depending on task complexity.

Logprob Evaluation
logprobs and top_logprobs expose token probability signals for confidence scoring.

Production Sampling
Temperature, top-p, max tokens, and stop sequences for controlled outputs.

How It Works
Deep analysis in four steps.
Set Thinking & Effort
Enable thinking and choose high or max depth.
Stream or Wait
Stream for long analyses; single-shot for batch evals.
Collect Logprobs
Use probability signals for confidence scoring.
Sample Responsibly
Tune temperature and top-p for tone control.
V4 Pro Domains
Where analytical depth is the product.
Formal Reasoning
Max effort for proofs and multi-hop analysis.
Eval Pipelines
Logprobs for confidence and regression tests.
Research Synthesis
High effort for literature and market synthesis.
Agent Decisions
Reasoned tool choice with probability signals.
Prompt Tips
Get more from V4 Pro reasoning.
High and Max effort spend tokens more usefully when you request proof paths.
Place documents before instructions for more reliable following.
Everyday Q&A is faster at high effort or with thinking off.
DeepSeek V4 Pro Quickstart
High / max reasoning with logprobs.
curl -X POST "https://api.powertokens.ai/v1/chat/completions" \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-v4-pro",
"messages": [
{"role": "user", "content": "Walk through a careful proof of the AM-GM inequality for two variables."}
],
"stream": true,
"temperature": 0.7
}'Technical Specifications
Confirmed parameters and runtime execution protocols.
价格详情
此模型的实际计费根据您在 API 请求中传入的具体参数动态计算。以下是具体的组合及其对应的价格:
- 在工作日(周一至周五) 09:00 - 12:00, 14:00 - 18:00 (UTC+8) 期间发起的请求,应用 2x 的价格倍率。
| 模态 | 输入积分 | 输出积分 | 输入价格 | 输出价格 | 隐式缓存命中 | 显式缓存命中 | 创建缓存 |
|---|---|---|---|---|---|---|---|
| Standard | 660/ 1M Tokens | 1,980/ 1M Tokens | $0.660 | $1.980 | $0.022 22/ 1M Tokens | $0.022 22/ 1M Tokens | -- |
Recommended Related Models
Explore complementary video and multimodal models with your unified API key.


GLM-5.3
glm-5.3
Zhipu AI's flagship text model optimized for complex software engineering and long-horizon Agent tasks. Features a 1M-token context window with mandatory reasoning (3 levels: low/high/max). Coding capability improved 50% over GLM-5.2 on Z.ai Code Bench; scores SOTA on Terminal Bench 3.0. Excels in cybersecurity tasks (cyber vulnerability discovery)


Qwen3.8 Flash
qwen3.8-flash
Alibaba's cost-efficient multimodal reasoning model. Supports text, image, and video inputs with text output. Features a native 1M-token context window for long documents, codebases, and agentic workflows. Excels in coding assistance, desktop interaction, chart analysis, and long-video understanding. Compatible with OpenAI/Anthropic protocols for seamless integration


GLM-5.2
glm-5.2
Flagship text model purpose-built for long-horizon agentic workflows. Features a 1M context window supporting project-level engineering in a single session. Excels at autonomous coding: can complete development, testing, and multi-platform deployment from a single prompt. Top open-weight model per Artificial Analysis; #1 globally on Code Arena. MIT-licensed and Day-0 optimized for domestic AI chips


Qwen3 Max
qwen3-max
Compared with the September 23, 2025 version, the newly upgraded Qwen-3 Max seamlessly integrates thinking and non-thinking modes, bringing an all-round obvious performance boost. Its thinking mode supports web search, web content extraction and code interpreter. It can conduct in-depth logical reasoning and call external tools to solve intricate problems more precisely
Frequently Asked Questions
Everything you need to know before integrating this model.
Start Building with DeepSeek V4 Pro Today
Create an account in seconds to receive 100 free credits and start generating immediately. No credit card or upfront contract required.