All AI models
Open-weightVerified August 11, 2026

MiniMax

MiniMax M3

MiniMax M3 is a strong open-weight candidate for teams that value a long context window, native multimodal input, agentic coding, and the ability to decide how much reasoning a request deserves. Start on the API, then evaluate self-hosting only if data control or sustained scale justifies the operational work.

License
Open-weight release, verify model-card terms
Context
1M tokens
Modalities
Text · Image · Video input
Parameters
428B total, 23B active

What the license means

MiniMax publishes the M3 weights and local deployment instructions. The official repository positions it as open-weight; confirm the current model-card terms before redistribution, offering it as a service, or building a commercial derivative.

Open-weight, open source, and no-cost API access mean different things. Always review the linked terms and your own data, hosting, and compliance requirements before shipping.

API and operating cost

Input / 1M
$0.30
Output / 1M
$1.20

Standard API price for input windows up to 512K tokens. The official long-context tier costs more.

Check current provider price

Where it fits

  • Agentic coding
  • Large codebases and document sets
  • Image and video-aware automation

Where it does not fit

  • Small local machines
  • Simple chat where 1M context adds no value
  • Teams that have not reviewed the current model-card terms

Local deployment reality

MiniMax documents SGLang, vLLM, Transformers, KTransformers, and Unsloth. Its 428B total parameter footprint makes managed access the practical starting point for most builders.

SGLangvLLMTransformersKTransformersUnsloth

Practical setup notes

  • Use adaptive reasoning before defaulting every request to the most expensive thinking setting.
  • Keep static instructions at the beginning of a prompt to improve cache reuse.

Capability notes

Reasoning
Enabled, adaptive, or disabled
Tools and output
Agent and computer-use workflows through MiniMax products and API

Limitations to plan around

  • Official pricing changes above 512K input tokens.
  • Self-hosting requires substantial GPU capacity despite sparse activation.

Continue comparing

Related model decisions

See all reviewed models
Editorial note: This is a dated decision page, not a benchmark leaderboard or a substitute for legal, security, or infrastructure review. Model and pricing facts can change after August 11, 2026; use the linked primary sources before committing budget or customer data.