MiniMaxAI
MiniMax-M3
Available on FriendliAI
Run this model inference on single tenant GPU with unmatched speed and reliability at scale.
Model Details
Model Provider
MiniMaxAI
Model Tree
Input Modalities
Output Modalities
Supported Functionality
GLM-5.2 is live. #1 throughput on OpenRouter, pay-per-token on FriendliAI. Try it today ➜
MiniMaxAI
Available on FriendliAI
Run this model inference on single tenant GPU with unmatched speed and reliability at scale.
Model Details
Model Provider
MiniMaxAI
Model Tree
Input Modalities
Output Modalities
Supported Functionality
MiniMax M3 is a 427B MoE model (23B active) with native multimodality across text, image, and video input, plus computer use. It pairs coding and agentic capability with a 1M context window. MiniMax Sparse Attention cuts per-token compute to 1/20 of the prior generation at 1M context, yielding 9x prefilling and 15x decoding speedups. A toggleable thinking mode switches between deep reasoning and fast responses. Scores 59.0% on SWE-Bench Pro, 66.0% on Terminal-Bench 2.1.