Quantized Model Summary
Table with columns: Item, Value| Item | Value |
|---|
| Base model | MiniMaxAI/MiniMax-M3 |
| Quantized repo | Pilcothink/MiniMax-M3-int4-AutoRound |
| Local source folder | MiniMax-M3-W4A16-RTN |
| Quantization method | AutoRound |
| Quantization format | int4 / W4A16 |
| Weight precision | 4-bit |
| Activation precision | 16-bit |
| Base model relation | quantized |
| License | minimax-community, inherited from upstream model |
Upload Status
The model card has been uploaded first.
The quantized model weight files are still being uploaded and may not be available yet.
Notes
This repository is intended to host int4 AutoRound quantized weights for MiniMax-M3.
Benchmark numbers, runtime requirements, local deployment notes, and usage examples in the upstream section below describe the original MiniMaxAI/MiniMax-M3 model unless explicitly stated otherwise.
Attribution and License
This model is a quantized derivative of MiniMaxAI/MiniMax-M3, which is released under the minimax-community license.
The original model card is preserved below for attribution and reference.
============================================================
Original Upstream Model Card
============================================================
The following section is copied from the upstream model card of MiniMaxAI/MiniMax-M3.
It describes the original model, not this int4 AutoRound quantized release.
MiniMax-M3 is a native multimodal model with 1M context. It has ~428B parameters and ~23B activated parameters.
Highlights:
- Native Multimodality: M3 undergoes mixed-modality training from the very first step, enabling deeper semantic fusion across text, image, and video.
- Context Scaling via Sparse Attention: M3 introduces MiniMax Sparse Attention (MSA) to improve long context efficiency. M3 delivers 9× prefill and 15× decode speedups compared to M2 at 1M context, reducing per-token compute to 1/20.
- Coding & Cowork Capability: M3 achieves frontier-level performance across long-horizon agentic benchmarks, excelling in both coding and cowork.
MiniMax Sparse Attention (MSA)
M3 is powered by MiniMax Sparse Attention (MSA), a high-performance sparse attention operator designed for million-token contexts. Compared with GQA, MSA dramatically reduces the attention compute and memory footprint while preserving model quality.
📄 Read the technical report: arXiv:2606.13392 · Hugging Face Papers
How to Use
M3 supports three reasoning modes through the thinking parameter:
enabled — Reasoning is always enabled.
adaptive — M3 automatically determines when additional reasoning is beneficial.
disabled — Reasoning is disabled to minimize latency and maximize throughput.
Local Deployment
Download the model:
hf download MiniMaxAI/MiniMax-M3 --local-dir MiniMax-M3
We recommend the following inference frameworks (listed alphabetically) to serve the model:
Inference Parameters
We recommend the following parameters for best performance: temperature=1.0, top_p=0.95, top_k=40.
Contact us at model@minimax.io.