Use with vLLM/SGLang
This model can be deployed efficiently using the vLLM and SGLang backends.
Evaluation
The model was evaluated on gsm8k benchmarks using the vllm framework.
Accuracy
Reproduction
The GSM8K result was obtained using the lm-evaluation-harness framework, based on the Docker image rocm/vllm-dev:nightly_main_20260603.
Install the lm-eval (Version: 0.4.12) in container first.
pip install lm-evalpip install lm-eval[api]
Launching server
VLLM_ROCM_USE_AITER=1 vllm serve amd/MiniMax-M2.5-NVFP4/ \ --tensor-parallel-size 2 \ --tool-call-parser minimax_m2 \ --reasoning-parser minimax_m2 \ --enable-auto-tool-choice \ --trust-remote-code
Evaluating model in a new terminal
lm_eval \ --model local-completions \ --model_args "model=amd/MiniMax-M2.5-NVFP4/,base_url=http://127.0.0.1:8000/v1/completions,tokenized_requests=False,tokenizer_backend=None,num_concurrent=32" \ --gen_kwargs temperature=1.0,top_p=0.95 \ --tasks gsm8k \ --num_fewshot 8 \ --batch_size 1
License
Modifications Copyright(c) 2026 Advanced Micro Devices, Inc. All rights reserved.