- December 11, 2025
- 3 min read
GLM-4.6, MiniMax-M2, and Ministral-3 Now Available on FriendliAI
- FriendliAI now supports GLM-4.6, MiniMax-M2, and Ministral-3, offering these high-capability models via OpenAI-compatible APIs on both Serverless and Dedicated Endpoints.
- These models provide advanced reasoning and long-context support (up to 200k tokens for GLM-4.6), making them ideal for complex agentic workflows and large-scale RAG applications.
- By deploying on FriendliAI, users gain optimized GPU kernels, 50% cost reductions via online quantization, and enterprise-grade reliability with 99.99% uptime SLAs.

As part of our ongoing commitment to supporting the latest frontier models, we’re excited to announce full support for GLM-4.6, MiniMax-M2, and Ministral-3 across Serverless and Dedicated Endpoints.
The open-source frontier is evolving faster than ever, and GLM-4.6, MiniMax-M2, and Ministral-3 series stand out as three of the most capable models in reasoning, long-context understanding, and tool usage.
All these models are now available on FriendliAI through our OpenAI-compatible API, with full tool-calling support for agentic workflows across Serverless Endpoints and Dedicated Endpoints. For teams building agentic AI systems, this means you can deploy these highly efficient, cutting-edge models with FriendliAI’s signature performance, reliability, and cost efficiency.
Why These New Models Matter
GLM-4.6: Longer Context and Superior Reasoning
- One of the most capable open large-context models available.
- Supports a 200k-token context window.
- Strong at reasoning, coding, and tool use.
- Excels in tasks requiring:
- Sustained long-context understanding
- Structured outputs
- Multi-step reasoning
- Analysis of long documents
- Coordination of complex workflows
- Efficient and practical for production systems needing transparency, customization, and scalability.
- Well-suited for tool-using, RAG, and agentic applications requiring both intelligence and operational flexibility.
MiniMax‑M2: Compact, Efficient, and Scalable
- Sparse Mixture-of-Experts (MoE) model: 230B total parameters, 10B active.
- Delivers state-of-the-art performance on reasoning, coding, and agentic tasks while staying efficient.
- Supports a 128k-token context window.
- Strong at tasks such as:
- Handling long documents or full codebases
- Multi-file edits
- Code generation and fixing
- Long-horizon planning and tool-using pipelines (shell, browser, code runner, retrieval, etc.)
- Offers cutting-edge capability with lower latency, lower cost, and easier scaling than dense frontier models.
Ministral-3 Series: Efficient, Multimodal Reasoning with Vision
- Optimized transformer-based model designed for efficient reasoning and instruction following.
- Delivers strong performance on reasoning, coding, and tool-using tasks with a focus on stability and control.
- Designed for low-latency, cost-efficient deployment compared to large dense frontier models.
- Strong at tasks such as:
- Multi-step reasoning with structured outputs
- Tool calling and function-driven workflows
- Agentic pipelines requiring predictable behavior
- RAG and automation systems operating at scale
- Offers a practical balance of intelligence, efficiency, and reliability for real-world production environments.
Why Run These Models on FriendliAI?
The models like GLM-4.6 and MiniMax-M2 excel at reasoning, coding, and tool use, real-world workflows require more than raw model power. FriendliAI provides the orchestration, integration, and scalability needed to turn these capabilities into practical, production-ready solutions.
By using FriendliAI, you get:
High Throughput & Low Latency
- Advanced batching & scheduling algorithms
- Optimized GPU kernels
- 50% cost reduction with online quantization
- Continuous batching for high-demand workloads
Production-Grade Reliability
- Customizable request-based autoscaling
- Full logs & metrics observability
- Enterprise-grade SLAs
- Globally geo-distributed infrastructure
- SOC2 certified
Flexible Deployment Options
- Serverless Endpoints: Use instantly with no infrastructure setup
- Dedicated Endpoints: Exclusive access to high-demand GPUs
- Container: Run on your public cloud or on-prem clusters
No matter how you deploy, you get FriendliAI’s hallmark scalability, reliability, speed, and cost efficiency.
Get Started with GLM-4.6 & MiniMax-M2 on FriendliAI
Try on Serverless (Instant Access)
Explore all available models immediately on Friendli Suite, no setup needed.
Deploy a Dedicated Endpoint
Choose your GPU, deploy in minutes, and scale to production workloads effortlessly.
Build Agents with Tools
Use our OpenAI-compatible API to connect your tools and power agentic workflows with reliability and speed.
👉 Try the models today on FriendliAI and start building your next generation of intelligent agents.
Written by
FriendliAI Tech & Research
Share
General FAQ
What is FriendliAI?
FriendliAI is the Frontier Inference Cloud for Agents, delivering high throughput, low latency, and reliability at scale for agentic workloads. Through vertically optimized inference infrastructure, it delivers 2–5× faster output token speed and a 99.99% uptime SLA for high-volume production traffic.
How does FriendliAI reduce inference costs?
FriendliAI reduces inference costs through higher GPU utilization and optimized inference performance. FriendliAI's patented continuous batching technique, along with quantization, speculative decoding, KV cache offloading, multi-LoRA serving, and autoscaling, helps you serve more tokens with fewer GPUs, lowering your infrastructure costs without sacrificing performance.
Why should I choose FriendliAI over other inference providers?
FriendliAI is built for production AI agents, combining speed, reliability, and efficiency at scale. It delivers low-latency streaming, reliable long-context inference, and robust tool calling without compromising stability. According to independent OpenRouter benchmarks, FriendliAI consistently ranks among the top providers for throughput, latency, and reliability across leading open-weight models. See why customers choose FriendliAI
Which open-weight models does FriendliAI support?
Run today’s frontier open-weight models—including GLM, MiniMax, Kimi, DeepSeek, Qwen, Gemma, and more—with a simple API call. FriendliAI Model API gives you instant access to the latest models with optimized inference performance for production workloads. Explore models and pricing
How do I get started?
Getting started takes just a few minutes. [1] Sign up for FriendliAI, [2] Generate your API key, and [3] Make your first inference request with frontier open-weight models.
Still have questions?
If you want a customized solution for that key issue that is slowing your growth, support@friendli.ai or click Talk to an engineer — our engineers (not a bot) will reply within one business day.

