- December 31, 2025
- 2 min read
K-EXAONE Is Now Available on Friendli Serverless Endpoints
- LG AI Research’s new K-EXAONE 236B-parameter Hybrid Attention MoE model is now available on Friendli Serverless Endpoints.
- Features a specialized architecture with 23B active parameters and Multi-Token Prediction (MTP) for 1.5x faster inference.
- This model outperforms competitors like Qwen3-235B and GPT-OSS 120B in complex logical tasks and long-context comprehension.
- Access K-EXAONE via OpenAI-compatible APIs with zero infrastructure overhead—free to use until January 28, 2026.

We’re excited to announce that K-EXAONE is now available on Friendli Serverless Endpoints, enabling developers and organizations to deploy and scale the LG AI Research’s new large-scale Hybrid Attention Mixture-of-Experts (MoE) model with ease.
Continuing our partnership, FriendliAI is yet again bringing K-EXAONE to production-ready environments with optimized inference, seamless APIs from day zero for free for a month, so anyone can try it out and start building right away.
About K-EXAONE
LGAI-EXAONE/K-EXAONE-236B-A23B is a 236B-parameter Hybrid Attention Mixture-of-Experts (MoE) model built for advanced reasoning, long-context understanding, and complex generative tasks. By routing data to specialized experts and combining this with selective full-attention layers, it delivers high efficiency without sacrificing performance.
The model excels at tasks requiring sustained reasoning and large-context comprehension, making it ideal for enterprise knowledge systems, multi-step workflows, and other applications that demand both depth and adaptability. Its mixture-of-experts design activates only the most relevant experts per token, while selective full-attention layers ensure key global context is captured.
In benchmark tests, K-EXAONE outperforms Qwen/Qwen3-235B-A22B-Thinking-2507 and openai/gpt-oss-120b, showing notable gains in reasoning and long-context tasks.

You can check out its detailed architecture specifications at https://huggingface.co/LGAI-EXAONE/K-EXAONE-236B-A23B. To highlight what sets K-EXAONE apart, the table below compares its key architectural characteristics with Qwen3-235B:
| K-EXAONE | Qwen3 235B | |
|---|---|---|
| num_hidden_layers | 48 | 94 |
| num_key_value_heads | 8 | 4 |
| sliding_window | O
(with full attention at every 4th layer) | X |
| num_experts | 128 + 1 | 128 |
| num_experts_per_tok | 8 + 1 | 8 |
Access K-EXAONE on Friendli Serverless Endpoints
With K-EXAONE on Friendli Serverless Endpoints, users can access the model through a fully managed, API-first experience:
- Zero infrastructure management, no GPU provisioning or tuning
- Automatic scaling, built for real-world traffic patterns
- Optimized inference, tuned for performance and cost efficiency
- OpenAI-compatible APIs, integrate quickly with existing stacks

Friendli Serverless Endpoints make it easy to take K-EXAONE from evaluation to production without operational overhead. Click here to access instantly.
Day-0 Support & Free Event
As part of this launch, FriendliAI will serve as the exclusive Day-0 support provider for K-EXAONE on Friendli Serverless Endpoints.
To help teams get started smoothly:
- Serverless Endpoint support will be provided free of charge for the first month, through January 28, 2026 PT
- Our team will assist with onboarding, deployment guidance, and performance optimization
This ensures developers can explore K-EXAONE’s capabilities confidently while building production-grade applications.
Start Building with K-EXAONE Today
K-EXAONE is now accessible via Friendli Serverless Endpoints, offering a streamlined path to accessing one of the latest Hybrid Attention MoE models.
If you’re looking to evaluate, prototype, or scale K-EXAONE in production, FriendliAI provides the infrastructure and support to help you move fast without compromise.
Ready to try it out for free? Click here (limited offer until Jan 28 PT)!
Written by
FriendliAI Tech & Research
Share
General FAQ
What is FriendliAI?
FriendliAI is the Frontier Inference Cloud for Agents, delivering high throughput, low latency, and reliability at scale for agentic workloads. Through vertically optimized inference infrastructure, it delivers 2–5× faster output token speed and a 99.99% uptime SLA for high-volume production traffic.
How does FriendliAI reduce inference costs?
FriendliAI reduces inference costs through higher GPU utilization and optimized inference performance. FriendliAI's patented continuous batching technique, along with quantization, speculative decoding, KV cache offloading, multi-LoRA serving, and autoscaling, helps you serve more tokens with fewer GPUs, lowering your infrastructure costs without sacrificing performance.
Why should I choose FriendliAI over other inference providers?
FriendliAI is built for production AI agents, combining speed, reliability, and efficiency at scale. It delivers low-latency streaming, reliable long-context inference, and robust tool calling without compromising stability. According to independent OpenRouter benchmarks, FriendliAI consistently ranks among the top providers for throughput, latency, and reliability across leading open-weight models. See why customers choose FriendliAI
Which open-weight models does FriendliAI support?
Run today’s frontier open-weight models—including GLM, MiniMax, Kimi, DeepSeek, Qwen, Gemma, and more—with a simple API call. FriendliAI Model API gives you instant access to the latest models with optimized inference performance for production workloads. Explore models and pricing
How do I get started?
Getting started takes just a few minutes. [1] Sign up for FriendliAI, [2] Generate your API key, and [3] Make your first inference request with frontier open-weight models.
Still have questions?
If you want a customized solution for that key issue that is slowing your growth, support@friendli.ai or click Talk to an engineer — our engineers (not a bot) will reply within one business day.

