- August 11, 2026
- 4 min read
NVIDIA Nemotron 3.5 Lightning on FriendliAI: Built for Agent Harnesses, Ready for Production, Available Day-0
- NVIDIA Nemotron 3.5 Lightning is a 30B hybrid MoE with 3B active parameters, distilled from NVIDIA's frontier Nemotron 3 Ultra
- Built for popular agent harnesses, it delivers leading accuracy for coding, tool calling, instruction following, and multi-turn workflows, with up to 4× higher throughput helping agents complete specialized tasks faster
- FriendliAI supports Nemotron 3.5 Lightning from Day 0 — available now on FriendliAI Dedicated Endpoints

Starting today, developers can run NVIDIA Nemotron 3.5 Lightning on FriendliAI Dedicated Endpoints. A fully customizable open 30B MoE model for powering always-on agents, it pairs leading agentic accuracy with the fastest token generation in its class, helping agents complete specialized tasks faster.
Agentic AI runs on a system of models
Always-on agents complete complex, multi-step tasks autonomously: researching, deciding, and acting without constant human input. Every step calls a language model, but not every step needs the same model. Complex reasoning and orchestration demand frontier-level capability, while high-volume, repetitive, domain-specific tasks are better handled by specialized models. Agentic AI needs systems of models that route each step to the right model for the task.
A fully customizable open model for always-on agents
NVIDIA Nemotron 3.5 Lightning is built for that high-volume specialized layer. A 30B MoE with 3B active parameters, distilled from NVIDIA's frontier Nemotron 3 Ultra and developed with the Nemotron Coalition, it is a fully customizable open model for powering always-on agents:
- Model control — an open model, trained with open datasets, that enterprises can control, customize, and deploy while managing model behavior, data handling, and agent workflows
- Customizable for highest accuracy — trained for popular agent harnesses, it delivers frontier-level accuracy in its class and and really shines when post-trained for specialized workflows to achieve leading accuracy in enterprise domains
- Fastest task completion — up to 4× higher throughput help always-on agents complete more steps and finish specialized tasks faster
Under the hood, a hybrid MoE architecture with multi-token prediction and up to 1M context supports the long-running, multi-turn sessions agents actually run.
Built for repeatable, domain-specific work
Lightning fits repeatable workflows where agents process large volumes of domain-specific work: PR summarization and test triage in software development, document extraction and policy checks in financial services, alert enrichment and incident classification in security operations, network alarm triage in telecom, catalog and order workflows in retail, and long-running personal agents for email, calendar, and projects.
The common thread is high call volume on well-scoped tasks, exactly where a fast, accurate, customizable model earns its place in a system of models, and where inference efficiency compounds with every call.
Inference is where efficiency compounds
Inference is where the model meets production. FriendliAI delivers frontier-level performance, throughput, and lower cost of inference to complete agentic tasks, and the execution layer is exactly the workload our stack is built for. A single agent task can fan out into dozens of Lightning calls, so every millisecond and every token-per-dollar compounds. And because agent steps run in sequence, per-call latency adds up across the chain, FriendliAI's inference stack shave time off each step, shortening the entire task.
What FriendliAI brings to Nemotron 3.5 Lightning:
High-throughput MoE inference. FriendliAI's inference engine maximizes throughput per GPU through continuous batching and kernel-level optimizations, so Lightning's speed advantage carries through to production under bursty, high-volume agent traffic.
Low-latency streaming for multi-turn agents. Reliable long-context inference and low-latency streaming keep long-running sessions moving between steps, a natural fit for Lightning's 1M context and multi-turn training.
Robust tool calling. Tool calls sit at the heart of agent workflows. FriendliAI's reliable tool calling support means Lightning's agent-harness training translates into dependable behavior under real traffic.
Cost-efficient inference at scale. Specialized models only make economic sense when they're inexpensive enough to call constantly. FriendliAI delivers higher tokens-per-dollar than alternative OSS serving stacks, making it practical to route high-volume work to Lightning as often as agents need.
Production-grade reliability. Auto-scaling, operational monitoring, and a 99.99% uptime SLA, because an agent workflow fails when any step fails.
Put Nemotron 3.5 Lightning into production
Nemotron 3.5 Lightning is available on Friendli Dedicated Endpoints starting today.
1️⃣ Navigate to the dedicated endpoint creation page.
2️⃣ Choose the model: [nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16], [nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4], [nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-Base-BF16]
3️⃣ Click "Create"
Prerequisites
- A FriendliAI account
- A FriendliAPI Key from Friendli Suite settings
Install the package
Environment Setup
Example Code
For teams with specialized tools, policies, or domain data, Lightning's open weights and open training datasets provide a clear path to post-train the model for your own workflows, and deploy the customized version on FriendliAI.
Deploy NVIDIA Nemotron at Scale with FriendliAI
FriendliAI is proud to partner with NVIDIA to bring the Nemotron family of models to developers and businesses building production-ready, agentic AI systems.
NVIDIA Nemotron 3.5 Lightning gives agents a fast, accurate, and fully customizable model for the high-volume work inside every workflow — and FriendliAI is your platform to take it to production.
👉 [Run Nemotron 3.5 Lightning on FriendliAI] today and give your agents an execution layer that keeps up with them.
Written by
FriendliAI Tech & Research
Share
General FAQ
What is FriendliAI?
FriendliAI is the Frontier Inference Cloud for Agents, delivering high throughput, low latency, and reliability at scale for agentic workloads. Through vertically optimized inference infrastructure, it delivers 2–5× faster output token speed and a 99.99% uptime SLA for high-volume production traffic.
How does FriendliAI reduce inference costs?
FriendliAI reduces inference costs through higher GPU utilization and optimized inference performance. FriendliAI's patented continuous batching technique, along with quantization, speculative decoding, KV cache offloading, multi-LoRA serving, and autoscaling, helps you serve more tokens with fewer GPUs, lowering your infrastructure costs without sacrificing performance.
Why should I choose FriendliAI over other inference providers?
FriendliAI is built for production AI agents, combining speed, reliability, and efficiency at scale. It delivers low-latency streaming, reliable long-context inference, and robust tool calling without compromising stability. According to independent OpenRouter benchmarks, FriendliAI consistently ranks among the top providers for throughput, latency, and reliability across leading open-weight models. See why customers choose FriendliAI
Which open-weight models does FriendliAI support?
Run today’s frontier open-weight models—including GLM, MiniMax, Kimi, DeepSeek, Qwen, Gemma, and more—with a simple API call. FriendliAI Model API gives you instant access to the latest models with optimized inference performance for production workloads. Explore models and pricing
How do I get started?
Getting started takes just a few minutes. [1] Sign up for FriendliAI, [2] Generate your API key, and [3] Make your first inference request with frontier open-weight models.
Still have questions?
If you want a customized solution for that key issue that is slowing your growth, support@friendli.ai or click Talk to an engineer — our engineers (not a bot) will reply within one business day.

