- July 25, 2025
- 2 min read
Announcing Online Quantization: Faster, Cheaper Inference with Same Accuracy

We're introducing Online Quantization — a new feature in Dedicated Endpoints that lowers inference costs while maintaining accuracy without requiring any model prep.
This feature automatically quantizes your models as it loads, so there’s no need to store separate quantized versions. With online quantization, Dedicated Endpoints can run your workloads using fewer GPUs, reducing infrastructure costs and accelerating speed while maintaining reliable inference.
What Is Online Quantization?
Online quantization automatically converts model weights and activations from its original weight (such as FP16) to lower-precision formats like FP8 or FP4 during model loading. Unlike traditional quantization methods that require preprocessing, retraining, or model changes, online quantization adjusts precision on the fly, providing a seamless and efficient experience, such as:
- Zero setup: No calibration data, retraining, or model changes required.
- Faster inference: Improves both Time-to-First-Token (TTFT) and Time-Per-Output-Token (TPOT).
- Lower GPU usage: Cut GPU needs by 2-4x, significantly reducing costs.
- On-the-fly conversion: Quantization happens automatically during model load.
- Preserved accuracy: Accuracy remains nearly identical to the original.
Powered by Friendli Inference, online quantization further optimizes compute efficiency alongside our cutting-edge technologies like iteration batching (a.k.a. continuous batching), multimodal caching, multi-LoRA, speculative decoding, and more.
How Online Quantization Cuts Costs
Quantization reduces precision (e.g., FP16 → FP8), cutting compute requirements and memory bandwidth during inference.
With online quantization now integrated into Dedicated Endpoints:
- Models run with less GPU memory and compute power per request
- Endpoint throughput increases, enabling more requests per GPU
- You can serve the same workload with fewer GPUs, reducing your cloud or data center costs
Fast Inference. Lower Costs. No Retraining.
Skip the slowdowns of traditional quantization. Our online approach runs automatically at initialization—no retraining, no calibration—delivering minimal accuracy loss with maximum speed.
Accelerate Time-to-First-Token (TTFT) and Time-per-Output-Token (TPOT) while cutting GPU costs instantly. No need to manage multiple model versions or pipelines. Speed meets simplicity with our online quantization technology.
Getting Started
To enable Online Quantization, just turn it on in your Dedicated Endpoint configuration — no model changes or pipeline updates required. It supports all major model formats and GPU types, and integrates directly into your existing AI workflows.
After selecting an eligible model on the endpoint creation page, you can toggle "Online Quantization" on or off under the "Endpoint features" section.

Qwen/Qwen2.5-72B-Instruct, for instance, would require 4x NVIDIA H100 GPUs, but with Online Quantization, it can run with only 2x NVIDIA H100 GPUs, cutting the cost by half.


To see whether online quantization is enabled for an endpoint, simply check the endpoint overview.

To learn more about how to configure your Dedicated Endpoints, please refer to our docs.
Written by
FriendliAI Tech & Research
Share
General FAQ
What is FriendliAI?
FriendliAI is the Frontier Inference Cloud for Agents, delivering high throughput, low latency, and reliability at scale for agentic workloads. Through vertically optimized inference infrastructure, it delivers 2–5× faster output token speed and a 99.99% uptime SLA for high-volume production traffic.
How does FriendliAI reduce inference costs?
FriendliAI reduces inference costs through higher GPU utilization and optimized inference performance. FriendliAI's patented continuous batching technique, along with quantization, speculative decoding, KV cache offloading, multi-LoRA serving, and autoscaling, helps you serve more tokens with fewer GPUs, lowering your infrastructure costs without sacrificing performance.
Why should I choose FriendliAI over other inference providers?
FriendliAI is built for production AI agents, combining speed, reliability, and efficiency at scale. It delivers low-latency streaming, reliable long-context inference, and robust tool calling without compromising stability. According to independent OpenRouter benchmarks, FriendliAI consistently ranks among the top providers for throughput, latency, and reliability across leading open-weight models. See why customers choose FriendliAI
Which open-weight models does FriendliAI support?
Run today’s frontier open-weight models—including GLM, MiniMax, Kimi, DeepSeek, Qwen, Gemma, and more—with a simple API call. FriendliAI Model API gives you instant access to the latest models with optimized inference performance for production workloads. Explore models and pricing
How do I get started?
Getting started takes just a few minutes. [1] Sign up for FriendliAI, [2] Generate your API key, and [3] Make your first inference request with frontier open-weight models.
Still have questions?
If you want a customized solution for that key issue that is slowing your growth, support@friendli.ai or click Talk to an engineer — our engineers (not a bot) will reply within one business day.

