• August 8, 2025
  • 2 min read

Introducing N-gram Speculative Decoding: Faster Inference for Structured Tasks

Introducing N-gram Speculative Decoding: Faster Inference for Structured Tasks thumbnail

We’re excited to introduce N-gram Speculative Decoding, a new feature in Dedicated Endpoints that speeds up LLM responses for structured and predictable tasks — such as code generation, legal drafting, or templated writing.

This is the first in a series of speculative decoding techniques coming to Dedicated Endpoints. N-gram speculative decoding accelerates inference by leveraging common N-gram patterns — with no changes required to your model or pipelines. It’s now available as a free-to-try feature for all plans.

What Is N-gram Speculative Decoding?

N-gram speculative decoding uses known token patterns (or N-grams) to look ahead in the generation process and predict likely next tokens. This approach enables the system to generate multiple tokens in parallel — greatly improving latency for deterministic or structured outputs.

Unlike draft-model speculative decoding, which relies on a separate lightweight model to propose future tokens, N-gram speculative decoding leverages the model’s own output patterns, making it faster to initialize and simpler to deploy.

With N-gram speculative decoding, you get:

  • Faster outputs: Reduces Time-per-Output-Token (TPOT)
  • No need to train draft models: It works without draft models, simplifying the setup
  • Seamless integration: Just toggle it on during endpoint creation
  • Optimized performance: Especially powerful when combined with Friendli Inference

This makes N-gram speculative decoding ideal for applications like:

  • Code generation
  • Formatted emails or reports
  • Legal contracts
  • Structured JSON generation
  • Templated summaries

Getting Started

To enable N-gram speculative decoding:

  1. Create a new Dedicated Endpoint
  2. Toggle on “N-gram speculative decoding” under “Endpoint features
  3. (Optional) Set minimum and maximum N-gram size for customization
  4. Deploy — no model or code changes needed
Create an endpoint with N-gram speculative decoding enabled.
Figure 1: Create an endpoint with N-gram speculative decoding enabled.

To see whether N-gram speculative decoding is enabled or not for an endpoint, you can check the overview page.

An endpoint overview with N-gram speculative decoding enabled.
Figure 2: An endpoint overview with N-gram speculative decoding enabled.

When enabled, N-gram speculative decoding automatically detects frequently occurring token sequences and uses those patterns to pre-generate likely continuations. The model then verifies these predictions in parallel, skipping unnecessary steps and speeding up generation.

If the predicted N-grams are correct, they're committed instantly. If not, the model falls back to standard decoding. This yields faster inference with minimal overhead and no accuracy tradeoff.

N-gram speculative decoding further accelerates generation with lookahead techniques, integrating seamlessly with our other advanced technologies powered by Friendli Inference.

To learn more about N-gram speculative decoding, check out our documentation!


Written by

FriendliAI Tech & Research


Share


General FAQ

What is FriendliAI?

FriendliAI is the Frontier Inference Cloud for Agents, delivering high throughput, low latency, and reliability at scale for agentic workloads. Through vertically optimized inference infrastructure, it delivers 2–5× faster output token speed and a 99.99% uptime SLA for high-volume production traffic.

How does FriendliAI reduce inference costs?

FriendliAI reduces inference costs through higher GPU utilization and optimized inference performance. FriendliAI's patented continuous batching technique, along with quantization, speculative decoding, KV cache offloading, multi-LoRA serving, and autoscaling, helps you serve more tokens with fewer GPUs, lowering your infrastructure costs without sacrificing performance.

Why should I choose FriendliAI over other inference providers?

FriendliAI is built for production AI agents, combining speed, reliability, and efficiency at scale. It delivers low-latency streaming, reliable long-context inference, and robust tool calling without compromising stability. According to independent OpenRouter benchmarks, FriendliAI consistently ranks among the top providers for throughput, latency, and reliability across leading open-weight models. See why customers choose FriendliAI

Which open-weight models does FriendliAI support?

Run today’s frontier open-weight models—including GLM, MiniMax, Kimi, DeepSeek, Qwen, Gemma, and more—with a simple API call. FriendliAI Model API gives you instant access to the latest models with optimized inference performance for production workloads. Explore models and pricing

How do I get started?

Getting started takes just a few minutes. [1] Sign up for FriendliAI, [2] Generate your API key, and [3] Make your first inference request with frontier open-weight models.

Still have questions?

If you want a customized solution for that key issue that is slowing your growth, support@friendli.ai or click Talk to an engineer — our engineers (not a bot) will reply within one business day.


Explore FriendliAI today