Pricing build to scale with your growth

Fast, reliable, and affordable inference at any scale. Get started instantly with self-serve, or contact us for enterprise deployments.

Model APIs

Run the fastest frontier model inference with a simple API call.

See pricing

Dedicated Endpoints

Run dedicated inference with unmatched speed and reliability at scale.

See pricing

Container

Run inference with full control and performance in your environment.

Contact us

Looking for Enterprise Capabilities?

The Enterprise plan is offered as a customizable framework, not a fixed bundle.
Features and capabilities are enabled based on your contract.
Let’s talk about an Enterprise plan designed for you.

Scale & Reliability

  • Custom Model APIs rate limits
  • Priority access to high-demand GPU types
  • Reserved GPU capacity

Control & Deployment

  • Custom region deployments
  • VPC deployments
  • On-prem deployment options

Enterprise Commitments

  • Dedicated support channels
  • Named Customer Success ownership
  • Custom commercial terms
Contact Sales

Model APIs Pricing

Get instant access to the fastest frontier model inference with a simple API call.

Dedicated Endpoints Pricing

Get instant access to the fastest frontier model inference with a simple API call.

On-demand deployment

Only pay for the compute you use, down to the second, with no extra charges for start-up times

GPU Type

$ / hour (billed per second)

A100 80GB GPU

$2.9

H100 80GB GPU

$3.9

H200 141GB GPU

$4.5

B200 180GB GPU

$8.9

B300 288GB GPU

$12.0

Container Pricing

Run inference with full control and performance in your environment.

Contact us

FAQ

1.How do I get started?
Getting started takes just a few minutes. [1] Sign up for FriendliAI, [2] Generate your API key, and [3] Make your first inference request with frontier open-weight models.
2.How can I issue an API Key?
Sign up to Friendli Suite and visit the API Keys settings page, click 'Create API Key'. This personal key will be associated with your default team for tier and rate limit purposes.
3.Which models are available?

FriendliAI supports today’s frontier open-weight models—including GLM, MiniMax, Kimi, DeepSeek, Qwen, Gemma, and more—with a simple API call.

  • For Model API served ones, see here.
  • For FDE served ones, see here.
4.Do I pay for idle GPU time?
No. You are billed only while the GPU is active, metered per second, with no extra charge for startup time.
5.How secure is FriendliAI?
FriendliAI is SOC 2 Type II and HIPAA compliant. Read more
6.How does FriendliAI compare to other inference providers?
FriendliAI is built for production AI agents, combining speed, reliability, and efficiency at scale. It delivers low-latency streaming, reliable long-context inference, and robust tool calling without compromising stability. According to independent OpenRouter benchmarks, FriendliAI consistently ranks among the top providers for throughput, latency, and reliability across leading open-weight models. See why customers choose FriendliAI
7.Still have questions?
Didn't find what you're looking for? Check out FriendliAI Help Center, or reach out to our engineering team directly.

Explore FriendliAI today