Pricing built to scale with your growth
Fast, reliable, and affordable inference at any scale. Get started instantly with self-serve, or contact us for enterprise deployments.
Model APIs
Run the fastest frontier model inference with a simple API call.
Dedicated Endpoints
Run dedicated inference with unmatched speed and reliability at scale.
BYOG
Run Friendli inference stack on your own GPUs.
Looking for Enterprise Capabilities?
The Enterprise plan is offered as a customizable framework, not a fixed bundle.
Features and capabilities are enabled based on your contract.
Let’s talk about an Enterprise plan designed for you.
Scale & Reliability
- Custom Model APIs rate limits
- Priority access to high-demand GPU types
- Reserved GPU capacity
Control & Deployment
- Custom region deployments
- VPC deployments
- On-prem deployment options
Enterprise Commitments
- Dedicated support channels
- Named Customer Success ownership
- Custom commercial terms
Model APIs Pricing
Get instant access to the fastest frontier model inference with a simple API call.
Dedicated Endpoints Pricing
Get instant access to the fastest frontier model inference with a simple API call.
On-demand deployment
Only pay for the compute you use, down to the second, with no extra charges for start-up times
GPU Type
$ / hour (billed per second)
Through Sep 30
Effective Oct 1
A100 80GB
$2.9
$4.0
H100 80GB
$3.9
$5.0
H200 141GB
$4.5
$7.0
B200 180GB
$8.9
$9.0
B300 288GB
$12.0
$12.0
BYOG Pricing
BYOG pricing is tailored to your infrastructure and deployment needs. Talk to us to find the right setup for your workload.