| Histogram | Metric Name | Description |
|---|---|---|
| Friendli TCache hit ratio (0≤value≤1) | friendli\_tcache\_hit\_ratio\_bucket | Bucketized number of histogram samples for TCache hit ratio, with le label |
| friendli\_tcache\_hit\_ratio\_count | Total number of histogram samples for TCache hit ratio | |
| friendli\_tcache\_hit\_ratio\_sum | Sum of histogram sample values for TCache hit ratio | |
| The length of input tokens (Experimental metric) | friendli\_input\_lengths\_bucket | Bucketized number of histogram samples for length of input tokens, with le label |
| friendli\_input\_lengths\_count | Total number of histogram samples for length of input tokens | |
| friendli\_input\_lengths\_sum | Sum of histogram sample values for length of input tokens | |
| The length of output tokens (Experimental metric) | friendli\_output\_lengths\_bucket | Bucketized number of histogram samples for length of output tokens, with le label |
| friendli\_output\_lengths\_count | Total number of histogram samples for length of output tokens | |
| friendli\_output\_lengths\_sum | Sum of histogram sample values for length of output tokens |
| Quantiles | Metric Name | Description |
|---|---|---|
| Request completion latency (in nanoseconds) | friendli\_requests\_latencies | Percentile value for request completion latency (quantile label is either 0.5, 0.9, or 0.99) |
| friendli\_requests\_latencies\_count | Total number of samples for request completion latency | |
| friendli\_requests\_latencies\_sum | Sum of sample values for request completion latency | |
| Time to first token (TTFT) (in nanoseconds) | friendli\_requests\_ttft | Percentile value for time to first token (TTFT) (quantile label is either 0.5, 0.9, or 0.99) |
| friendli\_requests\_ttft\_count | Total number of samples for time to first token (TTFT) | |
| friendli\_requests\_ttft\_sum | Sum of sample values for time to first token (TTFT) | |
| Request queueing delay (in nanoseconds) | friendli\_requests\_queueing\_delays | Percentile value for queueing delay (quantile label is either 0.5, 0.9, or 0.99) |
| friendli\_requests\_queueing\_delays\_count | Total number of samples for queueing delay | |
| friendli\_requests\_queueing\_delays\_sum | Sum of sample values for queueing delay |
With FriendliAI, your team gets fast, affordable, and reliable AI inference at scale.
Start with Model APIs, a curated set of popular, open-weight models that are ready for you today.
If you want to run any model — including your own — on dedicated GPUs, try Dedicated Endpoints.
Learn about inference and FriendliAI by reading how-tos and best-practice articles.
Read the GuidesExplore real-world use cases and see what you can do with FriendliAI.
Browse the ExamplesLook up the API's operations and parameters. For example, learn how to use the Chat Completions API.
View the ReferenceCreate an API key and send your first request. Once you complete these steps, you're ready to use your agent or SDK with FriendliAI.
Use OpenAI-compatible tool calling with broad model support, strict schema enforcement, and parallel tool calls.
View pricing per model. Compare token-based and audio-based rates across text and audio models.
Use the FriendliLink CLI to connect your agent to FriendliAI with one command. Once you complete these steps, your agent will send requests to FriendliAI.
Use Z.ai's flagship open-weight model with FriendliAI. Review the model's properties and pricing, then choose a feature to start building.
Set up Hermes Agent to connect to FriendliAI. Once you complete these steps, your agent will send requests to FriendliAI.