- May 20, 2022
- 2 min read
Introducing GPT-FAI 13B: A Large-scale Language Model Trained with FriendliAI’s Friendli Training

We are happy to announce that we are releasing GPT-FAI 13B, a large-scale language model trained with Friendli Training (formerly known as PeriFlow). GPT-FAI 13B is a 13-billion parameter version of GPT-3 trained on publicly available datasets. The release allows AI model researchers to perform research on various topics in large-scale language models for non-commercial purposes.
We trained GPT-FAI 13B with Friendli Training, an end-to-end large-scale AI training and serving cloud service. Although training large models like GPT-FAI 13B requires lots of GPU power, with a cloud service, one can use the cloud to train to avoid the effort and costs of building and maintaining server clusters. However, such training requires lots of GPUs and training can take anywhere from a few days to a few months. Over long training periods, it becomes ever more paramount to optimize the execution and swiftly handle any faults that occur. Luckily, training GPT-FAI 13B was a breeze. :)
FriendliAI has built Friendli Training for any client who would like to build large-scale AI models. Friendli Training is FriendliAI’s product that realizes the company’s vision, “Make large-scale AI simple for the enterprise.” With Friendli Training, one can simplify the process of training large-scale AI models on hundreds of GPUs and serving them. Friendli Training employs various optimization technologies developed by FriendliAI. It can train faster and handle faults smoother. Friendli Training is multi-cloud; currently, it supports popular cloud vendors such as AWS, Azure, and GCP.
Once the trained model is available, one can deploy the model using our Friendli Inference for inference serving. Friendli Inference shows significant improvement over state-of-the-art inference systems such as NVIDIA FasterTransformer, a well-known inference system for Transformer models. We provide a comparison of performance results below. With Friendli Inference, anyone can create and serve large-scale AI models with ease.
GPT-FAI 13B performance
We evaluated our model on various downstream tasks with lm-evaluation-harness. Note that our model is not fine-tuned to the downstream tasks, nor did we use any sophisticated prompt engineering. The following zero-shot results may not exactly represent the performance of our model.

Illustrating training monitoring and fault handling
Figure 1 represents a metric collected during GPT-FAI 13B training with Friendli Training. The different colors in the graph demonstrate that despite various faults occurring, the training process was able to automatically recover.

Serving performance of the Friendli Inference

Figure 2. Inference latency and throughput of GPT 13B/175B on Friendli Inference and FasterTransformer
Figure 2 compares inference serving performance when using FasterTransformer versus the Friendli Inference on GPT 13B and 175B models with respect to text generation latency and throughput. The experiments were conducted on a single A100 GPU for the 13B model and 16 A100 GPUs for the 175B model. Friendli Inference significantly outperforms FasterTransformer, showing an order of magnitude higher throughput at the same level of latency. We plan to share more details in a few months. Stay tuned!

Contact us if you want to try out Friendli Inference!
Written by
FriendliAI Tech & Research
Share
General FAQ
What is FriendliAI?
FriendliAI is the Frontier Inference Cloud for Agents, delivering high throughput, low latency, and reliability at scale for agentic workloads. Through vertically optimized inference infrastructure, it delivers 2–5× faster output token speed and a 99.99% uptime SLA for high-volume production traffic.
How does FriendliAI reduce inference costs?
FriendliAI reduces inference costs through higher GPU utilization and optimized inference performance. FriendliAI's patented continuous batching technique, along with quantization, speculative decoding, KV cache offloading, multi-LoRA serving, and autoscaling, helps you serve more tokens with fewer GPUs, lowering your infrastructure costs without sacrificing performance.
Why should I choose FriendliAI over other inference providers?
FriendliAI is built for production AI agents, combining speed, reliability, and efficiency at scale. It delivers low-latency streaming, reliable long-context inference, and robust tool calling without compromising stability. According to independent OpenRouter benchmarks, FriendliAI consistently ranks among the top providers for throughput, latency, and reliability across leading open-weight models. See why customers choose FriendliAI
Which open-weight models does FriendliAI support?
Run today’s frontier open-weight models—including GLM, MiniMax, Kimi, DeepSeek, Qwen, Gemma, and more—with a simple API call. FriendliAI Model API gives you instant access to the latest models with optimized inference performance for production workloads. Explore models and pricing
How do I get started?
Getting started takes just a few minutes. [1] Sign up for FriendliAI, [2] Generate your API key, and [3] Make your first inference request with frontier open-weight models.
Still have questions?
If you want a customized solution for that key issue that is slowing your growth, support@friendli.ai or click Talk to an engineer — our engineers (not a bot) will reply within one business day.

