• December 16, 2025
  • 2 min read

A Faster, Convenient Way to Discover and Deploy AI Models on FriendliAI

TL;DR
  • FriendliAI launched a refreshed model page with a new table-based view replaces card layouts, allowing teams to scan, filter, and compare the catalog of 484K+ models with ease.
  • Faster Deployment → New dedicated model pages feature one-click "Deploy" buttons that pre-fill configurations, alongside an instant Playground for zero-setup serverless testing.
  • Model specifications are now hosted on permanent, linkable pages to simplify technical evaluation and internal sharing.
A Faster, Convenient Way to Discover and Deploy AI Models on FriendliAI thumbnail

As the open-source model ecosystem continues to grow at an unprecedented pace, model discovery and deployment should be just as fast and scalable as inference itself.

Today, we’re excited to share a refreshed Model page experience on FriendliAI designed to make browsing, evaluating, and deploying models significantly faster and more intuitive for developers and ML teams.

👉 Explore the updated Model page: https://friendli.ai/model

Built for Scale: A New Table-Based Model List

With hundreds of thousands of models available, clarity matters. The new Model page introduces a table-based model list, making it easier to:

  • Scan large model catalogs at a glance
  • Filter and navigate models by key attributes
  • Quickly identify the right model for your workload

Instead of scrolling through cards or opening multiple modals, you can now compare models efficiently in a single, structured view, optimized for speed and usability.

Rich Model Detail Pages

Each model now has its own dedicated detail page, replacing the previous modal-based experience.

These pages provide:

  • Clear model specifications and supported features
  • Deployment guidance and usage context
  • A more permanent, linkable reference for teams evaluating models

This change makes it easier to share model information internally and supports more informed deployment decisions especially when working across teams.

Faster Deployment with One-Click Setup

We’ve streamlined the path from discovery to deployment. From the Model page, you can now:

  • Click Deploy to open the Suite FDE create page with the selected model already pre-filled
  • Skip repetitive configuration steps and move straight to provisioning

For teams deploying frequently or testing multiple models, this reduces friction and speeds up iteration.

Instant Model Testing for Serverless Endpoints

For serverless models, exploration is now even faster. With a single click, you can:

  • Use the model directly in the Playground
  • Test prompts and behavior instantly. No setup required

This makes it easier to validate model quality and behavior before deploying into production workflows.

Designed for How Teams Actually Work

This refresh isn’t just a UI update, it’s a reflection of how developers and ML teams use FriendliAI every day:

  • Discovering new models quickly
  • Comparing capabilities at scale
  • Moving seamlessly from evaluation to deployment

As FriendliAI continues to support 484K+ open-source and custom models across text, vision, image, and audio, we’ll keep investing in tooling that makes high-performance inference easier to access and faster to operate.

Explore the New Model Experience

👉 Visit the updated Model page: https://friendli.ai/model


We’d love to hear your feedback as you explore the new experience, and as always, we’re building with real production workloads in mind.


Written by

FriendliAI Tech & Research


Share


General FAQ

What is FriendliAI?

FriendliAI is the Frontier Inference Cloud for Agents, delivering high throughput, low latency, and reliability at scale for agentic workloads. Through vertically optimized inference infrastructure, it delivers 2–5× faster output token speed and a 99.99% uptime SLA for high-volume production traffic.

How does FriendliAI reduce inference costs?

FriendliAI reduces inference costs through higher GPU utilization and optimized inference performance. FriendliAI's patented continuous batching technique, along with quantization, speculative decoding, KV cache offloading, multi-LoRA serving, and autoscaling, helps you serve more tokens with fewer GPUs, lowering your infrastructure costs without sacrificing performance.

Why should I choose FriendliAI over other inference providers?

FriendliAI is built for production AI agents, combining speed, reliability, and efficiency at scale. It delivers low-latency streaming, reliable long-context inference, and robust tool calling without compromising stability. According to independent OpenRouter benchmarks, FriendliAI consistently ranks among the top providers for throughput, latency, and reliability across leading open-weight models. See why customers choose FriendliAI

Which open-weight models does FriendliAI support?

Run today’s frontier open-weight models—including GLM, MiniMax, Kimi, DeepSeek, Qwen, Gemma, and more—with a simple API call. FriendliAI Model API gives you instant access to the latest models with optimized inference performance for production workloads. Explore models and pricing

How do I get started?

Getting started takes just a few minutes. [1] Sign up for FriendliAI, [2] Generate your API key, and [3] Make your first inference request with frontier open-weight models.

Still have questions?

If you want a customized solution for that key issue that is slowing your growth, support@friendli.ai or click Talk to an engineer — our engineers (not a bot) will reply within one business day.


Explore FriendliAI today