• January 4, 2024
  • 2 min read

Friendli Serverless Endpoints: Unleashing Generative AI for Everyone

Friendli Serverless Endpoints: Unleashing Generative AI for Everyone thumbnail

FriendliAI, the world’s leading generative AI engine company, has launched its Friendli Serverless Endpoints, unlocking a new era of accessibility for generative AI inference. Users can access open-source generative AI models with simple API calls with the lowest cost on the market. This innovative service brings the power of Friendli Inference, our GPU-optimized inference engine, to anyone, regardless of their technical expertise.

Say goodbye to deployment headaches: Gone are the days of wrestling with infrastructure and optimizing models on complex GPU machines. Friendli Serverless Endpoints takes care of everything, allowing you to harness the transformative potential of generative AI models right within your applications.

Who is it for? Whether you're:

  • A curious developer eager to experiment with cutting-edge LLMs like Llama-2 and image creation models like Stable Diffusion,
  • A product manager seeking to integrate text generation or image creation into your product, or
  • A researcher exploring preliminary LLM features before diving into deep-dive fine-tuning,

Friendli Serverless Endpoints provides the perfect platform to unlock the magic of generative AI.

No more barriers: Friendli Serverless Endpoints removes the technical hurdles that often block the adoption of generative AI. You no longer need to worry about setting up the infrastructure, optimizing the model serving, or even choosing the right GPU. Simply connect your application to Friendli's secure endpoints and start weaving generative AI magic into your workflow with the lowest cost on the market.

Power under the hood: While Friendli Serverless Endpoints simplifies your experience, Friendli Inference, the beating heart of the service, delivers unparalleled inference serving performance and cost-efficiency.

  • Reduced costs: $0.2/M0.2/Mtokens for Llama-2 13B and $0.8/M0.8/Mtokens for Llama-2 70B, thanks to Friendli Inference.
  • Low latency: 2-4x faster compared to other leading solutions that use vLLM, ensuring a smooth and responsive generative AI experience.

Open doors to diverse models: Get started with a curated selection of popular open-source models including:

  • Large language models: Llama-2 and Llama-2-chat (13B and 70B), Mistral 7B and Mistral-7B-instruct
  • Visual models: Stable Diffusion v1.5
  • And more models soon to come!
Diverse models are supported in Friendli Serverless Endpoints including Llama-2, Llama-2-Chat, Mistral 7B, Mistral-7B-instruct, Stable Diffusion v1.5, and more-FriendliAI

More choices, more power: While Friendli Serverless Endpoints democratizes generative AI, FriendliAI also offers Friendli Dedicated Endpoints for advanced users. This premium service provides dedicated GPU instances, allowing you to serve your customized models reliably with high performance and low costs.

Start your generative journey today: Friendli Serverless Endpoints is the key to unlocking the incredible potential of generative AI. Sign up today and start building applications that leverage the power of language, image, and code without getting bogged down in technical complexities.

The future is generative. With FriendliAI, it's more accessible than ever.


Written by

FriendliAI Tech & Research


Share


General FAQ

What is FriendliAI?

FriendliAI is the Frontier Inference Cloud for Agents, delivering high throughput, low latency, and reliability at scale for agentic workloads. Through vertically optimized inference infrastructure, it delivers 2–5× faster output token speed and a 99.99% uptime SLA for high-volume production traffic.

How does FriendliAI reduce inference costs?

FriendliAI reduces inference costs through higher GPU utilization and optimized inference performance. FriendliAI's patented continuous batching technique, along with quantization, speculative decoding, KV cache offloading, multi-LoRA serving, and autoscaling, helps you serve more tokens with fewer GPUs, lowering your infrastructure costs without sacrificing performance.

Why should I choose FriendliAI over other inference providers?

FriendliAI is built for production AI agents, combining speed, reliability, and efficiency at scale. It delivers low-latency streaming, reliable long-context inference, and robust tool calling without compromising stability. According to independent OpenRouter benchmarks, FriendliAI consistently ranks among the top providers for throughput, latency, and reliability across leading open-weight models. See why customers choose FriendliAI

Which open-weight models does FriendliAI support?

Run today’s frontier open-weight models—including GLM, MiniMax, Kimi, DeepSeek, Qwen, Gemma, and more—with a simple API call. FriendliAI Model API gives you instant access to the latest models with optimized inference performance for production workloads. Explore models and pricing

How do I get started?

Getting started takes just a few minutes. [1] Sign up for FriendliAI, [2] Generate your API key, and [3] Make your first inference request with frontier open-weight models.

Still have questions?

If you want a customized solution for that key issue that is slowing your growth, support@friendli.ai or click Talk to an engineer — our engineers (not a bot) will reply within one business day.


Explore FriendliAI today