- April 8, 2026
- 3 min read
FriendliAI Appoints Brian Yoo, Former Moloco COO, as Chief Business Officer to Drive Next Phase of Hypergrowth
AI industry veteran brings proven track record of scaling revenue 500x, securing $180M+ in capital, and leading global operations at the ~$4B machine learning company to accelerate adoption of The Frontier AI Inference Cloud.

FriendliAI, The Frontier AI Inference Cloud, today announced the appointment of Brian Yoo as Chief Business Officer. Joining from applied AI leader Moloco, Yoo will lead global commercial operations, go-to-market strategies, and partnerships as FriendliAI rapidly scales its high-performance, cost-efficient AI inference infrastructure worldwide.
Yoo brings exceptional operational and strategic scaling expertise to FriendliAI. He most recently served as Chief Operating Officer at Moloco, where he built its global operations entirely from the ground up. Yoo created, managed, and scaled the Finance, Marketing, HR, BizOps, Legal, IT, and Workplace Operations functions. His leadership was the engine behind the company's massive expansion: he successfully grew revenue more than 500x to over $250 million, scaled headcount from a 10-person startup to a global workforce of over 600 employees, and led pivotal fundraising efforts that secured more than $180 million in capital, helping drive Moloco to a ~$4 billion valuation.
"Brian’s track record of building the operational engine behind an AI-driven startup and scaling it into a multi-billion dollar global powerhouse is simply remarkable," said Byung-Gon Chun, CEO of FriendliAI. "As we experience surging demand for our frontier inference infrastructure, Brian's unmatched expertise in hypergrowth, global commercial operations, and capital strategy is exactly what we need to accelerate our expansion and deliver unparalleled value to our customers."
"As AI moves into production, performance at the inference layer directly determines how many tokens you can generate—and ultimately the margins you can capture," said Brian Yoo, Chief Business Officer at FriendliAI. "FriendliAI is positioned to maximize both, delivering industry-leading throughput and efficiency so our customers get the most out of every GPU. I am thrilled to join Byung-Gon and the incredible team here, especially now, as businesses that raced to build with large language models are beginning to heavily scrutinize their inference costs."
Accelerating Market Leadership in AI Inference
Yoo joins FriendliAI during a period of explosive momentum and enterprise adoption. FriendliAI's core technology, built by the researchers who pioneered the now-industry-standard "continuous batching" optimization technique, powers massive production workloads for rapidly growing AI companies.
Under Yoo's commercial leadership, the company will focus on rapidly expanding its market share among AI-native startups, SaaS companies, and enterprises looking to escape inefficient open-source setups or expensive closed model APIs. FriendliAI is already delivering transformative results for its partners: Twelve Labs, a leader in video understanding, partnered with FriendliAI for production inference at scale, while NextDay AI reported processing trillions of tokens monthly while cutting GPU requirements by 50% after migrating to the platform.
To learn more about FriendliAI and request an inference consultation, visit https://www.friendli.ai.
About FriendliAI
FriendliAI is The Frontier AI Inference Cloud. Built by the researchers who invented continuous batching, FriendliAI provides AI engineers with a highly optimized engine that runs state-of-the-art open-weight and custom models at production scale with 99.99% reliability. By maximizing GPU utilization, FriendliAI delivers speeds up to 3x faster than vLLM and 50% to 90% cost savings relative to closed model APIs, empowering engineers to deploy frontier AI with uncompromising speed and model ownership.
Written by
FriendliAI Tech & Research
Share
General FAQ
What is FriendliAI?
FriendliAI is the Frontier Inference Cloud for Agents, delivering high throughput, low latency, and reliability at scale for agentic workloads. Through vertically optimized inference infrastructure, it delivers 2–5× faster output token speed and a 99.99% uptime SLA for high-volume production traffic.
How does FriendliAI reduce inference costs?
FriendliAI reduces inference costs through higher GPU utilization and optimized inference performance. FriendliAI's patented continuous batching technique, along with quantization, speculative decoding, KV cache offloading, multi-LoRA serving, and autoscaling, helps you serve more tokens with fewer GPUs, lowering your infrastructure costs without sacrificing performance.
Why should I choose FriendliAI over other inference providers?
FriendliAI is built for production AI agents, combining speed, reliability, and efficiency at scale. It delivers low-latency streaming, reliable long-context inference, and robust tool calling without compromising stability. According to independent OpenRouter benchmarks, FriendliAI consistently ranks among the top providers for throughput, latency, and reliability across leading open-weight models. See why customers choose FriendliAI
Which open-weight models does FriendliAI support?
Run today’s frontier open-weight models—including GLM, MiniMax, Kimi, DeepSeek, Qwen, Gemma, and more—with a simple API call. FriendliAI Model API gives you instant access to the latest models with optimized inference performance for production workloads. Explore models and pricing
How do I get started?
Getting started takes just a few minutes. [1] Sign up for FriendliAI, [2] Generate your API key, and [3] Make your first inference request with frontier open-weight models.
Still have questions?
If you want a customized solution for that key issue that is slowing your growth, support@friendli.ai or click Talk to an engineer — our engineers (not a bot) will reply within one business day.

