1. Home
  2. Companies
  3. FriendliAI

About

FriendliAI was founded in 2021 as a spin-out from research at Seoul National University, with the aim of commercialising AI research to make deploying and running large language models fast, affordable, and reliable. The company opened an office in San Francisco in 2026.

The company's core product is an AI inference platform purpose-built to deliver faster inference through custom GPU kernels, smart caching, continuous batching, speculative decoding, and parallel inference. The platform provides instant access to over 590,000 Hugging Face models and claims to achieve 2× or greater inference speedups compared to standard approaches. Technical work spans AI inference optimization, large language models, and GPU kernel development.

FriendliAI also offers enterprise-focused products, including Dedicated Endpoints for custom or fine-tuned models and container solutions, both backed by 99.99% uptime SLAs. The company operates in the AI research and enterprise AI verticals.

Similar companies

Inference logoIN

Inference

Inference runs a distributed GPU platform that aggregates idle compute to deliver low-cost, high-performance AI inference and custom model training services.

Novita AI logoNA

Novita AI

Novita AI provides accessible, cost-effective AI infrastructure through model APIs, serverless GPU compute, and secure agent sandboxes for developers and enterprises.

FA

Fireworks AI

Fireworks AI provides a globally distributed inference platform enabling developers to build, tune, and scale generative AI applications using open-source models.

Sciforium logoSC

Sciforium

Sciforium develops byte-native multimodal AI foundation models and a proprietary, high-efficiency platform for serving and deploying AI models.

WEKA logoWE

WEKA

WEKA develops AI-native data infrastructure software that accelerates machine learning and high-performance computing workloads across cloud and hardware environments.

Moonlite logoMO

Moonlite

Moonlite provides high-performance, co-located AI infrastructure and cloud-native platforms for large-scale model training and compute-intensive workloads for enterprises and research institutions.