1. Home
  2. Companies
  3. Preference Model

About

Preference Model is an AI research company developing reinforcement learning environments intended to replicate real-world complexity. Its work focuses on creating robust training settings for machine learning models, with a core emphasis on building data infrastructure and constructing curated datasets. The company partners with leading AI labs to advance the field.

The founding team brings experience from organisations including Anthropic, Stripe, and Datology. Technical expertise spans reinforcement learning, machine learning research, data infrastructure, dataset construction, and cybersecurity.

Preference Model's primary product consists of RL Environments designed for training AI models. These environments feature diverse tasks and robust reward functions, aiming to bridge the gap between simulated training conditions and practical deployment scenarios.

Similar companies

Scale logoSC

Scale

Scale AI is a San Francisco-based data annotation platform that provides high-quality training data and full-stack AI infrastructure to power machine learning models for enterprises, governments, and AI labs worldwide.

1 job
Datacurve logoDA

Datacurve

Datacurve provides high-quality training and evaluation data for frontier AI models, using a gamified bounty platform that engages software engineers on complex coding challenges.

ID

Idler

Idler builds reinforcement learning environments that train AI models to write code at expert human levels using real-world coding scenarios.

DatologyAI logoDA

DatologyAI

DatologyAI develops automated tools to select and optimize training data for deep learning, enabling faster model training and reduced computational costs.

AfterQuery logoAF

AfterQuery

AfterQuery is an applied research lab that creates specialized data solutions, including reasoning traces and agent environments, to support frontier foundation model development.

Encord logoEN

Encord

Encord provides an AI data development platform covering data management, curation, annotation, workforce tooling, and model evaluation and observability.