Preference Model is an AI research company developing reinforcement learning environments intended to replicate real-world complexity. Its work focuses on creating robust training settings for machine learning models, with a core emphasis on building data infrastructure and constructing curated datasets. The company partners with leading AI labs to advance the field.
The founding team brings experience from organisations including Anthropic, Stripe, and Datology. Technical expertise spans reinforcement learning, machine learning research, data infrastructure, dataset construction, and cybersecurity.
Preference Model's primary product consists of RL Environments designed for training AI models. These environments feature diverse tasks and robust reward functions, aiming to bridge the gap between simulated training conditions and practical deployment scenarios.





