Datacurve builds data infrastructure for frontier AI models, supplying the high-quality datasets required to improve foundation model capabilities. The company provides post-training and evaluation data tailored to the needs of foundation model labs and enterprises, spanning supervised fine-tuning, reinforcement learning environments, and reinforcement learning with human feedback.
The company's product, Shipd, gamifies the data creation process by transforming complex coding challenges into competitive problem-solving opportunities. Engineers compete on bounty-based tasks, producing datasets of reasoning challenges, debugging exercises, and agentic workflow traces. This approach simultaneously generates high-fidelity training and evaluation data whilst engaging top software engineering talent through a rewards-based system.
Datacurve's datasets extend across multiple data types: supervised fine-tuning material, reinforcement learning environments, RLHF training data, private repository taskbenches, and reasoning or debugging challenges. The company works with leading foundation model labs and enterprises seeking to unlock new model capabilities through data-driven improvement.





