Judgment Labs develops infrastructure for monitoring and improving AI agent behaviour in production environments. The company's core offering, Agent Behavior Monitoring (ABM), is designed to analyse production data from deployed AI agents, identify behavioural anomalies, and drive continuous improvement. Its work sits at the intersection of AI agent reliability and anomaly detection, serving teams building and operating AI-native systems.
The company's product suite includes Agent Search, which enables behavioural-level trajectory querying; Agent Judge, for trajectory-level evaluation through harnesses; Behavior Discovery, which surfaces failure modes from unlabelled production data; and AutoRubrics, which automatically constructs evaluation rubrics from verifiable signals.
Judgment Labs has raised $32 million across seed and Series A funding rounds. It operates as an applied-research lab, working on problems relevant to teams deploying AI agents at scale.






