Data lab to train frontier models & evaluate agents

The data research company to accelerate AI in the real world

Research

Our research lab aspires to solve humanity’s greatest coordination challenge: deciding where each person should spend their time

PersonalAgentBench

We tested leading personal AI agents on complex everyday tasks, from drafting sensitive messages to finding flights and the best prices.

    CortexRetrievalBench

    An evaluation of AI retrieval systems' ability to find and rank authoritative sources for financial research

      Realm: Legal reasoning benchmark

      The standard for evaluating legal reasoning in AI systems