Solutions
Data and environment programs for frontier labs, applied AI teams, and enterprises: vertical eval programs, horizontal and collaborative scenarios, RL environment fleets, and in-VPC deployment.
Vertical data & eval programs
Deep, role-specific programs in occupations labs cannot staff. Payroll closes, claims files, requisitions, ledgers: task sets, rubrics, verifiers, and benchmarks authored by verified practitioners on real software, graded so the ledger ties. This is where rubric quality and real-world usability decide vendor evaluations, and where we over-invest.

Horizontal & collaborative scenarios
The least mature space in agent evaluation is not role-specific work. It is the diverse, everyday coordination of information workers, and multi-user collaboration on top of it. We source query sets across many kinds of users and simulate the other people in the loop, because these scenarios have no single right answer and need a richer corpus per use case.

RL environment fleets for post-training
Resettable, seeded, snapshot-restorable environments engineered for thousands of eval jobs a day. Real business software at pinned versions, shaped rewards compiled from expert rubrics, and variants at scale, so the same instrument serves evals today and the RLVR loop tomorrow.

Enterprise & in-VPC deployment
For organizations whose data cannot leave, the entire stack licenses into your own cloud. Post-train on your tenant's real workflows, build evals from your own operations, and keep every byte inside your boundary. Built for regulated verticals where this is the only acceptable shape.

How engagements actually run
Most work in this market starts with “we need this,” and we are built for that: requirement in, sample packet out, usually within days. The packet contains representative tasks, rubrics and verifiers, the expert baseline, and environment access, so your team evaluates the data by running agents against it rather than reading about it.
From there, engagements scale along whichever axis the requirement demands: more occupations, harder difficulty bands, exclusive task families, private holdout licensing, or a standing refresh contract tied to your release cadence. Flexibility is the point. Datasets alone, environments alone, grading alone, or the full instrument.
Most engagements start with “we need this.”
Tell us the requirement. We answer with samples: representative tasks, their rubrics and verifiers, the expert baseline, and environment access.