Task Datasets & Rubrics
Long-horizon task datasets elicited from verified practitioners, with expert-weighted rubrics and deterministic verifiers. Ground truth as a set of acceptable outcomes, not a single golden answer.


Conversational Data
Multi-turn conversational training data from real work: negotiations, multi-stakeholder coordination, and long-horizon dialogues with outcomes, not scripted chat transcripts.

Verifiers & Rubrics for Verticals
Expert-weighted rubrics and deterministic verifiers for vertical AI: payroll, claims, billing, recruiting, and back-office domains. The grading is what matters, and ours compiles to code.

Multimodal Data
Multimodal training data from real work: annotated screen trajectories, voice interactions, documents, and video demonstrations of professional workflows, captured with full consent.

Off-the-Shelf Datasets
Licensed, ready-to-run task datasets and RL environments off the shelf. Sample the cut, run your agents against it, and license shared or exclusive, without a months-long custom engagement.

Custom Evals
Custom evaluation suites and private benchmarks built to your requirement: agentic metrics that grade end outcomes and process, expert-baselined, refreshed against every frontier release.
Tasks written by the people who do the work
Most task datasets are written by contractors imagining a job. Ours are elicited from verified practitioners doing the job, capturing the decisions, exceptions, and tacit tricks that never made it to the internet, then confirmed and refined by the expert themselves.
Every task carries an explicit notion of done. Multi-hour, multi-application, long-horizon work, not single-turn prompts.
Rubrics that turn judgment into reward
Each task carries an expert-weighted rubric. Criteria compile into deterministic verifiers wherever the software allows it: database assertions, numeric reconciliation, file diffs, audit-log checks, and invariant assertions. The ledger ties, or the episode fails.
Where judgment is genuinely subjective, model-graded criteria are used sparingly, capped at a fifth of any corpus, always labeled as such, and always audited by a second expert.
Quality machinery, not spot checks
One expert authors a task. A second runs it blind. A third adjudicates divergence. Elite practitioners are promoted into a reviewer tier that cross-reviews every packet before release, and independent expert runs corroborate ground truth before a task earns its place in a sellable cut. If ten specialists give us a hundred tasks, our job is to find the five best.
What ships in the box
Every task is input, tools, and expected output. No ambiguity about what the agent gets, what it may use, and what done means.
Two competent practitioners legitimately diverge. Acceptable outcome sets with numeric tolerance bands capture that divergence as signal, not noise.
Invariants that must hold and mistakes that end the episode, weighted by the experts who live with the consequences.
Where agents fail, experts correct. Demonstrations, preference pairs, and process labels from the same verified practitioners who authored the tasks.
Task families ladder from standard to hard, with distractors drawn from real workflow noise rather than synthetic padding.
License shared cuts or commission exclusive task families in your capability area. Exclusive cuts never appear in any other customer's corpus.
See it before you buy it.
Sample task packets and environment access for evaluation. Tell us the capability you care about and we will send the relevant cut.