Task Datasets & Rubrics

Long-horizon task datasets elicited from verified practitioners, with expert-weighted rubrics and deterministic verifiers. Ground truth as a set of acceptable outcomes, not a single golden answer.

Archival task packets with expert rubrics and tolerance bands

Tasks written by the people who do the work

Most task datasets are written by contractors imagining a job. Ours are elicited from verified practitioners doing the job, capturing the decisions, exceptions, and tacit tricks that never made it to the internet, then confirmed and refined by the expert themselves.

Every task carries an explicit notion of done. Multi-hour, multi-application, long-horizon work, not single-turn prompts.

Rubrics that turn judgment into reward

Each task carries an expert-weighted rubric. Criteria compile into deterministic verifiers wherever the software allows it: database assertions, numeric reconciliation, file diffs, audit-log checks, and invariant assertions. The ledger ties, or the episode fails.

Where judgment is genuinely subjective, model-graded criteria are used sparingly, capped at a fifth of any corpus, always labeled as such, and always audited by a second expert.

Quality machinery, not spot checks

One expert authors a task. A second runs it blind. A third adjudicates divergence. Elite practitioners are promoted into a reviewer tier that cross-reviews every packet before release, and independent expert runs corroborate ground truth before a task earns its place in a sellable cut. If ten specialists give us a hundred tasks, our job is to find the five best.

What ships in the box

Three-part task schema

Every task is input, tools, and expected output. No ambiguity about what the agent gets, what it may use, and what done means.

Ground truth is a set

Two competent practitioners legitimately diverge. Acceptable outcome sets with numeric tolerance bands capture that divergence as signal, not noise.

Mandatory & disqualifying conditions

Invariants that must hold and mistakes that end the episode, weighted by the experts who live with the consequences.

Expert corrections corpus

Where agents fail, experts correct. Demonstrations, preference pairs, and process labels from the same verified practitioners who authored the tasks.

Difficulty bands & distractors

Task families ladder from standard to hard, with distractors drawn from real workflow noise rather than synthetic padding.

Exclusivity options

License shared cuts or commission exclusive task families in your capability area. Exclusive cuts never appear in any other customer's corpus.

See it before you buy it.

Sample task packets and environment access for evaluation. Tell us the capability you care about and we will send the relevant cut.