Methodology
How rlsupply turns the real work of verified experts into benchmarks and RL environments: voice elicitation, sandbox authoring, deterministic verification, and cross-review.
Why the input matters most
Everything downstream of a bad task is wasted compute. Our methodology over-invests at the source: reaching working practitioners in occupations frontier labs cannot staff, and capturing how their work actually gets done before a single environment is built.
Elicitation: conversations reach work nobody wrote down
Supply comes through our private expert network, which recruits working specialists into paid, identity-verified cohorts. Each engagement captures the expert’s real day-to-day: the software they use, the decisions they make, and why.
Experts are told plainly that their contributions train and evaluate AI systems, and are paid for every stage of participation.
Extraction and confirmation
Elicited work is distilled into structured task definitions with an explicit notion of done, and the expert reviews and corrects those definitions themselves. Nothing enters the corpus on a model’s word alone.
Authoring: the expert’s run is the ground truth
The expert then performs their own tasks inside a recorded sandbox environment on real software. That cold run becomes the human baseline, and it grounds every check that follows: assertions are derived from what the expert actually did, not from what a task writer imagined.
Verification: the ledger ties
Wherever the software allows it, grading is code: database assertions, numeric reconciliation within expert-set tolerance bands, file diffs against acceptable outcome sets, audit-log checks, and invariants whose violation ends the episode. Model-graded criteria are capped at a fifth of any corpus, labeled, and expert-audited.
Two rules are absolute:
- Every verifier must reproduce the authoring expert’s own recorded run before it ships.
- Ground truth is a set. When two independent expert runs diverge, the divergence is captured as tolerance, not averaged away.
Independence: author, performer, adjudicator
No expert grades their own work. One authors, a second runs the task blind, a third adjudicates divergence. Elite practitioners are promoted into a reviewer tier that cross-reviews every packet, and later cohorts independently corroborate earlier cohorts’ tasks before anything is sold.
Publication: public cuts, sacred holdouts
Benchmarks publish a curated public cut with a leaderboard, a reproducible harness, and failure analysis. Held-out sets never publish and are never seeded into training data. That separation is enforced in our registry, not by convention.
Refresh: priced to the model clock
Environments compound. Every frontier release gets re-run leaderboards, fresh variants, harder difficulty bands, and a new failure corpus. A pinned environment with a deterministic verifier gets more valuable every time a new model ships.