Blog
Notes from the rlsupply team on environments, verification, benchmarks, and the economics of expert data.
- →
Introducing ATS-Bench: grading agents on what the system says, not what the code looks like
Our first public benchmark evaluates agent integration engineering against a live applicant tracking system, with webhook-verified grading and a 200-task private holdout.
- →
Environments compound. Datasets deplete.
Why we sell pinned environments with deterministic verifiers instead of one-shot preference data, and what that means for teams buying post-training data.