Research
Benchmarks, leaderboards, and findings from the rlsupply data bench. Public cuts are open source; held-out sets are licensed.
BenchmarkBenchmark
Recruiter-Bench
Long-horizon, multi-stakeholder recruiting workflows: sourcing, screening, coordination, and recovery when humans behave like humans. Most agents fail the hard band.
10+hard-band scenarios
2difficulty bands
multistakeholder simulation
ATS-Bench
Can agents build, fix, harden, and migrate integrations against a live applicant tracking system? 50 public tasks, 200 private, graded by webhook-verified system state.
50public tasks
200private held-out tasks
4task categories
In development
An entity-resolution benchmark over a large synthetic candidate corpus is in progress, alongside new occupation cohorts. Every frontier model release triggers a leaderboard re-run.