Recruiter-Bench
Long-horizon, multi-stakeholder recruiting workflows: sourcing, screening, coordination, and recovery when humans behave like humans. Most agents fail the hard band.
Leaderboard
| Rank | Model | Hard-band pass rate | Completion | |
|---|---|---|---|---|
| 1 | 30 | 3/10 hard tasks | ||
| – | running | – | ||
| – | running | – | ||
| – | running | – |
Preliminary results from the hard difficulty band. Most agents fail complex variants. Full results publish with the benchmark release.
Task categories
Source and screen against a live requisition
Coordinate candidates who stall, refuse, or reply gibberish
Recover a pipeline after a candidate blocks the agent
Close the loop across hiring stakeholders
Why recruiting workflows
Recruiting is a canonically hard agent domain: long horizons, many stakeholders, and success criteria that depend on outcomes across an entire pipeline rather than a single artifact. It is also work we know at production depth, drawn from years of operating recruiting software used by real staffing teams.
How it is built
Each episode drops the agent into a live requisition with simulated stakeholders who behave like people: candidates stall, refuse, reply with gibberish, or block the agent outright. Tasks are graded end to end on pipeline outcomes, with rubric criteria for process quality along the way.
Two difficulty bands ladder from standard coordination to adversarial recovery scenarios. The hard band is calibrated so that current frontier agents fail most variants, keeping headroom legible as models improve.
Early findings
Agents handle happy-path coordination but degrade quickly when stakeholders deviate: a single uncooperative candidate is often enough to derail an entire pipeline. Recovery behavior, not planning, is the discriminating capability. Full results publish with the benchmark release.