Recruiter-Bench

Long-horizon, multi-stakeholder recruiting workflows: sourcing, screening, coordination, and recovery when humans behave like humans. Most agents fail the hard band.

10+hard-band scenarios
2difficulty bands
multistakeholder simulation
e2eoutcome-graded episodes

Leaderboard

RankModelHard-band pass rateCompletion
1 GPT-5.6 Sol303/10 hard tasks
Claude Fable 5running
Claude Opus 5running
Gemini 3.1 Prorunning

Preliminary results from the hard difficulty band. Most agents fail complex variants. Full results publish with the benchmark release.

Task categories

Source and screen against a live requisition

Coordinate candidates who stall, refuse, or reply gibberish

Recover a pipeline after a candidate blocks the agent

Close the loop across hiring stakeholders

Why recruiting workflows

Recruiting is a canonically hard agent domain: long horizons, many stakeholders, and success criteria that depend on outcomes across an entire pipeline rather than a single artifact. It is also work we know at production depth, drawn from years of operating recruiting software used by real staffing teams.

How it is built

Each episode drops the agent into a live requisition with simulated stakeholders who behave like people: candidates stall, refuse, reply with gibberish, or block the agent outright. Tasks are graded end to end on pipeline outcomes, with rubric criteria for process quality along the way.

Two difficulty bands ladder from standard coordination to adversarial recovery scenarios. The hard band is calibrated so that current frontier agents fail most variants, keeping headroom legible as models improve.

Early findings

Agents handle happy-path coordination but degrade quickly when stakeholders deviate: a single uncooperative candidate is often enough to derail an entire pipeline. Recovery behavior, not planning, is the discriminating capability. Full results publish with the benchmark release.