Verifiers & Rubrics for Verticals
Expert-weighted rubrics and deterministic verifiers for vertical AI: payroll, claims, billing, recruiting, and back-office domains. The grading is what matters, and ours compiles to code.

The grading is what matters
Researchers who buy environments say it plainly: pixel-perfect realism is negotiable, the grading is not. A quarter of flagship agent benchmarks accept incorrect solutions, and LLM-as-judge scoring drifts with the judge. Rubric quality, and how usable the graded data is against real-world workflows, is how serious buyers now evaluate vendors.
So we sell the grading as its own product. Expert-weighted rubrics compiled into deterministic verifiers, portable into your own eval stack.
What a vertical rubric actually contains
A payroll close, a claims adjustment, or a requisition has an explicit notion of done: required outcomes with tolerance bands, mandatory invariants that must hold throughout, and disqualifying mistakes that end the episode. Practitioners weight each criterion. The result compiles to code: the record exists in the right state, the ledger ties, the required intermediate steps appear in the audit trail.
Why determinism is a discipline, not a preference
We hold two rules absolute. Model-graded criteria never exceed a fifth of any corpus, are always labeled, and are always audited by a second expert. And every verifier must pass the authoring expert’s own recorded run before it ships. That second rule regularly catches grading bugs that would otherwise ship silently, which is precisely the point.
What ships in the box
Criteria and weights set by the experts who live with the consequences of the work, not by contractors guessing what matters.
Database assertions, numeric reconciliation, file diffs, audit-log checks, and invariant assertions that grade agents in code.
Where judgment is genuinely subjective, model-graded criteria are capped at a fifth of any corpus, always labeled, always expert-audited.
Every verifier must reproduce the authoring expert's own recorded run before it ships. If it cannot recognize expert work, it does not ship.
Rubrics and verifiers license independently, so you can bring audited grading to environments and datasets you already run.
Payroll and recruiting are live. New verticals are sourced through the same expert pipeline, from claims adjusting to medical billing.
See it before you buy it.
Sample task packets and environment access for evaluation. Tell us the capability you care about and we will send the relevant cut.