RL Environments
Resettable, seeded, deterministically verified RL environments built on real business software. Designed to survive thousands of eval runs a day, priced per environment with variants at scale.


Computer-Use & Browser Environments
Computer-use and browser-use RL environments for GUI agents: full Windows and Linux desktops, real business software, simulated users, deterministic resets, and verifiable rewards at eval scale.

Code Generation & IDE Environments
Agentic coding environments that grade system state, not diffs: integration engineering, fault injection, data migration, and terminal workflows against live containerized services.
Environments built for training, not demos
Most agent environments fall over in production evaluation: state pollutes between runs, graders drift, and the “environment” is a mockup of the software rather than the software. Ours are built the other way around.
Every rlsupply environment is real business software, self-hosted at a pinned version inside an isolated sandbox. An agent gets the same starting state a working professional would inherit: populated databases, historical records, messy edge cases carried in from the authoring expert’s own workflows.
The reward is the hard part, so it ships first
An environment without a trustworthy reward signal is a screenshot generator. Each environment ships with its reward blueprint: expert-weighted rubrics compiled into deterministic verifiers that assert against the environment’s database, reconcile its ledgers, and check its audit trail. Model-graded criteria are capped, labeled, and expert-audited.
Before any environment ships, its verifiers must reproduce the authoring expert’s own recorded run. If the check cannot recognize expert work as correct, it does not ship.
Built against a stated operational pain
Teams running serious post-training tell us the bottleneck is operational: running thousands of eval jobs a day means environments must reset, restore, and rerun without human intervention. That constraint shaped the architecture, not the other way around.
What ships in the box
HR, payroll, ATS, and other business platforms running as pinned versions inside isolated sandboxes. Linux and full Windows desktops.
Every episode starts from a known seed. State resets between runs, snapshots restore any prior state, and parallel episodes never share a database.
Expert rubric weights compile into reward functions. Required facts become intermediate checkpoints, disqualifiers become terminal penalties. RLVR-ready.
Every click, keystroke, API call, and state change is recorded and exported. Episode dumps ship as JSONL through a standard harness adapter.
The same task under different seeds, where the interesting variants change which acceptable outcome applies, not just the numbers.
Environments deploy into your own cloud via Terraform and Helm for teams whose data cannot leave. Licensing available per environment or per fleet.
See it before you buy it.
Sample task packets and environment access for evaluation. Tell us the capability you care about and we will send the relevant cut.