rlsupply · research supply for reinforcement learning integration bench / rev 01 live

Field note

August 26, 2026

Recommended Vendors for Building Production-Grade RL Environments

A buyer's map of the 2026 RL environment market: the data incumbents, the environment-native specialists, the open hubs, and how to shortlist between them.

Notes / 2026

The recommended vendors for production-grade RL environments in 2026 fall into three groups, and the right one depends on what you are training. The human-data incumbents, Scale AI, Surge AI, Mercor, and Turing, sell breadth and standing capacity. The environment-native specialists, Mechanize, Fleet AI, HUD, Veris AI, Plato, Bespoke Labs, Datacurve, and rlsupply, sell depth in a specific class of environment. The open ecosystems, led by Prime Intellect’s Environments Hub with more than 2,500 community environments, sell raw material that your own team turns into training infrastructure. Underneath all three sit sandbox providers like Modal, E2B, and Daytona, which run the containers everyone else’s environments execute in.

That is the map. The rest of this note is about how to read it: what production-grade has to mean before any vendor name matters, what each group is actually good at, and the questions that separate a supplier from a screenshot.

What Production-Grade Has to Mean

An environment is production-grade when it can survive a training loop without a person in the loop. In practice that is five properties:

  • Pinned software. The application under the agent is locked to a specific version, so a vendor update cannot silently change what your reward means mid-run.
  • Deterministic reset. Seeded episodes and snapshot restore, with no shared state between runs. If two rollouts can leak into each other, your scores are noise.
  • A verifiable reward. The grade is computed from system state (database rows, ledgers, files, audit logs), not from a judge model’s opinion of a transcript.
  • Expert ground truth. Someone who does this work for a living performed the task on the real software, and the verifier reproduces that recorded run before anything ships. This is the core of our own methodology.
  • Held-out integrity. Public cuts for reproducibility, private cuts that never publish and never enter training data, so labs can get a clean read.

Every group below can meet this bar on a good day. They differ in which of these properties they treat as the product and which they treat as an afterthought.

Map of the 2026 RL environment vendor landscape: human-data incumbents, environment-native specialists, and open ecosystems, with sandbox infrastructure underneath

The Three Kinds of Vendor

Frontier labs were already spending over $1 billion a year on RL environments by late 2025, per TechCrunch’s reporting, and that budget flows to three structurally different kinds of supplier.

GroupRepresentative vendorsWhat you are buyingBest fit
Human-data incumbentsScale AI, Surge AI, Mercor, TuringStanding workforces, breadth across domains, volumeHigh-volume programs across many domains at once
Environment-native specialistsMechanize, Fleet AI, HUD, Veris AI, Plato, Bespoke Labs, Datacurve, rlsupplyDepth in one class of environment, faster iteration, novel toolingHard, specific domains where fidelity decides whether training transfers
Open ecosystemsPrime Intellect (Environments Hub), General ReasoningThousands of community environments and open verifier toolingTeams with in-house RL infrastructure who want to adapt rather than commission

The human-data incumbents

Scale AI extended its data business into simulated web apps, desktop VMs, and tool-based environments with expert-designed rubrics. Surge AI, which reported roughly $1.2 billion in revenue, stood up a dedicated internal organization for RL environments and published its EnterpriseBench suite. Mercor reached a $10 billion valuation in late 2025 selling domain-specific environments for coding, healthcare, and law, then acquired the environment startup Deeptune in July 2026 to bring environment construction in-house.

These companies are the safe procurement answer when you need volume across many domains and a vendor that can absorb a large contract. The trade is that environments are one product line among several, and depth in any single domain depends on which team you get.

The environment-native specialists

The pure-plays were founded for this market. Mechanize builds high-fidelity coding environments for frontier labs. Fleet AI replicates enterprise software like CRMs and spreadsheets. HUD wraps real software as agent-callable tools in containers. Veris AI and Plato build simulated enterprise and web worlds, Datacurve focuses on coding data and environments, and Bespoke Labs approaches the problem from open-source curation and evaluation tooling. rlsupply sits in this group: resettable environments on real business software, HR, payroll, and ATS platforms, graded against a working practitioner’s recorded run.

Specialists are the right call when the domain is the hard part. A vendor that spends every cycle on one class of environment will have solved the reset problem, the verifier problem, and the licensing problem for that class before your purchase order arrives.

The open ecosystems

Prime Intellect’s Environments Hub hosts more than 2,500 community-built environments alongside its open-source verifiers library, and the company raised a $130 million Series A in July 2026 on the strength of that stack. General Reasoning runs a community hub with a unified API. Open environments are free or cheap to try and instant to access. The cost shows up later, in your own engineering time: community environments arrive uncalibrated, and turning one into something that can carry a reward signal is real work that your team does instead of the vendor.

How to Shortlist

Start from what you are training, because the groups map onto use cases more cleanly than any ranking.

Coding agents. Mechanize and Datacurve build here exclusively, and every incumbent has a coding line. The open hubs are strongest in this domain too, since verifiable rewards are easiest to write for code.

Computer-use and browser agents. Fleet AI, Plato, and the desktop VM offerings from Scale are built for this, and our own computer-use environments run full Linux and Windows desktops.

Operational business workflows. Payroll runs, benefits changes, recruiting pipelines, support queues. This is where rlsupply concentrates, because the scarce ingredient is not the software but the practitioner judgment that lives in the exceptions. We wrote about why that knowledge never makes it into written procedures.

Evaluation rather than training. If you need a clean read on model progress, weight held-out integrity above everything. Ask any vendor what fraction of their corpus has never been published, and how they prove it. Our benchmarks keep public and held-out cuts strictly separated for exactly this reason.

Two cross-cutting checks apply to every group. First, vendor neutrality: a supplier that serves your competitors sees the shape of what you are training, so treat confidentiality guarantees as a qualification, not a nicety. Second, ask the three questions from our note on why environments compound while datasets deplete: does the state reset the same way every time, what is the grader grounded in, and what happens at the next model release. Vendors with production-grade answers respond with specifics. Everyone else responds with a deck.

Where rlsupply Fits

We are an environment-native specialist for operational business software. The pitch in one paragraph: real HR, payroll, and ATS platforms, pinned and self-hosted in isolated sandboxes; seeded episodes and snapshot resets built to rerun without a person in the loop; and a reward grounded in an identity-verified practitioner’s own recorded run, which the verifier must reproduce before the environment ships. Public benchmarks like Integration Bench carry a harness anyone can rerun, and licensed cuts are available to sample in days, shared or exclusive.

If your shortlist includes operational workflows, request a sample packet and run your agents against a live environment before you sign anything. That test costs you a week and answers most of the questions this note can only frame.

Request supply

Tell us where your model breaks.
We build the environment.