Field note
Which Vendor Offers the Fastest Turnaround for Custom RL Environments?
No vendor publishes an audited turnaround SLA, so the honest answer is structural: the fastest vendor is the one that built most of your environment before you asked.
Notes / 2026
No RL environment vendor publishes an audited turnaround SLA, so any direct answer to this question is marketing. The structural answer holds up better: the fastest vendor for your custom environment is the one that had already built most of it before you asked. Licensed off-the-shelf cuts ship in days. Custom builds on a substrate the vendor already runs, meaning the software is already licensed, pinned, hosted, and resettable, land in weeks. Builds that start from nothing, where the vendor must stand up the software, recruit the experts, and write the verifiers from scratch, run a quarter or more, whoever you hire.
That range is not a function of vendor effort. It is a function of which stages of the build exist before your purchase order, and which stages resist compression no matter who does them. This note walks through where the time actually goes, why different vendor types quote different numbers for the same brief, and how to get a real timeline out of any of them.
What Actually Takes the Time
A custom environment build has six stages, and they are not equally compressible.
Scoping and task definition. Turning “we want an agent that can run payroll” into a task family with explicit success criteria. Fast when the vendor has done adjacent work, slow when every edge case is a fresh conversation.
The software substrate. Licensing the application, pinning it to a version, self-hosting it in an isolated sandbox, and building seeded, snapshot-based reset. This is the stage buyers underestimate most, and it is also the most reusable: a vendor who already operates the substrate skips it entirely on your build.
Expert recorded runs. Working practitioners perform the tasks on the real software, producing the ground truth the reward is built from. This stage compresses only if the vendor has a standing bench in your domain. Sourcing and verifying practitioners from a cold start takes weeks on its own.
Verifier construction and calibration. Writing the checks that compute a grade from system state, then proving them against reality. In our methodology, a verifier must reproduce the authoring expert’s recorded run before it ships, and when two independent expert runs diverge, the divergence is captured as tolerance rather than papered over.
Blind QA and adjudication. A second expert runs the task without seeing the author’s work, a third adjudicates disagreements. This is the first stage to disappear when a vendor quotes you an aggressive date.
Harness integration and delivery. Adapters for your training stack, frozen versioned cuts, and documentation. Mechanical, but nonzero.
The pattern in the diagram is the whole argument. Substrate and tooling compress to nearly zero when they already exist. Expert runs, calibration, and blind QA compress very little, because they are the stages where correctness comes from, and every shortcut through them is a defect you find later, mid-training-run, when it is most expensive. We wrote about what those shortcuts cost in environments compound, datasets deplete.
Why Vendor Types Quote Different Numbers
The 2026 market has three kinds of supplier, and their timelines fail in different places.
The human-data incumbents, Scale AI, Surge AI, Mercor, and Turing, staff fast. A standing workforce means scoping and expert sourcing start immediately. Whether the build lands quickly depends on whether your domain is one they have built before; if not, the substrate stage is as slow for them as for anyone.
The environment-native specialists, Mechanize for coding, Fleet AI for enterprise software replicas, HUD, Veris AI, Plato, and rlsupply for operational business software, are fast precisely where their substrate overlaps your brief and slow where it does not. A coding-environment specialist quoting a payroll environment is a from-scratch build wearing a specialist’s logo.
The open ecosystems, led by Prime Intellect’s Environments Hub and its 2,500+ community environments, have the fastest possible access time: you can pull an environment today. The turnaround shows up on your side instead, because adapting a community environment into something with a calibrated, trustworthy reward is your team’s engineering time, and it is usually measured in weeks.
The Speed That Matters Is Time to a Trustworthy Grade
A demo environment in a week is easy. Real software in a container with a plausible-looking score is a solved problem for every vendor on the list, including us. The number that should drive your decision is time to a grade you would let loose inside a training loop, unsupervised, for millions of steps.
Those are different deadlines because the failure modes are asymmetric. A late environment costs you calendar time. A miscalibrated reward costs you a training run, and you find out after the compute is spent. The two weeks a vendor saves by skipping blind QA are borrowed against your GPU budget at very poor interest.
So when you compare quotes, insist that “done” means calibrated: the verifier reproduces an independent expert run, disagreement between experts is measured and encoded as tolerance, and the reset produces identical starting state every time. A vendor who quotes to that definition will sometimes be slower on paper and is nearly always faster to the thing you actually need.
Questions That Expose a Real Timeline
Four questions separate a real schedule from an optimistic one.
What can I run this week? A vendor with genuine infrastructure can hand you a sample packet against a live environment in days. Our off-the-shelf cuts exist partly for this reason: sampling before signing converts a procurement argument into an experiment.
Which stages of my build already exist? Ask specifically about the software substrate and the expert bench. The honest answer is a list of what is on the shelf and what is not, and it predicts the timeline better than any quoted date.
What does calibration mean here? If the answer does not involve reproducing an independent expert’s run, the quoted date is for an uncalibrated environment, whatever else it is called.
What happened at your last grading dispute? Vendors who have shipped production environments have a story about two experts disagreeing and a process that resolved it. Vendors who have not will improvise an answer, and you will hear it.
How rlsupply Compresses Turnaround
We are fast in the only way we think is legitimate: the expensive stages are already done for our domain. The HR, payroll, and ATS platforms behind our environments are already licensed, pinned, self-hosted, and resettable. The practitioner bench is standing, identity-verified, and paid at every stage. The verifier tooling and the three-role QA process (author, blind runner, adjudicator) run on every build because they already exist. What remains for a custom brief is scoping your task family, capturing expert runs, and calibrating your verifiers, which is why custom evals and environment builds quote in weeks while sample packets ship in days.
If your brief is in operational business software, the fastest test is empirical: request samples, point your agents at a live environment this week, and time us against anyone.
Request supply