AI model improvement index · 2026

RL environments, training data & evaluations.

An index of companies building RL environments, training data, benchmarks and evaluations for frontier labs.

Training a frontier model now means buying tasks, environments and graders from the AI training data and RL environments market. This index tracks 42 companies selling into that market and ranks 30 of them on five factors: substance, realism, verification, training value and outside validation. The other twelve are listed without a rank, because we couldn’t find enough finished, attributable work to score them. All of this is read off published material: publishing little results in lower confidence. Each entry links out to what it was scored on.

Leaderboard · September 2026

AI model improvement providers

30 ranked providers · 12 monitored (NR). Expand an entry for its role, evidence and sources. Equal assessments share a rank.

01Scale AIBuilds expert-designed frontier evaluations and training data, with public methods, held-out controls, and broad model coverage.EnterpriseCyberHumanity's Last ExamScaled providerSan Francisco, USExtensive

Builds expert-designed frontier evaluations and training data, with public methods, held-out controls, and broad model coverage.

Humanity's Last Exam publishes benchmark construction, evaluation methods, results, and a held-out integrity design. The work is jointly attributed to Scale AI and the Center for AI Safety; Scale Labs also operates a broader set of continuously updated model evaluations. Scale's ASPI and PropensityBench cyber track are scored separately on RL Index Cyber.

EnterpriseCyber
Evidence snapshot
Humanity's Last Exam
Attributed role
Frontier evaluation and training-data provider
Provider type
Scaled provider
Location
San Francisco, US
Evidence
Extensive
Visit company website
02AfterQueryDevelops expert data and agent evaluations, including a task-linked vulnerability-analysis benchmark with public graded cases.CodeEnterpriseCyberVADER; public graded casesSpecialistSan Francisco, USDocumented

Develops expert data and agent evaluations, including a task-linked vulnerability-analysis benchmark with public graded cases.

VADER reports methods and model results across real code vulnerability cases. The repository exposes task metadata, graded outputs, and scoring code, allowing the reported evaluation to be inspected beyond a marketing claim.

CodeEnterpriseCyber
Evidence snapshot
VADER; public graded cases
Attributed role
Evaluation, environment, and expert-data developer
Provider type
Specialist
Location
San Francisco, US
Evidence
Documented
Visit company website
03Surge AIDelivers human feedback, expert data, and agent environments used in frontier-model training and evaluation.EnterpriseLong HorizonAnthropic RLHF; workplace RLScaled providerSan Francisco, USExtensive

Delivers human feedback, expert data, and agent environments used in frontier-model training and evaluation.

Surge documents Anthropic's use of its RLHF platform and publishes measured evaluations from a 150-task workplace environment. Its research record also attributes joint work with Anthropic, OpenAI, Meta, and academic groups; each collaboration is counted once rather than as separate validation signals.

EnterpriseLong Horizon
Evidence snapshot
Anthropic RLHF; workplace RL
Attributed role
Scaled RLHF, expert-data, and environment provider
Provider type
Scaled provider
Location
San Francisco, US
Evidence
Extensive
Visit company website
04MechanizeBuilds software-engineering environments and evaluations using executable tasks, deterministic emulation, and outcome-based grading.CodeGBA EvalSpecialistSan Francisco, USDocumented

Builds software-engineering environments and evaluations using executable tasks, deterministic emulation, and outcome-based grading.

GBA Eval publishes its scope, harness design, grading approach, and model results. It is a bounded coding evaluation rather than evidence for every software-engineering capability, but it provides an inspectable example of Mechanize's environment methodology.

Code
Evidence snapshot
GBA Eval
Attributed role
Software-engineering environment and evaluation developer
Provider type
Specialist
Location
San Francisco, US
Evidence
Documented
Visit company website
05MercorCombines a large expert network with public professional-work benchmarks and commercial RL-environment delivery.CodeEnterpriseLong HorizonAPEX; APEX-Agents; APEX-SWEScaled providerSan Francisco, USDocumented

Combines a large expert network with public professional-work benchmarks and commercial RL-environment delivery.

Mercor publishes the APEX benchmark family with task descriptions, data or code links, and measured model results across professional, agentic, and software-engineering work. Customer-count claims are not treated as independent validation without attributable evidence.

CodeEnterpriseLong Horizon
Evidence snapshot
APEX; APEX-Agents; APEX-SWE
Attributed role
Expert-data, benchmark, and environment provider
Provider type
Scaled provider
Location
San Francisco, US
Evidence
Documented
Visit company website
06Bespoke LabsDevelops open datasets, executable agent tasks, and post-training tooling with reproducible public releases.CodeLong HorizonCyberOpenThoughts-TBLiteSpecialistMountain View, USExtensive

Develops open datasets, executable agent tasks, and post-training tooling with reproducible public releases.

OpenThoughts-TBLite publishes task definitions, tests, and aggregate results with joint attribution. Nous Research separately distributes the dataset, corroborating outside integration while preserving collaborators' credit. OpenThoughts-TBLite's security-task subset is scored separately on RL Index Cyber.

CodeLong HorizonCyber
Evidence snapshot
OpenThoughts-TBLite
Attributed role
Joint dataset, task, and post-training tooling contributor
Provider type
Specialist
Location
Mountain View, US
Evidence
Extensive
Visit company website
07Handshake AICombines an expert network with open verifier research and realistic professional-work evaluation environments.EnterpriseLong HorizonGandalf; ATLAS-FinanceScaled providerSan Francisco, USDocumented

Combines an expert network with open verifier research and realistic professional-work evaluation environments.

Handshake publishes Gandalf, an environment-aligned verifier, with results on 3,204 expert-graded judgments from BankerVerifierBench. ATLAS-Finance extends that work to 100 tasks across 13 workplace environments with expert-authored rubrics and public code.

EnterpriseLong Horizon
Evidence snapshot
Gandalf; ATLAS-Finance
Attributed role
Expert-data, verifier, and professional-environment developer
Provider type
Scaled provider
Location
San Francisco, US
Evidence
Documented
Visit company website
08DatacurveDevelops long-horizon coding data and evaluations with original repository tasks and programmatic verification.CodeDeepSWESpecialistSan Francisco, USDocumented

Develops long-horizon coding data and evaluations with original repository tasks and programmatic verification.

DeepSWE publishes 113 original tasks across active repositories, isolated environments, and program-based verifiers. The repository exposes task manifests and grading artifacts, making the benchmark inspectable at task level.

Code
Evidence snapshot
DeepSWE
Attributed role
Coding benchmark and training-data developer
Provider type
Specialist
Location
San Francisco, US
Evidence
Documented
Visit company website
09Collinear AIBuilds task-linked evaluations and runnable training environments, including secure code repair with functional and security checks.CodeCyberLong HorizonCWE-Bench; YC-BenchSpecialistSan Francisco, USExtensive

Builds task-linked evaluations and runnable training environments, including secure code repair with functional and security checks.

CWE-Bench publishes methods and task-linked results. Google reports a completed Gemini evaluation run by Collinear on the benchmark. YC-Bench evaluates long-horizon agent planning over a simulated one-year startup run, with published methods, results across 12 proprietary and open-source models at three seeds each, and an open-source implementation.

CodeCyberLong Horizon
Evidence snapshot
CWE-Bench; YC-Bench
Attributed role
Benchmark and model-improvement environment developer
Provider type
Specialist
Location
San Francisco, US
Evidence
Extensive
Visit company website
10TuringBuilds expert-authored coding, terminal, and STEM training tasks with executable verification and published benchmark results.CodeEnterpriseCyberSWE-Bench++; Terminal-Bench dataScaled providerPalo Alto, USDocumented

Builds expert-authored coding, terminal, and STEM training tasks with executable verification and published benchmark results.

SWE-Bench++ publishes more than 500 public and 7,000 additional repository tasks with executable tests and an evaluation and training framework. Turing also documents verifier-graded terminal and frontier STEM data, including successful trajectories used for fine-tuning. Turing's CyberStrike evaluation is scored separately on RL Index Cyber.

CodeEnterpriseCyber
Evidence snapshot
SWE-Bench++; Terminal-Bench data
Attributed role
Expert-data and benchmark developer
Provider type
Scaled provider
Location
Palo Alto, US
Evidence
Documented
Visit company website
11Gray Swan AIBuilds adversarial evaluations and protective methods with substantial joint research and outside model-assessment use.CyberAgentHarm; ARTEMIS; IPISpecialistPittsburgh, USExtensive

Builds adversarial evaluations and protective methods with substantial joint research and outside model-assessment use.

Published agent evaluations, protective tests, and implemented attack methods establish technical depth. NIST and Anthropic materials corroborate outside use; joint research retains its collaborators' credit.

Cyber
Evidence snapshot
AgentHarm; ARTEMIS; IPI
Attributed role
Joint research contributor and adversarial evaluation developer
Provider type
Specialist
Location
Pittsburgh, US
Evidence
Extensive
Visit company website
12Andon LabsPublishes long-horizon, computer-use, and physical-agent evaluations with inspectable setups and measured model results.Long HorizonUser SimulationVending-Bench 2; Drone-Bench; Blueprint-Bench 2SpecialistNot listedDocumented

Publishes long-horizon, computer-use, and physical-agent evaluations with inspectable setups and measured model results.

Andon Labs publishes detailed environments, scoring methods, and repeated model runs across Vending-Bench 2, Drone-Bench, and Blueprint-Bench 2. The benchmarks cover distinct outcomes and are not combined merely because they share a publisher.

Long HorizonUser Simulation
Evidence snapshot
Vending-Bench 2; Drone-Bench; Blueprint-Bench 2
Attributed role
Agent benchmark and evaluation developer
Provider type
Specialist
Location
Not listed
Evidence
Documented
Visit company website
13ProximalDevelops original ultra-long-horizon coding tasks with external contributors and measured frontier-model results.CodeLong HorizonFrontierSWESpecialistNot listedDocumented

Develops original ultra-long-horizon coding tasks with external contributors and measured frontier-model results.

FrontierSWE publishes original tasks, model results, long execution budgets, and contributions from academic and industry partners. The evidence supports difficult software-engineering evaluation, not every model-improvement domain.

CodeLong Horizon
Evidence snapshot
FrontierSWE
Attributed role
Long-horizon coding benchmark and data developer
Provider type
Specialist
Location
Not listed
Evidence
Documented
Visit company website
14Prime IntellectTrains and open-sources large RL post-trained models and operates a public hub of community RL environments.CodeComputer UseEnterpriseINTELLECT-3; Environments HubOpen / hybridSan Francisco, USDocumented

Trains and open-sources large RL post-trained models and operates a public hub of community RL environments.

INTELLECT-3 is a 106B-parameter open-weight model post-trained with large-scale reinforcement learning, released with a technical report and published benchmark scores (MATH-500 98.1, AIME 2024 90.8, AIME 2025 88.0, GPQA 74.4). The Environments Hub separately hosts more than 2,500 community-contributed RL environments used for training and evaluation; environment count is treated as platform breadth, not independent validation. Prime Intellect's Security Verifiers and hosted Strix-XSS training are scored separately on RL Index Cyber.

CodeComputer UseEnterpriseLong HorizonCyberUser SimulationMulti-Agent
Evidence snapshot
INTELLECT-3; Environments Hub
Attributed role
Open post-training model and RL-environment developer
Provider type
Open / hybrid
Location
San Francisco, US
Evidence
Documented
Visit company website
15RefreshPublishes verifiable coding datasets and computer-use tooling with measured out-of-distribution training gains.CodeComputer UseGauntlet 4K; computer-1SpecialistNot listedDocumented

Publishes verifiable coding datasets and computer-use tooling with measured out-of-distribution training gains.

Gauntlet 4K reports a threefold Terminal-Bench 2.0 improvement after post-training on 4,000 pytest-verifiable environments. The larger Gauntlet release and open computer-1 harness provide inspectable tasks, verification, and execution tooling.

CodeComputer Use
Evidence snapshot
Gauntlet 4K; computer-1
Attributed role
Coding-data, environment, and computer-use tooling developer
Provider type
Specialist
Location
Not listed
Evidence
Documented
Visit company website
16Snorkel AIBuilds expert-curated training and evaluation data, including verifiable professional computer-use tasks.Computer UseEnterpriseCyberCUA-Bench+; KiCadScaled providerRedwood City, USExtensive

Builds expert-curated training and evaluation data, including verifiable professional computer-use tasks.

Snorkel publishes a 25-task KiCad benchmark built on Cua-Bench with objective saved-artifact checks and measured model results. CUA-Bench+ extends the work into a commercial data series with expert review, programmatic validation, and difficulty calibration. Snorkel's ShadowRelay and OpenThoughts-TBLite security tasks are scored separately on RL Index Cyber.

Computer UseEnterpriseCyber
Evidence snapshot
CUA-Bench+; KiCad
Attributed role
Expert-data and computer-use benchmark developer
Provider type
Scaled provider
Location
Redwood City, US
Evidence
Extensive
Visit company website
17Invisible TechnologiesDelivers scaled expert training and evaluation programs, with named customer work and outcome-linked case evidence.EnterpriseCohere eval; targeted trainingScaled providerDistributedExtensive

Delivers scaled expert training and evaluation programs, with named customer work and outcome-linked case evidence.

Invisible identifies Cohere as an evaluation customer and reports expert annotation across specialized domains and languages. A separate case record reports evaluation of 45 internal models and an 87% quality improvement after targeted training; full task-level methods are not public.

Enterprise
Evidence snapshot
Cohere eval; targeted training
Attributed role
Scaled expert-data and evaluation provider
Provider type
Scaled provider
Location
Distributed
Evidence
Extensive
Visit company website
=18CuaProvides open computer-use environments, fleets, and a benchmark framework for verifiable desktop and mobile tasks.Computer UseCua-BenchInfrastructureNot listedExtensive

Provides open computer-use environments, fleets, and a benchmark framework for verifiable desktop and mobile tasks.

Cua-Bench publishes a task model, runner, evaluators, registry, and measured datasets across desktop and mobile surfaces. Snorkel's independent KiCad study uses the framework, adding outside execution beyond Cua's own demonstrations.

Computer Use
Evidence snapshot
Cua-Bench
Attributed role
Computer-use environment and benchmark developer
Provider type
Infrastructure
Location
Not listed
Evidence
Extensive
Visit company website
=18HUDBuilds expert-validated financial-analyst and robotics benchmarks with public datasets and measured model results.CodeComputer UseEnterpriseSheetBench-50; Assemble BenchSpecialistNot listedDocumented

Builds expert-validated financial-analyst and robotics benchmarks with public datasets and measured model results.

SheetBench-50 is a 50-task financial-analyst spreadsheet benchmark built with Sepal AI and validated by expert reviewers from PwC, Cisco, Charles Schwab, and Fannie Mae, with each task constrained to a single reproducible answer. Assemble Bench separately publishes 14 contact-rich robot-assembly tasks with an open dataset, repository, and measured dense-reward results on a DROID checkpoint. HUD's ZeroDayBench evaluation is scored separately on RL Index Cyber.

CodeComputer UseEnterpriseCyber
Evidence snapshot
SheetBench-50; Assemble Bench
Attributed role
Benchmark and robotics-environment developer
Provider type
Specialist
Location
Not listed
Evidence
Documented
Visit company website
=18Vals AIRuns independent held-out benchmarks across finance, legal, and coding, publishing methods and measured model scores.CodeCyberVals Index; VLAIRSpecialistSan Francisco, USExtensive

Runs independent held-out benchmarks across finance, legal, and coding, publishing methods and measured model scores.

The Vals Index is a GDP-weighted composite of finance, coding, and legal benchmarks with a public leaderboard of measured model scores. VLAIR separately tests four legal-AI products against a practicing-lawyer baseline across seven tasks, with results published per vendor and per task. An independent arXiv paper builds directly on Vals' Finance Agent v2 and calls it a reference benchmark. Vals AI's CyberBench evaluation is scored separately on RL Index Cyber.

CodeCyber
Evidence snapshot
Vals Index; VLAIR
Attributed role
Benchmark and evaluation developer
Provider type
Specialist
Location
San Francisco, US
Evidence
Extensive
Visit company website
21QuesmaPublishes open agentic coding benchmarks with public leaderboards reporting per-model pass rates, cost, and speed.CodeLong HorizonCyberCompileBench; OTelBenchSpecialistWarsaw, PolandDocumented

Publishes open agentic coding benchmarks with public leaderboards reporting per-model pass rates, cost, and speed.

CompileBench scores 26 models on 15 real-world build and compilation tasks, tracking pass@1, pass@3, cost, and completion time on a public leaderboard. OTelBench separately publishes 23 observability-instrumentation tasks across 11 languages with executable grading and results for 14 models. Quesma's BinaryAudit backdoor-detection benchmark is scored separately on RL Index Cyber.

CodeLong HorizonCyber
Evidence snapshot
CompileBench; OTelBench
Attributed role
Open coding-benchmark developer
Provider type
Specialist
Location
Warsaw, Poland
Evidence
Documented
Visit company website
=22LabelboxCombines scaled training-data delivery with public agent evaluations and simulated enterprise-work environments.EnterpriseLong HorizonCyberImplicit Intelligence; WorldSimScaled providerSan Francisco, USDocumented

Combines scaled training-data delivery with public agent evaluations and simulated enterprise-work environments.

Implicit Intelligence publishes 205 agent scenarios, scoring methods, and model results across contextual constraints. Horizon documents WorldSim environments and reward design; general platform scale is not counted as additional technical validation. A privacy-and-security scenario subset of Implicit Intelligence is scored separately on RL Index Cyber.

EnterpriseLong HorizonCyber
Evidence snapshot
Implicit Intelligence; WorldSim
Attributed role
Training-data, environment, and evaluation provider
Provider type
Scaled provider
Location
San Francisco, US
Evidence
Documented
Visit company website
=22TolokaDevelops private agent benchmarks, RL gyms, and expert-annotated data across business workflows and research domains.Long HorizonUser SimulationToloka Arena; Tau-bench extensionScaled providerNot listedDocumented

Develops private agent benchmarks, RL gyms, and expert-annotated data across business workflows and research domains.

Toloka Arena publishes model results from hidden, non-contaminated workflow tasks and offers the underlying RL gyms for training. Its Tau-bench extension documents databases, tools, policies, golden trajectories, and calibrated difficulty.

Long HorizonUser Simulation
Evidence snapshot
Toloka Arena; Tau-bench extension
Attributed role
Training-data, RL-gym, and evaluation provider
Provider type
Scaled provider
Location
Not listed
Evidence
Documented
Visit company website
=22VmaxPublishes RL post-training research on automated task generation and self-play, with measured gains on math, code, and SWE benchmarks.CodeCyberPROPEL; PopuLoRA; unix-ctfSpecialistSan Francisco, USExtensive

Publishes RL post-training research on automated task generation and self-play, with measured gains on math, code, and SWE benchmarks.

PROPEL trains task generators at a targeted solve rate, lifting learnable-frontier coding tasks from 10.1% to 20.0% (Qwen2.5-3B) and software-engineering tasks on unseen repositories from 9.8% to 19.6% (Qwen3.5-27B). PopuLoRA evolves populations of LoRA adapters under programmatic verification and beats single-agent baselines across ten math and coding benchmarks; an independent open-source reimplementation corroborates the method. unix-ctf separately reports procedural command-line environment generation with held-out, baseline-controlled training gains and is scored on RL Index Cyber.

CodeCyber
Evidence snapshot
PROPEL; PopuLoRA; unix-ctf
Attributed role
RL post-training method developer
Provider type
Specialist
Location
San Francisco, US
Evidence
Extensive
Visit company website
25Good Start LabsBuilds game-based benchmarks and training environments with a named publisher partner and measured deployed outcomes.Long HorizonArkadium Game Lab; Gin RummyOpen / hybridNot listedExtensive

Builds game-based benchmarks and training environments with a named publisher partner and measured deployed outcomes.

Good Start Labs attributes benchmark and training work to Arkadium, including Game Lab and a compact expert model that beat casual human players in reported trials. The public record includes the customer, objective, and measured outcome, though not the complete training dataset.

Long Horizon
Evidence snapshot
Arkadium Game Lab; Gin Rummy
Attributed role
Domain benchmark and agent-training developer
Provider type
Open / hybrid
Location
Not listed
Evidence
Extensive
Visit company website
26Chakra LabsBuilds deterministic computer-use environments with frame-accurate control and publishes model results on a demonstration benchmark.Computer UseDojo; dojo-bench-miniSpecialistNot listedDocumented

Builds deterministic computer-use environments with frame-accurate control and publishes model results on a demonstration benchmark.

Chakra documents the Dojo environment suite, deterministic state controls, and model performance on dojo-bench-mini. The public evidence is provider-authored and the full task set was not independently executed in this review.

Computer Use
Evidence snapshot
Dojo; dojo-bench-mini
Attributed role
Computer-use environment and trajectory-data developer
Provider type
Specialist
Location
Not listed
Evidence
Documented
Visit company website
27BenchFlowPublishes stateful workplace environments, agent benchmarks, datasets, and an environment framework for reproducible evaluation.CodeEnterpriseSkillsBench; ClawsBenchSpecialistNot listedDocumented

Publishes stateful workplace environments, agent benchmarks, datasets, and an environment framework for reproducible evaluation.

BenchFlow publishes SkillsBench, ClawsBench, and PostTrain with papers, datasets, repositories, and live environments. ClawsBench supplies five wire-compatible workplace services and task manifests rather than a static prompt set.

CodeEnterprise
Evidence snapshot
SkillsBench; ClawsBench
Attributed role
Agent benchmark and environment developer
Provider type
Specialist
Location
Not listed
Evidence
Documented
Visit company website
28General ReasoningMaintains open RL-environment infrastructure and publishes a long-horizon decision benchmark with measured model results.Long HorizonOpenReward; KellyBenchOpen / hybridNot listedDocumented

Maintains open RL-environment infrastructure and publishes a long-horizon decision benchmark with measured model results.

OpenReward serves more than 330 community environments through a common API, while KellyBench publishes a sequential football-market environment and evaluation results. Community environment counts are treated as platform breadth, not as General Reasoning-authored benchmarks.

Long Horizon
Evidence snapshot
OpenReward; KellyBench
Attributed role
Open environment infrastructure and benchmark developer
Provider type
Open / hybrid
Location
Not listed
Evidence
Documented
Visit company website
29Patronus AIDevelops evaluation datasets and judge models with public methods and results across financial reasoning and hallucination detection.Long HorizonFinanceBench; HaluBench; LynxSpecialistNot listedDocumented

Develops evaluation datasets and judge models with public methods and results across financial reasoning and hallucination detection.

FinanceBench publishes a 10,000-question financial evaluation and measured retrieval-system results. HaluBench and Lynx add public data, code, and comparative judge-model results; benchmark publication does not by itself establish training delivery.

Long Horizon
Evidence snapshot
FinanceBench; HaluBench; Lynx
Attributed role
Evaluation benchmark and judge-model developer
Provider type
Specialist
Location
Not listed
Evidence
Documented
Visit company website
30ArenaOperates large-scale human-preference evaluations and publishes open datasets, ranking methods, and commercial evaluation services.Chatbot Arena; Arena-HardData providerNot listedExtensive

Operates large-scale human-preference evaluations and publishes open datasets, ranking methods, and commercial evaluation services.

Arena publishes its leaderboard policy, ranking pipeline, open datasets, and academic methods. Its commercial evaluation product uses community feedback for model labs and developers; the public record demonstrates evaluation delivery rather than RL-environment production.

Evidence snapshot
Chatbot Arena; Arena-Hard
Attributed role
Human-preference evaluation platform
Provider type
Data provider
Location
Not listed
Evidence
Extensive
Visit company website
NRAIChampMarket relevance is plausible, but the reviewed public record does not establish a completed model-improvement artifact or delivery.EnterpriseLong HorizonQualifying evidence not verifiedSpecialistNot listedLimited

Market relevance is plausible, but the reviewed public record does not establish a completed model-improvement artifact or delivery.

No named benchmark, environment, dataset, or attributable customer result with sufficient public methods and outcomes was verified for this edition.

EnterpriseLong Horizon
Evidence snapshot
Qualifying evidence not verified
Attributed role
Provider under review
Provider type
Specialist
Location
Not listed
Evidence
Limited
Visit company website
NRAndromedeDevelops programmatically generated RL environments, tasks, and verifiers, but remains in private beta without a public completed result.Long HorizonPrivate beta; completed result not verifiedSpecialistNot listedLimited

Develops programmatically generated RL environments, tasks, and verifiers, but remains in private beta without a public completed result.

The public site describes a post-training and evaluation pipeline and a small partner beta. No named task set, measured evaluation, or independently attributable delivery was available for scoring in this review.

Long Horizon
Evidence snapshot
Private beta; completed result not verified
Attributed role
RL data lab under review
Provider type
Specialist
Location
Not listed
Evidence
Limited
Visit company website
NRFleet AIPublishes post-training infrastructure and describes high-fidelity environment delivery, but lacks task-linked public outcome evidence in this review.CodeEnterpriseQualifying outcome evidence not verifiedSpecialistNew York, USLimited

Publishes post-training infrastructure and describes high-fidelity environment delivery, but lacks task-linked public outcome evidence in this review.

Fleet's public site and Miles repository establish an active environment and post-training platform. This review did not verify a provider-attributable task set with published methods and measured outcomes, so the company remains monitored rather than scored.

CodeEnterprise
Evidence snapshot
Qualifying outcome evidence not verified
Attributed role
Environment platform under review
Provider type
Specialist
Location
New York, US
Evidence
Limited
Visit company website
NRHabitat IncIs described as a work-automation environment vendor, but no inspectable company-authored technical artifact was verified.Computer UseEnterpriseQualifying artifact not verifiedSpecialistNot listedLimited

Is described as a work-automation environment vendor, but no inspectable company-authored technical artifact was verified.

The reviewed record supports category fit and an active company. It does not expose a named environment, public methods, or measured results attributable to Habitat Inc; third-party directory inclusion does not qualify as evidence.

Computer UseEnterprise
Evidence snapshot
Qualifying artifact not verified
Attributed role
Work-automation environment provider under review
Provider type
Specialist
Location
Not listed
Evidence
Limited
Visit company website
NRHalluminateAppears in the model-improvement market, but a completed named environment, benchmark, or attributable delivery was not verified.EnterpriseQualifying artifact not verifiedSpecialistNot listedLimited

Appears in the model-improvement market, but a completed named environment, benchmark, or attributable delivery was not verified.

The current public record establishes category relevance but does not expose task-level methods, measurable outcomes, or a named customer delivery. Offering language and list inclusion do not satisfy the ranking gate.

Enterprise
Evidence snapshot
Qualifying artifact not verified
Attributed role
Provider under review
Provider type
Specialist
Location
Not listed
Evidence
Limited
Visit company website
NRHillclimbTargets AI-research training data and RL-environment creation, but has not published a current company-attributable artifact with outcomes.Current attributable artifact not verifiedSpecialistNot listedLimited

Targets AI-research training data and RL-environment creation, but has not published a current company-attributable artifact with outcomes.

Hillclimb states that it sells training data to frontier labs and is scaling RL-environment creation. Prior-team work and investor backing do not substitute for a named current artifact or delivered result.

Evidence snapshot
Current attributable artifact not verified
Attributed role
Research-data provider under review
Provider type
Specialist
Location
Not listed
Evidence
Limited
Visit company website
NRHuzzle LabsDescribes professional-task RL environments, but no named task set with public methods and measured outcomes was verified.Long HorizonQualifying artifact not verifiedSpecialistLondon, UKLimited

Describes professional-task RL environments, but no named task set with public methods and measured outcomes was verified.

The public offering describes environment construction and domain access. The reviewed sources do not yet expose a completed, attributable artifact or delivered evaluation with enough methods and outcomes to apply the scoring rubric.

Long Horizon
Evidence snapshot
Qualifying artifact not verified
Attributed role
Environment provider under review
Provider type
Specialist
Location
London, UK
Evidence
Limited
Visit company website
NRMatricesStates that it builds computer-use training environments for frontier labs, but publishes no named evaluation result in the reviewed record.Computer UseAttributable result not verifiedSpecialistNot listedLimited

States that it builds computer-use training environments for frontier labs, but publishes no named evaluation result in the reviewed record.

Matrices describes active frontier-lab partnerships and real-work computer-use environments. Without a named artifact, inspectable methods, or attributable delivered outcome, those claims establish relevance but not ranking eligibility.

Computer Use
Evidence snapshot
Attributable result not verified
Attributed role
Computer-use environment provider under review
Provider type
Specialist
Location
Not listed
Evidence
Limited
Visit company website
NRmicro1Offers expert-led agent evaluation and targeted training data, but the reviewed record lacks a named benchmark or attributable outcome.Long HorizonNamed delivery outcome not verifiedScaled providerNot listedLimited

Offers expert-led agent evaluation and targeted training data, but the reviewed record lacks a named benchmark or attributable outcome.

Cortex documents evaluation design, failure diagnosis, expert training data, and monitoring. No task-level result, named model deployment, or independent customer evidence was verified for scoring in this edition.

Long Horizon
Evidence snapshot
Named delivery outcome not verified
Attributed role
Expert evaluation provider under review
Provider type
Scaled provider
Location
Not listed
Evidence
Limited
Visit company website
NRPlatoOffers a documented SDK for simulated application environments, but no completed benchmark or training result was verified.Computer UseEnterpriseWorking SDK; benchmark outcome not publishedSpecialistNot listedLimited

Offers a documented SDK for simulated application environments, but no completed benchmark or training result was verified.

Plato's documentation shows runnable environment sessions, application simulators, task loading, and rollout evaluation. A working product alone does not establish the methods and outcomes needed for a comparative ranking.

Computer UseEnterprise
Evidence snapshot
Working SDK; benchmark outcome not published
Attributed role
Simulation platform under review
Provider type
Specialist
Location
Not listed
Evidence
Limited
Visit company website
NRPreference ModelDescribes RL environments for machine-learning research, but no completed named artifact or measured evaluation was verified.Completed artifact not verifiedSpecialistUndisclosedLimited

Describes RL environments for machine-learning research, but no completed named artifact or measured evaluation was verified.

The company states that it builds diverse tasks and reward functions with frontier labs. The reviewed public materials do not identify a completed environment, attributable customer delivery, or measured outcome that can be scored.

Evidence snapshot
Completed artifact not verified
Attributed role
ML-research environment provider under review
Provider type
Specialist
Location
Undisclosed
Evidence
Limited
Visit company website
NRVeris AIDocuments a simulation platform and customer benchmark workflow, but the reviewed public record does not yet support a scored provider-level assessment.EnterpriseUser SimulationTask-level public outcome record incompleteSpecialistNot listedLimited

Documents a simulation platform and customer benchmark workflow, but the reviewed public record does not yet support a scored provider-level assessment.

Veris describes stateful simulations, RL integration, and more than 40 customer benchmarks. This review did not verify a provider-authored public task set with inspectable task-level results or independently attributable delivery, so the company remains monitored.

EnterpriseUser Simulation
Evidence snapshot
Task-level public outcome record incomplete
Attributed role
Simulation and benchmark platform under review
Provider type
Specialist
Location
Not listed
Evidence
Limited
Visit company website

Methodology

RL Index scores work, not companies. Every rank comes from reading what a provider has actually put out (a benchmark and the harness that runs it, a paper, a released dataset, an evaluation somebody else executed) and asking the same five questions of it.

Scope and limits

The index sees only what gets published, and a lot of the strongest work in this field is done under NDA for one lab and never surfaces. Open, qualified and confidential access count the same when enough of the record can be verified. A quiet provider will still rank below what it has earned, which is a limit of the method rather than a judgment on the provider.

  1. 01

    Technical substance, depth & scope

    Finished work we can attribute, with methods and results. Depth means measured differences across mechanisms, conditions or skill stages, rather than a larger pile of tasks. Scope means work that spans data, environments, benchmarks, expert feedback and training instead of one narrow slice. An announced plan isn’t weighed against a shipped artifact.

  2. 02

    Technical realism

    Real tasks, real tools, real interfaces, and the awkwardness of the job as it’s actually done. Multistep dependencies and operating constraints count. A large task count on its own doesn’t.

  3. 03

    Verification rigor

    Whether the grading would catch a wrong answer. Controls that fail when they should, runs that repeat, and some stated handling of shortcuts, leakage, contamination and grader error.

  4. 04

    Training utility

    Whether the work moves a model or only measures one. Repeatable evaluation, assets a team could train on, integration into a training loop, gains reported on held-out tasks. A model score by itself doesn’t show this.

  5. 05

    External validation

    Somebody outside the company ran it and said what happened. A logo, a citation, or a second write-up of the same evaluation is one data point twice.

Evidence labels

Each entry carries one of three labels. They describe how much of a provider’s record we could get to, not how good the work is, and they don’t add to the score.

  • Extensive a broad record, with qualifying evidence across several pieces of work.
  • Documented qualifying technical evidence we could read.
  • Limited not much we could find in public material.

Limited comes up often and isn’t a criticism. Early companies, quiet companies, and companies whose main work sits behind an NDA all land there.

What stays blank

The figures a lab actually asks for in an RFI (task and sample counts, unique environment counts, pass rates and difficulty splits, data-type breakdowns, harness and format details, pricing) aren’t published by anyone in this market. We don’t estimate them, so those fields stay empty.

How to read a rank

A rank reflects how much a provider’s published work shows and how well it holds up. That tends to track product quality only loosely, so this is a reasonable place to start a shortlist.

A few practical notes. We score each provider’s strongest qualifying work for a given factor, which doesn’t mean all of it lives in one product. Equal assessments share a rank. A specialist can finish above a much larger provider, and often does. Breadth only helps when it represents genuinely separate outcomes rather than one piece of work described several ways.

Freshness

Evidence reviewed through 18 September 2026, and re-checked on a rolling basis.

AI model improvement FAQ

Technical answers about training and evaluating AI systems with interactive, verifiable tasks, data and feedback.

What is an RL environment?

An RL environment lets an AI system take actions on a task, observe the resulting state and receive rewards used during training. Feedback can come from programmatic checks, expert judgment or other grading methods. Reliable resets and outcome checks make an environment more useful.

How is an RL environment different from a dataset or benchmark?

A dataset supplies examples, while a benchmark measures performance. An RL environment supports repeated interaction through an action space, observable state and reward or feedback. A benchmark or dataset can support training, but neither establishes an interactive environment by itself.

What training data is needed to improve an AI agent?

Useful training assets include demonstrations, successful and failed tool-use trajectories, expert feedback, realistic tasks and checks of final outcomes. The right mixture depends on the target capability. Improvement should be measured on held-out tasks rather than assumed from dataset size or variety.

How do AI agents learn from environments?

Agents can be trained on demonstrations or rewarded attempts at tasks. Training updates the model or policy from those examples or rewards; running a workflow alone does not train it. Held-out tests measure whether improvements transfer to unfamiliar tasks and conditions.

What makes a model-improvement task reliably verifiable?

A verifiable task has an outcome that can be checked against the resulting state, artifact or expert rubric. Graders should be tested against known successes, failures and shortcuts. Passing a verifier is evidence within its coverage, not proof of general capability.

How can teams reduce reward hacking and evaluation contamination?

Teams can separate training and held-out tasks, inspect final state, and audit shortcuts that satisfy a grader without completing the objective. Depending on the task, safeguards include hidden tests, controlled access, refreshed cases, overlap checks and independent review.

When should an AI lab build versus buy training environments?

An AI lab should build when an environment encodes a proprietary capability or must be tightly coupled to internal infrastructure. Buying is useful when the lab needs faster task production, specialist expertise, broader coverage or independent held-out evaluation. Many programs combine an internal harness with externally supplied tasks, data and audits.

How should an AI lab evaluate a model-improvement provider?

An AI lab should evaluate technical realism, verifier quality, task diversity, reset reliability, difficulty calibration and resistance to shortcuts. It should also examine integration with the training stack, separation of training and held-out evaluation, security controls and delivery within a model-training cycle. A public artifact is useful evidence, but not sufficient by itself.

Why are established companies included?

This is an index of the market, not of startups. A provider qualifies on completed, attributable work with methods and results. Company age, size, funding and prestige are not ranking factors.

Can a specialist outrank a much larger provider?

Yes, and several do. The five factors weigh the quality of the published work rather than the size of the company, so a narrow provider with a strong record can finish above a broad one.