AI model improvement index · 2026
RL environments, training data & evaluations.
An index of companies building RL environments, training data, benchmarks and evaluations for frontier labs.
Training a frontier model now means buying tasks, environments and graders from the AI training data and RL environments market. This index tracks 42 companies selling into that market and ranks 30 of them on five factors: substance, realism, verification, training value and outside validation. The other twelve are listed without a rank, because we couldn’t find enough finished, attributable work to score them. All of this is read off published material: publishing little results in lower confidence. Each entry links out to what it was scored on.
Leaderboard · September 2026
AI model improvement providers
30 ranked providers · 12 monitored (NR). Expand an entry for its role, evidence and sources. Equal assessments share a rank.
01Scale AIBuilds expert-designed frontier evaluations and training data, with public methods, held-out controls, and broad model coverage.EnterpriseCyberHumanity's Last ExamScaled providerSan Francisco, USExtensive
Builds expert-designed frontier evaluations and training data, with public methods, held-out controls, and broad model coverage.
Humanity's Last Exam publishes benchmark construction, evaluation methods, results, and a held-out integrity design. The work is jointly attributed to Scale AI and the Center for AI Safety; Scale Labs also operates a broader set of continuously updated model evaluations. Scale's ASPI and PropensityBench cyber track are scored separately on RL Index Cyber.
02AfterQueryDevelops expert data and agent evaluations, including a task-linked vulnerability-analysis benchmark with public graded cases.CodeEnterpriseCyberVADER; public graded casesSpecialistSan Francisco, USDocumented
Develops expert data and agent evaluations, including a task-linked vulnerability-analysis benchmark with public graded cases.
VADER reports methods and model results across real code vulnerability cases. The repository exposes task metadata, graded outputs, and scoring code, allowing the reported evaluation to be inspected beyond a marketing claim.
03Surge AIDelivers human feedback, expert data, and agent environments used in frontier-model training and evaluation.EnterpriseLong HorizonAnthropic RLHF; workplace RLScaled providerSan Francisco, USExtensive
Delivers human feedback, expert data, and agent environments used in frontier-model training and evaluation.
Surge documents Anthropic's use of its RLHF platform and publishes measured evaluations from a 150-task workplace environment. Its research record also attributes joint work with Anthropic, OpenAI, Meta, and academic groups; each collaboration is counted once rather than as separate validation signals.
04MechanizeBuilds software-engineering environments and evaluations using executable tasks, deterministic emulation, and outcome-based grading.CodeGBA EvalSpecialistSan Francisco, USDocumented
Builds software-engineering environments and evaluations using executable tasks, deterministic emulation, and outcome-based grading.
GBA Eval publishes its scope, harness design, grading approach, and model results. It is a bounded coding evaluation rather than evidence for every software-engineering capability, but it provides an inspectable example of Mechanize's environment methodology.
05MercorCombines a large expert network with public professional-work benchmarks and commercial RL-environment delivery.CodeEnterpriseLong HorizonAPEX; APEX-Agents; APEX-SWEScaled providerSan Francisco, USDocumented
Combines a large expert network with public professional-work benchmarks and commercial RL-environment delivery.
Mercor publishes the APEX benchmark family with task descriptions, data or code links, and measured model results across professional, agentic, and software-engineering work. Customer-count claims are not treated as independent validation without attributable evidence.
06Bespoke LabsDevelops open datasets, executable agent tasks, and post-training tooling with reproducible public releases.CodeLong HorizonCyberOpenThoughts-TBLiteSpecialistMountain View, USExtensive
Develops open datasets, executable agent tasks, and post-training tooling with reproducible public releases.
OpenThoughts-TBLite publishes task definitions, tests, and aggregate results with joint attribution. Nous Research separately distributes the dataset, corroborating outside integration while preserving collaborators' credit. OpenThoughts-TBLite's security-task subset is scored separately on RL Index Cyber.
07Handshake AICombines an expert network with open verifier research and realistic professional-work evaluation environments.EnterpriseLong HorizonGandalf; ATLAS-FinanceScaled providerSan Francisco, USDocumented
Combines an expert network with open verifier research and realistic professional-work evaluation environments.
Handshake publishes Gandalf, an environment-aligned verifier, with results on 3,204 expert-graded judgments from BankerVerifierBench. ATLAS-Finance extends that work to 100 tasks across 13 workplace environments with expert-authored rubrics and public code.
08DatacurveDevelops long-horizon coding data and evaluations with original repository tasks and programmatic verification.CodeDeepSWESpecialistSan Francisco, USDocumented
Develops long-horizon coding data and evaluations with original repository tasks and programmatic verification.
DeepSWE publishes 113 original tasks across active repositories, isolated environments, and program-based verifiers. The repository exposes task manifests and grading artifacts, making the benchmark inspectable at task level.
09Collinear AIBuilds task-linked evaluations and runnable training environments, including secure code repair with functional and security checks.CodeCyberLong HorizonCWE-Bench; YC-BenchSpecialistSan Francisco, USExtensive
Builds task-linked evaluations and runnable training environments, including secure code repair with functional and security checks.
CWE-Bench publishes methods and task-linked results. Google reports a completed Gemini evaluation run by Collinear on the benchmark. YC-Bench evaluates long-horizon agent planning over a simulated one-year startup run, with published methods, results across 12 proprietary and open-source models at three seeds each, and an open-source implementation.
10TuringBuilds expert-authored coding, terminal, and STEM training tasks with executable verification and published benchmark results.CodeEnterpriseCyberSWE-Bench++; Terminal-Bench dataScaled providerPalo Alto, USDocumented
Builds expert-authored coding, terminal, and STEM training tasks with executable verification and published benchmark results.
SWE-Bench++ publishes more than 500 public and 7,000 additional repository tasks with executable tests and an evaluation and training framework. Turing also documents verifier-graded terminal and frontier STEM data, including successful trajectories used for fine-tuning. Turing's CyberStrike evaluation is scored separately on RL Index Cyber.
11Gray Swan AIBuilds adversarial evaluations and protective methods with substantial joint research and outside model-assessment use.CyberAgentHarm; ARTEMIS; IPISpecialistPittsburgh, USExtensive
Builds adversarial evaluations and protective methods with substantial joint research and outside model-assessment use.
Published agent evaluations, protective tests, and implemented attack methods establish technical depth. NIST and Anthropic materials corroborate outside use; joint research retains its collaborators' credit.
12Andon LabsPublishes long-horizon, computer-use, and physical-agent evaluations with inspectable setups and measured model results.Long HorizonUser SimulationVending-Bench 2; Drone-Bench; Blueprint-Bench 2SpecialistNot listedDocumented
Publishes long-horizon, computer-use, and physical-agent evaluations with inspectable setups and measured model results.
Andon Labs publishes detailed environments, scoring methods, and repeated model runs across Vending-Bench 2, Drone-Bench, and Blueprint-Bench 2. The benchmarks cover distinct outcomes and are not combined merely because they share a publisher.
13ProximalDevelops original ultra-long-horizon coding tasks with external contributors and measured frontier-model results.CodeLong HorizonFrontierSWESpecialistNot listedDocumented
Develops original ultra-long-horizon coding tasks with external contributors and measured frontier-model results.
FrontierSWE publishes original tasks, model results, long execution budgets, and contributions from academic and industry partners. The evidence supports difficult software-engineering evaluation, not every model-improvement domain.
14Prime IntellectTrains and open-sources large RL post-trained models and operates a public hub of community RL environments.CodeComputer UseEnterpriseINTELLECT-3; Environments HubOpen / hybridSan Francisco, USDocumented
Trains and open-sources large RL post-trained models and operates a public hub of community RL environments.
INTELLECT-3 is a 106B-parameter open-weight model post-trained with large-scale reinforcement learning, released with a technical report and published benchmark scores (MATH-500 98.1, AIME 2024 90.8, AIME 2025 88.0, GPQA 74.4). The Environments Hub separately hosts more than 2,500 community-contributed RL environments used for training and evaluation; environment count is treated as platform breadth, not independent validation. Prime Intellect's Security Verifiers and hosted Strix-XSS training are scored separately on RL Index Cyber.
15RefreshPublishes verifiable coding datasets and computer-use tooling with measured out-of-distribution training gains.CodeComputer UseGauntlet 4K; computer-1SpecialistNot listedDocumented
Publishes verifiable coding datasets and computer-use tooling with measured out-of-distribution training gains.
Gauntlet 4K reports a threefold Terminal-Bench 2.0 improvement after post-training on 4,000 pytest-verifiable environments. The larger Gauntlet release and open computer-1 harness provide inspectable tasks, verification, and execution tooling.
16Snorkel AIBuilds expert-curated training and evaluation data, including verifiable professional computer-use tasks.Computer UseEnterpriseCyberCUA-Bench+; KiCadScaled providerRedwood City, USExtensive
Builds expert-curated training and evaluation data, including verifiable professional computer-use tasks.
Snorkel publishes a 25-task KiCad benchmark built on Cua-Bench with objective saved-artifact checks and measured model results. CUA-Bench+ extends the work into a commercial data series with expert review, programmatic validation, and difficulty calibration. Snorkel's ShadowRelay and OpenThoughts-TBLite security tasks are scored separately on RL Index Cyber.
17Invisible TechnologiesDelivers scaled expert training and evaluation programs, with named customer work and outcome-linked case evidence.EnterpriseCohere eval; targeted trainingScaled providerDistributedExtensive
Delivers scaled expert training and evaluation programs, with named customer work and outcome-linked case evidence.
Invisible identifies Cohere as an evaluation customer and reports expert annotation across specialized domains and languages. A separate case record reports evaluation of 45 internal models and an 87% quality improvement after targeted training; full task-level methods are not public.
=18CuaProvides open computer-use environments, fleets, and a benchmark framework for verifiable desktop and mobile tasks.Computer UseCua-BenchInfrastructureNot listedExtensive
Provides open computer-use environments, fleets, and a benchmark framework for verifiable desktop and mobile tasks.
Cua-Bench publishes a task model, runner, evaluators, registry, and measured datasets across desktop and mobile surfaces. Snorkel's independent KiCad study uses the framework, adding outside execution beyond Cua's own demonstrations.
=18HUDBuilds expert-validated financial-analyst and robotics benchmarks with public datasets and measured model results.CodeComputer UseEnterpriseSheetBench-50; Assemble BenchSpecialistNot listedDocumented
Builds expert-validated financial-analyst and robotics benchmarks with public datasets and measured model results.
SheetBench-50 is a 50-task financial-analyst spreadsheet benchmark built with Sepal AI and validated by expert reviewers from PwC, Cisco, Charles Schwab, and Fannie Mae, with each task constrained to a single reproducible answer. Assemble Bench separately publishes 14 contact-rich robot-assembly tasks with an open dataset, repository, and measured dense-reward results on a DROID checkpoint. HUD's ZeroDayBench evaluation is scored separately on RL Index Cyber.
=18Vals AIRuns independent held-out benchmarks across finance, legal, and coding, publishing methods and measured model scores.CodeCyberVals Index; VLAIRSpecialistSan Francisco, USExtensive
Runs independent held-out benchmarks across finance, legal, and coding, publishing methods and measured model scores.
The Vals Index is a GDP-weighted composite of finance, coding, and legal benchmarks with a public leaderboard of measured model scores. VLAIR separately tests four legal-AI products against a practicing-lawyer baseline across seven tasks, with results published per vendor and per task. An independent arXiv paper builds directly on Vals' Finance Agent v2 and calls it a reference benchmark. Vals AI's CyberBench evaluation is scored separately on RL Index Cyber.
21QuesmaPublishes open agentic coding benchmarks with public leaderboards reporting per-model pass rates, cost, and speed.CodeLong HorizonCyberCompileBench; OTelBenchSpecialistWarsaw, PolandDocumented
Publishes open agentic coding benchmarks with public leaderboards reporting per-model pass rates, cost, and speed.
CompileBench scores 26 models on 15 real-world build and compilation tasks, tracking pass@1, pass@3, cost, and completion time on a public leaderboard. OTelBench separately publishes 23 observability-instrumentation tasks across 11 languages with executable grading and results for 14 models. Quesma's BinaryAudit backdoor-detection benchmark is scored separately on RL Index Cyber.
=22LabelboxCombines scaled training-data delivery with public agent evaluations and simulated enterprise-work environments.EnterpriseLong HorizonCyberImplicit Intelligence; WorldSimScaled providerSan Francisco, USDocumented
Combines scaled training-data delivery with public agent evaluations and simulated enterprise-work environments.
Implicit Intelligence publishes 205 agent scenarios, scoring methods, and model results across contextual constraints. Horizon documents WorldSim environments and reward design; general platform scale is not counted as additional technical validation. A privacy-and-security scenario subset of Implicit Intelligence is scored separately on RL Index Cyber.
=22TolokaDevelops private agent benchmarks, RL gyms, and expert-annotated data across business workflows and research domains.Long HorizonUser SimulationToloka Arena; Tau-bench extensionScaled providerNot listedDocumented
Develops private agent benchmarks, RL gyms, and expert-annotated data across business workflows and research domains.
Toloka Arena publishes model results from hidden, non-contaminated workflow tasks and offers the underlying RL gyms for training. Its Tau-bench extension documents databases, tools, policies, golden trajectories, and calibrated difficulty.
=22VmaxPublishes RL post-training research on automated task generation and self-play, with measured gains on math, code, and SWE benchmarks.CodeCyberPROPEL; PopuLoRA; unix-ctfSpecialistSan Francisco, USExtensive
Publishes RL post-training research on automated task generation and self-play, with measured gains on math, code, and SWE benchmarks.
PROPEL trains task generators at a targeted solve rate, lifting learnable-frontier coding tasks from 10.1% to 20.0% (Qwen2.5-3B) and software-engineering tasks on unseen repositories from 9.8% to 19.6% (Qwen3.5-27B). PopuLoRA evolves populations of LoRA adapters under programmatic verification and beats single-agent baselines across ten math and coding benchmarks; an independent open-source reimplementation corroborates the method. unix-ctf separately reports procedural command-line environment generation with held-out, baseline-controlled training gains and is scored on RL Index Cyber.
25Good Start LabsBuilds game-based benchmarks and training environments with a named publisher partner and measured deployed outcomes.Long HorizonArkadium Game Lab; Gin RummyOpen / hybridNot listedExtensive
Builds game-based benchmarks and training environments with a named publisher partner and measured deployed outcomes.
Good Start Labs attributes benchmark and training work to Arkadium, including Game Lab and a compact expert model that beat casual human players in reported trials. The public record includes the customer, objective, and measured outcome, though not the complete training dataset.
26Chakra LabsBuilds deterministic computer-use environments with frame-accurate control and publishes model results on a demonstration benchmark.Computer UseDojo; dojo-bench-miniSpecialistNot listedDocumented
Builds deterministic computer-use environments with frame-accurate control and publishes model results on a demonstration benchmark.
Chakra documents the Dojo environment suite, deterministic state controls, and model performance on dojo-bench-mini. The public evidence is provider-authored and the full task set was not independently executed in this review.
27BenchFlowPublishes stateful workplace environments, agent benchmarks, datasets, and an environment framework for reproducible evaluation.CodeEnterpriseSkillsBench; ClawsBenchSpecialistNot listedDocumented
Publishes stateful workplace environments, agent benchmarks, datasets, and an environment framework for reproducible evaluation.
BenchFlow publishes SkillsBench, ClawsBench, and PostTrain with papers, datasets, repositories, and live environments. ClawsBench supplies five wire-compatible workplace services and task manifests rather than a static prompt set.
28General ReasoningMaintains open RL-environment infrastructure and publishes a long-horizon decision benchmark with measured model results.Long HorizonOpenReward; KellyBenchOpen / hybridNot listedDocumented
Maintains open RL-environment infrastructure and publishes a long-horizon decision benchmark with measured model results.
OpenReward serves more than 330 community environments through a common API, while KellyBench publishes a sequential football-market environment and evaluation results. Community environment counts are treated as platform breadth, not as General Reasoning-authored benchmarks.
29Patronus AIDevelops evaluation datasets and judge models with public methods and results across financial reasoning and hallucination detection.Long HorizonFinanceBench; HaluBench; LynxSpecialistNot listedDocumented
Develops evaluation datasets and judge models with public methods and results across financial reasoning and hallucination detection.
FinanceBench publishes a 10,000-question financial evaluation and measured retrieval-system results. HaluBench and Lynx add public data, code, and comparative judge-model results; benchmark publication does not by itself establish training delivery.
30ArenaOperates large-scale human-preference evaluations and publishes open datasets, ranking methods, and commercial evaluation services.Chatbot Arena; Arena-HardData providerNot listedExtensive
Operates large-scale human-preference evaluations and publishes open datasets, ranking methods, and commercial evaluation services.
Arena publishes its leaderboard policy, ranking pipeline, open datasets, and academic methods. Its commercial evaluation product uses community feedback for model labs and developers; the public record demonstrates evaluation delivery rather than RL-environment production.
NRAIChampMarket relevance is plausible, but the reviewed public record does not establish a completed model-improvement artifact or delivery.EnterpriseLong HorizonQualifying evidence not verifiedSpecialistNot listedLimited
Market relevance is plausible, but the reviewed public record does not establish a completed model-improvement artifact or delivery.
No named benchmark, environment, dataset, or attributable customer result with sufficient public methods and outcomes was verified for this edition.
NRAndromedeDevelops programmatically generated RL environments, tasks, and verifiers, but remains in private beta without a public completed result.Long HorizonPrivate beta; completed result not verifiedSpecialistNot listedLimited
Develops programmatically generated RL environments, tasks, and verifiers, but remains in private beta without a public completed result.
The public site describes a post-training and evaluation pipeline and a small partner beta. No named task set, measured evaluation, or independently attributable delivery was available for scoring in this review.
NRFleet AIPublishes post-training infrastructure and describes high-fidelity environment delivery, but lacks task-linked public outcome evidence in this review.CodeEnterpriseQualifying outcome evidence not verifiedSpecialistNew York, USLimited
Publishes post-training infrastructure and describes high-fidelity environment delivery, but lacks task-linked public outcome evidence in this review.
Fleet's public site and Miles repository establish an active environment and post-training platform. This review did not verify a provider-attributable task set with published methods and measured outcomes, so the company remains monitored rather than scored.
NRHabitat IncIs described as a work-automation environment vendor, but no inspectable company-authored technical artifact was verified.Computer UseEnterpriseQualifying artifact not verifiedSpecialistNot listedLimited
Is described as a work-automation environment vendor, but no inspectable company-authored technical artifact was verified.
The reviewed record supports category fit and an active company. It does not expose a named environment, public methods, or measured results attributable to Habitat Inc; third-party directory inclusion does not qualify as evidence.
NRHalluminateAppears in the model-improvement market, but a completed named environment, benchmark, or attributable delivery was not verified.EnterpriseQualifying artifact not verifiedSpecialistNot listedLimited
Appears in the model-improvement market, but a completed named environment, benchmark, or attributable delivery was not verified.
The current public record establishes category relevance but does not expose task-level methods, measurable outcomes, or a named customer delivery. Offering language and list inclusion do not satisfy the ranking gate.
NRHillclimbTargets AI-research training data and RL-environment creation, but has not published a current company-attributable artifact with outcomes.Current attributable artifact not verifiedSpecialistNot listedLimited
Targets AI-research training data and RL-environment creation, but has not published a current company-attributable artifact with outcomes.
Hillclimb states that it sells training data to frontier labs and is scaling RL-environment creation. Prior-team work and investor backing do not substitute for a named current artifact or delivered result.
NRHuzzle LabsDescribes professional-task RL environments, but no named task set with public methods and measured outcomes was verified.Long HorizonQualifying artifact not verifiedSpecialistLondon, UKLimited
Describes professional-task RL environments, but no named task set with public methods and measured outcomes was verified.
The public offering describes environment construction and domain access. The reviewed sources do not yet expose a completed, attributable artifact or delivered evaluation with enough methods and outcomes to apply the scoring rubric.
NRMatricesStates that it builds computer-use training environments for frontier labs, but publishes no named evaluation result in the reviewed record.Computer UseAttributable result not verifiedSpecialistNot listedLimited
States that it builds computer-use training environments for frontier labs, but publishes no named evaluation result in the reviewed record.
Matrices describes active frontier-lab partnerships and real-work computer-use environments. Without a named artifact, inspectable methods, or attributable delivered outcome, those claims establish relevance but not ranking eligibility.
NRmicro1Offers expert-led agent evaluation and targeted training data, but the reviewed record lacks a named benchmark or attributable outcome.Long HorizonNamed delivery outcome not verifiedScaled providerNot listedLimited
Offers expert-led agent evaluation and targeted training data, but the reviewed record lacks a named benchmark or attributable outcome.
Cortex documents evaluation design, failure diagnosis, expert training data, and monitoring. No task-level result, named model deployment, or independent customer evidence was verified for scoring in this edition.
NRPlatoOffers a documented SDK for simulated application environments, but no completed benchmark or training result was verified.Computer UseEnterpriseWorking SDK; benchmark outcome not publishedSpecialistNot listedLimited
Offers a documented SDK for simulated application environments, but no completed benchmark or training result was verified.
Plato's documentation shows runnable environment sessions, application simulators, task loading, and rollout evaluation. A working product alone does not establish the methods and outcomes needed for a comparative ranking.
NRPreference ModelDescribes RL environments for machine-learning research, but no completed named artifact or measured evaluation was verified.Completed artifact not verifiedSpecialistUndisclosedLimited
Describes RL environments for machine-learning research, but no completed named artifact or measured evaluation was verified.
The company states that it builds diverse tasks and reward functions with frontier labs. The reviewed public materials do not identify a completed environment, attributable customer delivery, or measured outcome that can be scored.
NRVeris AIDocuments a simulation platform and customer benchmark workflow, but the reviewed public record does not yet support a scored provider-level assessment.EnterpriseUser SimulationTask-level public outcome record incompleteSpecialistNot listedLimited
Documents a simulation platform and customer benchmark workflow, but the reviewed public record does not yet support a scored provider-level assessment.
Veris describes stateful simulations, RL integration, and more than 40 customer benchmarks. This review did not verify a provider-authored public task set with inspectable task-level results or independently attributable delivery, so the company remains monitored.
Methodology
RL Index scores work, not companies. Every rank comes from reading what a provider has actually put out (a benchmark and the harness that runs it, a paper, a released dataset, an evaluation somebody else executed) and asking the same five questions of it.
The index sees only what gets published, and a lot of the strongest work in this field is done under NDA for one lab and never surfaces. Open, qualified and confidential access count the same when enough of the record can be verified. A quiet provider will still rank below what it has earned, which is a limit of the method rather than a judgment on the provider.
- 01
Technical substance, depth & scope
Finished work we can attribute, with methods and results. Depth means measured differences across mechanisms, conditions or skill stages, rather than a larger pile of tasks. Scope means work that spans data, environments, benchmarks, expert feedback and training instead of one narrow slice. An announced plan isn’t weighed against a shipped artifact.
- 02
Technical realism
Real tasks, real tools, real interfaces, and the awkwardness of the job as it’s actually done. Multistep dependencies and operating constraints count. A large task count on its own doesn’t.
- 03
Verification rigor
Whether the grading would catch a wrong answer. Controls that fail when they should, runs that repeat, and some stated handling of shortcuts, leakage, contamination and grader error.
- 04
Training utility
Whether the work moves a model or only measures one. Repeatable evaluation, assets a team could train on, integration into a training loop, gains reported on held-out tasks. A model score by itself doesn’t show this.
- 05
External validation
Somebody outside the company ran it and said what happened. A logo, a citation, or a second write-up of the same evaluation is one data point twice.
Evidence labels
Each entry carries one of three labels. They describe how much of a provider’s record we could get to, not how good the work is, and they don’t add to the score.
- Extensive a broad record, with qualifying evidence across several pieces of work.
- Documented qualifying technical evidence we could read.
- Limited not much we could find in public material.
Limited comes up often and isn’t a criticism. Early companies, quiet companies, and companies whose main work sits behind an NDA all land there.
What stays blank
The figures a lab actually asks for in an RFI (task and sample counts, unique environment counts, pass rates and difficulty splits, data-type breakdowns, harness and format details, pricing) aren’t published by anyone in this market. We don’t estimate them, so those fields stay empty.
How to read a rank
A rank reflects how much a provider’s published work shows and how well it holds up. That tends to track product quality only loosely, so this is a reasonable place to start a shortlist.
A few practical notes. We score each provider’s strongest qualifying work for a given factor, which doesn’t mean all of it lives in one product. Equal assessments share a rank. A specialist can finish above a much larger provider, and often does. Breadth only helps when it represents genuinely separate outcomes rather than one piece of work described several ways.
Freshness
Evidence reviewed through 18 September 2026, and re-checked on a rolling basis.
AI model improvement FAQ
Technical answers about training and evaluating AI systems with interactive, verifiable tasks, data and feedback.
What is an RL environment?+
An RL environment lets an AI system take actions on a task, observe the resulting state and receive rewards used during training. Feedback can come from programmatic checks, expert judgment or other grading methods. Reliable resets and outcome checks make an environment more useful.
How is an RL environment different from a dataset or benchmark?+
A dataset supplies examples, while a benchmark measures performance. An RL environment supports repeated interaction through an action space, observable state and reward or feedback. A benchmark or dataset can support training, but neither establishes an interactive environment by itself.
What training data is needed to improve an AI agent?+
Useful training assets include demonstrations, successful and failed tool-use trajectories, expert feedback, realistic tasks and checks of final outcomes. The right mixture depends on the target capability. Improvement should be measured on held-out tasks rather than assumed from dataset size or variety.
How do AI agents learn from environments?+
Agents can be trained on demonstrations or rewarded attempts at tasks. Training updates the model or policy from those examples or rewards; running a workflow alone does not train it. Held-out tests measure whether improvements transfer to unfamiliar tasks and conditions.
What makes a model-improvement task reliably verifiable?+
A verifiable task has an outcome that can be checked against the resulting state, artifact or expert rubric. Graders should be tested against known successes, failures and shortcuts. Passing a verifier is evidence within its coverage, not proof of general capability.
How can teams reduce reward hacking and evaluation contamination?+
Teams can separate training and held-out tasks, inspect final state, and audit shortcuts that satisfy a grader without completing the objective. Depending on the task, safeguards include hidden tests, controlled access, refreshed cases, overlap checks and independent review.
When should an AI lab build versus buy training environments?+
An AI lab should build when an environment encodes a proprietary capability or must be tightly coupled to internal infrastructure. Buying is useful when the lab needs faster task production, specialist expertise, broader coverage or independent held-out evaluation. Many programs combine an internal harness with externally supplied tasks, data and audits.
How should an AI lab evaluate a model-improvement provider?+
An AI lab should evaluate technical realism, verifier quality, task diversity, reset reliability, difficulty calibration and resistance to shortcuts. It should also examine integration with the training stack, separation of training and held-out evaluation, security controls and delivery within a model-training cycle. A public artifact is useful evidence, but not sufficient by itself.
Why are established companies included?+
This is an index of the market, not of startups. A provider qualifies on completed, attributable work with methods and results. Company age, size, funding and prestige are not ranking factors.
Can a specialist outrank a much larger provider?+
Yes, and several do. The five factors weigh the quality of the published work rather than the size of the company, so a narrow provider with a strong record can finish above a broad one.
No providers match that search or filter.