RL Environments Startups & Companies

An index of 92 startups building RL environments, agent evals and benchmarks, RLHF and human data pipelines, and sandboxed training infrastructure for AI labs. Corrections and additions welcome via the get listed page.

92 startups

NameDomainsLocationTeamFundingRaising
AfterQuery
Expert human data and RL environments across code, finance, and computer use
Multi-Domain, Coding, FinanceSan Francisco, New York, Seattle51-100$30.5M total (reported)-
AIChamp
Custom enterprise-workflow RL environments with expert grading
Enterprise, Long Horizon, Custom EnvironmentsSan Francisco1-10--
Akhara
Enterprise and code RL environments
Enterprise, CodingSan Francisco1-10--
Anchor Browser
Reliable browser automation platform for agentic AI
Browser, Agents InfrastructureTel Aviv1-10Seed, $6M (Blumberg Capital, Gradient Ventures, Nov 2025)-
Andon Labs
Long-horizon autonomy benchmarks like Vending-Bench
Long Horizon, AlignmentSan Francisco1-10Seed (Y Combinator)-
Andromede
Programmatic generation of long-horizon RL environments
Long HorizonLausanne1-10--
Anthromind
Medical and long-horizon environments and expert data
Medical, Long Horizon, Data LabelingSan Francisco1-10--
Applied Compute
Ex-OpenAI trio applying RL to build specialist enterprise models
Enterprise, Machine Learning, Custom EnvironmentsSan Francisco11-25$80M total at $700M valuation (Oct 2025)Yes
ARIMLABS
Security and long-horizon environments for agentic AI
Cybersecurity, Long HorizonWarsaw11-25--
Artificial Analysis
Independent benchmarking of AI models across intelligence, speed, and price
Multi-Domain, Machine LearningSan Francisco11-25$2.6M (2024)-
Axiom Math
Self-improving AI mathematician with formally verified reasoning
MathSan Francisco11-25Series A, $200M at $1.6B valuation (2026); $264M total-
BenchFlow
Open-source benchmark hub and eval infrastructure for agents
Enterprise, Browser, CodingSan Francisco1-10~$1M (reported)-
Besimple AI
Human-in-the-loop annotation and evals, with a voice data focus
Data Labeling, Voice, RLHFSan Francisco1-10~$3.5M total-
Bespoke Labs
Data curation and RL environment recipes from ex-Google DeepMind researchers
Coding, Machine LearningMountain View, Menlo Park, Bangalore, San Francisco11-25~$40M (reported)-
Browserbase
Headless browser infrastructure powering web agents
Browser, Agents InfrastructureSan Francisco26-50Series B, $40M at $300M valuation (June 2025); $67.5M total-
Chakra Labs
Dojo: a hub of computer-use and tool-use environments
Computer Use, Tool UseBrooklyn11-25~$10.1M (reported)-
Collinear
Enterprise simulation, judges, and long-horizon trajectory generation
Enterprise, Long Horizon, Machine Learning, SimulationMountain View, Sunnyvale11-25--
Cua
Open-source computer-use agent infrastructure and environments
Coding, Computer Use, Agents InfrastructureSan Francisco1-10Pre-seed/Seed (YC X25)-
Datacurve
Frontier coding data and repository RL environments via the Shipd bounty platform
Coding, RLHFSan Francisco26-50Series A, $15M led by Chemistry (Oct 2025); $17.7M total-
Deeptune
Code and computer-use environments; acquired by Mercor
Coding, Computer UseNew York26-50Series A, $43M; acquired by Mercor (2026)No
Diffuse Labs
ML and long-horizon RL environments
Machine Learning, Long HorizonPalo Alto, San Francisco1-10--
Dissei
Finance-domain RL environments
FinanceLondon1-10--
dmodel
Alignment-oriented environments focused on reward quality
Machine Learning, AlignmentSan Francisco11-25--
Duality AI
Falcon: reality-grade digital twin simulation for AI and robotics
Robotics, SimulationSan Mateo26-50--
E2B
Open-source cloud sandboxes for AI agents
Agents Infrastructure, CodingSan Francisco26-50Series A, $21M (Insight Partners, July 2025); $32M total-
EdotEnv
Long-horizon planning environments for frontier models
Long Horizon, Machine LearningSan Francisco1-10--
Emulated
Code and ML RL environments
Coding, Machine LearningSan Francisco1-10--
Epoch AI
Nonprofit research institute behind FrontierMath and AI capability benchmarks
Math, Machine LearningRemote11-25Philanthropic grants (nonprofit)-
Exabite
Code RL environments with realistic software execution
CodingRemote1-10--
Fleet
RL gyms replicating real enterprise software like Salesforce and Excel
Multi-Domain, Enterprise, Computer Use, SimulationSan Francisco, New York26-50Seed, $15M (Sequoia, Menlo Ventures, SV Angel)Yes
General Reasoning
Open reasoning data and reward models from the ex-Meta AI reasoning lead
Finance, Long Horizon, Machine LearningLondon, San Francisco1-10~$10.9M (reported)-
Genesis AI
Physics simulation engine and foundation model for robotics
Robotics, Simulation, Machine LearningSan Francisco, Paris26-50Seed, $105M (Eclipse Ventures, Khosla Ventures, July 2025)-
Good Start Labs
Game-based RL environments and benchmarks
Games, Long HorizonBrooklyn, New York, Toronto1-10~$3.6M (reported)-
Gray Swan AI
Adversarial red-teaming arenas and safety evals for frontier models
Cybersecurity, AlignmentPittsburgh11-25~$40M (reported)-
Habitat
Code and desktop interaction RL environments
Coding, Computer UseNew York1-10--
Halluminate
Sandboxed RL environments for finance and enterprise workflows
Finance, Enterprise, BrowserSan Francisco26-50--
Handshake
Career network turned human-data and RL environments provider via Handshake AI
Multi-Domain, Data Labeling, RLHFSan Francisco, New York, Bangalore, Berlin250+Series F, $200M (2022, ~$3.5B valuation)-
Harmonic
Mathematical superintelligence via formally verified reasoning
MathPalo Alto26-50Series B, $100M at ~$900M valuation (Kleiner Perkins, 2025)-
Hillclimb
Math environments emphasizing verifiable correctness
MathSan Francisco1-10--
HUD
Evals and RL environments platform for computer-use agents
Computer Use, Coding, Long Horizon, Agents Infrastructure, Custom EnvironmentsSan Francisco, Singapore11-25$15M raised (YC W25, Exceptional Capital)-
Huzzle Labs
Long-horizon code, tool-use, and enterprise workflow environments
Long Horizon, Coding, EnterpriseLondon, Berlin, San Francisco26-50~$6M (reported)-
Idler
Code environments with realistic execution constraints
CodingSan Francisco11-25--
Incalmo
Offensive-security environments for AI agents
CybersecuritySan Mateo11-25--
Invariant Labs
Security testing and analysis for AI agents; acquired by Snyk
Cybersecurity, AlignmentZurich1-10Acquired by Snyk (June 2025)No
Kaizen
Browser agents that simulate and automate legacy web portal work
Browser, EnterpriseSan Francisco1-10Seed, $500K (Y Combinator, Pioneer Fund)-
Kernel
Browser infrastructure for AI agents
Browser, Agents InfrastructureSan Francisco1-10Seed + Series A, $22M (Accel, 2025)-
Latch
Biology data infrastructure turned life-science environments and benchmarks
ScienceSan Francisco11-25Series A, $15M (2022)-
LMArena
Crowdsourced model leaderboards from the Chatbot Arena team
Multi-Domain, Machine LearningSan Francisco, Berkeley26-50$100M seed (a16z, UC Investments, 2025); $150M at $1.7B valuation (Jan 2026)-
Math, Inc.
Autoformalization agents for verified mathematics
MathPalo Alto1-10Seed, $15M (Torch Capital, Robot Ventures, Chapter One)-
Matrices
Browser-native training environments for web agents
Browser, Computer UseSan Francisco1-10~$5M (reported)-
Mechanize
RL environments to automate software engineering, founded by ex-Epoch AI researchers
CodingSan Francisco51-100~$9.1M (reported)-
Mercor
Expert marketplace powering evals and RL environments for frontier labs
Multi-Domain, Data Labeling, RLHFSan Francisco250+Series C, $350M at $10B valuation (Oct 2025)Yes
Metaphi
Code and enterprise RL environments
Coding, EnterpriseSan Francisco, New York1-10--
Micro1
Vetted domain experts for AI training data and evals
Data Labeling, Multi-Domain, RLHFLos Angeles101-250Series A, $35M at $500M valuation (Sept 2025)-
Morph
Infinibranch VMs: snapshot and branch entire environments for agents
Agents InfrastructureSan Francisco11-25~$19M total (seed led by Khosla Ventures)-
Normal
Hardware engineering environments for AI models
Hardware EngineeringSan Francisco1-10--
Nous Research
Open AI lab behind the Atropos RL environments framework
Machine Learning, Agents InfrastructureNew York11-25Series A, $50M led by Paradigm (~$1B valuation, 2025)-
Originator
Long-horizon computer-use environments
Computer Use, Coding, Long HorizonLondon, Paris1-10--
Osmosis
Forward-deployed reinforcement learning for AI agents
Machine LearningSan Francisco1-10Seed, $7M (CRV, Audacious Ventures, YC)-
Pareto
Expert data workforce for RLHF, evals, and tool-use environments
Multi-Domain, Tool Use, Data Labeling, RLHFSan Francisco26-50--
Phinity
Chip-design RL environments
Chip DesignSan Francisco, Palo Alto1-10--
Plato
High-fidelity replicas of websites and software for agent training
Browser, Enterprise, SimulationSan Francisco1-10--
pre.dev
Software-planning platform offering coding and long-horizon RL environments
Coding, Long HorizonDelaware1-10--
Preference Model
Stealth startup working on preference and reward modeling
Machine Learning, CodingSan Francisco, Toronto, Seattle11-25--
Prime Intellect
Open superintelligence stack: compute, RL environments hub, and sandboxes
Machine Learning, Agents Infrastructure, Multi-DomainSan Francisco26-50Series A, $130M at $1B valuation (2026); $150M+ total-
Proximal
Long-horizon coding RL environments built from real codebases
Coding, Long HorizonSan Francisco, Bangalore26-50--
Quesma
Security-domain RL environments and binary analysis evals
Cybersecurity, CodingWarsaw11-25--
ReasonCore
Science and code reasoning environments and benchmarks
Science, CodingSan Francisco11-25--
Refresh
Simulation engines with verifiable rewards for coding and computer use
Coding, Computer Use, SimulationSan Francisco1-10--
Rise Data Labs
US-based expert human data and custom RL environments for enterprise AI
Computer Use, Coding, Finance, Cybersecurity, Legal, Data Labeling, Enterprise, Multi-Domain, RLHF, Custom EnvironmentsUnited States--Yes
Runloop
Devboxes for training and evaluating coding agents
Coding, Agents InfrastructureSan Francisco1-10Seed, ~$7M-
Scale
Data-labeling incumbent extending into agent evals and RL environments
Multi-Domain, Coding, Data Labeling, RLHFSan Francisco, New York, Washington DC, London250+$1.6B+ raised; Meta invested $14.3B at ~$29B valuation (June 2025)-
Scaled Foundations
GRID: a simulation-first platform for robot learning
Robotics, Simulation, Machine LearningSeattle1-10--
Scrapybara
Virtual desktops for computer-use agents
Computer Use, Agents InfrastructureSan Francisco1-10Seed, $500K (CRV)-
Sepal AI
Science-domain environments; acquired by Mercor
ScienceSan Francisco1-10Acquired by Mercor (Feb 2026)No
Silverstream AI
Infrastructure and training data for reliable autonomous web agents
Browser, Machine LearningSan Francisco1-10Pre-seed, $1.2M (Gradient Ventures)-
Snorkel
Programmatic data platform expanding into expert evals and RL environments
Multi-Domain, Coding, Machine Learning, Data Labeling, RLHFSan Francisco, Redwood City, New York101-250Series D, $100M at $1.3B valuation (2025)-
Steel
Open-source browser API for AI agents
Browser, Agents Infrastructure-1-10~$17M (reported)-
Surge
Bootstrapped human-data leader with a dedicated RL environments org
Multi-Domain, Data Labeling, RLHFSan Francisco, New York, Seattle101-250Bootstrapped; reported in talks to raise ~$1B at $25B+ valuation (2025)Yes
SynthLabs
Post-training research: synthetic data and scalable RL alignment
Machine Learning, Alignment, RLHFSan Francisco1-10Seed (M12 and First Spark Ventures, 2024)-
Tacit Labs
Life-science and long-horizon environments for AI models
Science, Long HorizonSan Francisco1-10--
Taste Labs
Design-domain environments and evals for AI models
DesignNew York26-50--
The LLM Data Company
Evals and reward data tooling for LLM training
Machine Learning, Data Labeling, RLHF-1-10--
Theta
Enterprise RL environments
EnterpriseSan Francisco1-10--
Trajectory Labs
Alignment-focused environments for safe agent trajectories
AlignmentBerkeley, Toronto1-10--
Turing
AGI infrastructure: coding data and RL environments at scale
Coding, Multi-Domain, Data Labeling, RLHFSan Francisco, Palo Alto, Gurugram250+Series E, $111M at $2.2B valuation (2025)-
Ulam
Math RL environments and RLVR trajectories
MathWarsaw, London1-10--
Vals AI
Independent domain-specific benchmarks for legal, finance, and tax AI
Legal, Finance, MedicalSan Francisco1-10--
Veris AI
High-fidelity simulated environments to train enterprise AI agents
Enterprise, Custom Environments, SimulationNew York1-10Seed, $8.5M (Decibel and Acrew, June 2025)-
Verita AI
Design and UX environments with human preference feedback
DesignSan Francisco, Vancouver26-50--
Vetto AI
Code and computer-use environments from ex-DeepMind/Instagram founders
Coding, Computer UseSan Francisco, São Paulo, London1-10--
Vmax
Converts proprietary data into RL environments
Machine Learning, Custom EnvironmentsSan Francisco, New York1-10--