RL Environments Startups & Companies
An index of 92 startups building RL environments, agent evals and benchmarks, RLHF and human data pipelines, and sandboxed training infrastructure for AI labs. Corrections and additions welcome via the get listed page.
92 startups
| Name | Domains | Location | Team | Funding | Raising |
|---|---|---|---|---|---|
| AfterQuery Expert human data and RL environments across code, finance, and computer use | Multi-Domain, Coding, Finance | San Francisco, New York, Seattle | 51-100 | $30.5M total (reported) | - |
| AIChamp Custom enterprise-workflow RL environments with expert grading | Enterprise, Long Horizon, Custom Environments | San Francisco | 1-10 | - | - |
| Akhara Enterprise and code RL environments | Enterprise, Coding | San Francisco | 1-10 | - | - |
| Anchor Browser Reliable browser automation platform for agentic AI | Browser, Agents Infrastructure | Tel Aviv | 1-10 | Seed, $6M (Blumberg Capital, Gradient Ventures, Nov 2025) | - |
| Andon Labs Long-horizon autonomy benchmarks like Vending-Bench | Long Horizon, Alignment | San Francisco | 1-10 | Seed (Y Combinator) | - |
| Andromede Programmatic generation of long-horizon RL environments | Long Horizon | Lausanne | 1-10 | - | - |
| Anthromind Medical and long-horizon environments and expert data | Medical, Long Horizon, Data Labeling | San Francisco | 1-10 | - | - |
| Applied Compute Ex-OpenAI trio applying RL to build specialist enterprise models | Enterprise, Machine Learning, Custom Environments | San Francisco | 11-25 | $80M total at $700M valuation (Oct 2025) | Yes |
| ARIMLABS Security and long-horizon environments for agentic AI | Cybersecurity, Long Horizon | Warsaw | 11-25 | - | - |
| Artificial Analysis Independent benchmarking of AI models across intelligence, speed, and price | Multi-Domain, Machine Learning | San Francisco | 11-25 | $2.6M (2024) | - |
| Axiom Math Self-improving AI mathematician with formally verified reasoning | Math | San Francisco | 11-25 | Series A, $200M at $1.6B valuation (2026); $264M total | - |
| BenchFlow Open-source benchmark hub and eval infrastructure for agents | Enterprise, Browser, Coding | San Francisco | 1-10 | ~$1M (reported) | - |
| Besimple AI Human-in-the-loop annotation and evals, with a voice data focus | Data Labeling, Voice, RLHF | San Francisco | 1-10 | ~$3.5M total | - |
| Bespoke Labs Data curation and RL environment recipes from ex-Google DeepMind researchers | Coding, Machine Learning | Mountain View, Menlo Park, Bangalore, San Francisco | 11-25 | ~$40M (reported) | - |
| Browserbase Headless browser infrastructure powering web agents | Browser, Agents Infrastructure | San Francisco | 26-50 | Series B, $40M at $300M valuation (June 2025); $67.5M total | - |
| Chakra Labs Dojo: a hub of computer-use and tool-use environments | Computer Use, Tool Use | Brooklyn | 11-25 | ~$10.1M (reported) | - |
| Collinear Enterprise simulation, judges, and long-horizon trajectory generation | Enterprise, Long Horizon, Machine Learning, Simulation | Mountain View, Sunnyvale | 11-25 | - | - |
| Cua Open-source computer-use agent infrastructure and environments | Coding, Computer Use, Agents Infrastructure | San Francisco | 1-10 | Pre-seed/Seed (YC X25) | - |
| Datacurve Frontier coding data and repository RL environments via the Shipd bounty platform | Coding, RLHF | San Francisco | 26-50 | Series A, $15M led by Chemistry (Oct 2025); $17.7M total | - |
| Deeptune Code and computer-use environments; acquired by Mercor | Coding, Computer Use | New York | 26-50 | Series A, $43M; acquired by Mercor (2026) | No |
| Diffuse Labs ML and long-horizon RL environments | Machine Learning, Long Horizon | Palo Alto, San Francisco | 1-10 | - | - |
| Dissei Finance-domain RL environments | Finance | London | 1-10 | - | - |
| dmodel Alignment-oriented environments focused on reward quality | Machine Learning, Alignment | San Francisco | 11-25 | - | - |
| Duality AI Falcon: reality-grade digital twin simulation for AI and robotics | Robotics, Simulation | San Mateo | 26-50 | - | - |
| E2B Open-source cloud sandboxes for AI agents | Agents Infrastructure, Coding | San Francisco | 26-50 | Series A, $21M (Insight Partners, July 2025); $32M total | - |
| EdotEnv Long-horizon planning environments for frontier models | Long Horizon, Machine Learning | San Francisco | 1-10 | - | - |
| Emulated Code and ML RL environments | Coding, Machine Learning | San Francisco | 1-10 | - | - |
| Epoch AI Nonprofit research institute behind FrontierMath and AI capability benchmarks | Math, Machine Learning | Remote | 11-25 | Philanthropic grants (nonprofit) | - |
| Exabite Code RL environments with realistic software execution | Coding | Remote | 1-10 | - | - |
| Fleet RL gyms replicating real enterprise software like Salesforce and Excel | Multi-Domain, Enterprise, Computer Use, Simulation | San Francisco, New York | 26-50 | Seed, $15M (Sequoia, Menlo Ventures, SV Angel) | Yes |
| General Reasoning Open reasoning data and reward models from the ex-Meta AI reasoning lead | Finance, Long Horizon, Machine Learning | London, San Francisco | 1-10 | ~$10.9M (reported) | - |
| Genesis AI Physics simulation engine and foundation model for robotics | Robotics, Simulation, Machine Learning | San Francisco, Paris | 26-50 | Seed, $105M (Eclipse Ventures, Khosla Ventures, July 2025) | - |
| Good Start Labs Game-based RL environments and benchmarks | Games, Long Horizon | Brooklyn, New York, Toronto | 1-10 | ~$3.6M (reported) | - |
| Gray Swan AI Adversarial red-teaming arenas and safety evals for frontier models | Cybersecurity, Alignment | Pittsburgh | 11-25 | ~$40M (reported) | - |
| Habitat Code and desktop interaction RL environments | Coding, Computer Use | New York | 1-10 | - | - |
| Halluminate Sandboxed RL environments for finance and enterprise workflows | Finance, Enterprise, Browser | San Francisco | 26-50 | - | - |
| Handshake Career network turned human-data and RL environments provider via Handshake AI | Multi-Domain, Data Labeling, RLHF | San Francisco, New York, Bangalore, Berlin | 250+ | Series F, $200M (2022, ~$3.5B valuation) | - |
| Harmonic Mathematical superintelligence via formally verified reasoning | Math | Palo Alto | 26-50 | Series B, $100M at ~$900M valuation (Kleiner Perkins, 2025) | - |
| Hillclimb Math environments emphasizing verifiable correctness | Math | San Francisco | 1-10 | - | - |
| HUD Evals and RL environments platform for computer-use agents | Computer Use, Coding, Long Horizon, Agents Infrastructure, Custom Environments | San Francisco, Singapore | 11-25 | $15M raised (YC W25, Exceptional Capital) | - |
| Huzzle Labs Long-horizon code, tool-use, and enterprise workflow environments | Long Horizon, Coding, Enterprise | London, Berlin, San Francisco | 26-50 | ~$6M (reported) | - |
| Idler Code environments with realistic execution constraints | Coding | San Francisco | 11-25 | - | - |
| Incalmo Offensive-security environments for AI agents | Cybersecurity | San Mateo | 11-25 | - | - |
| Invariant Labs Security testing and analysis for AI agents; acquired by Snyk | Cybersecurity, Alignment | Zurich | 1-10 | Acquired by Snyk (June 2025) | No |
| Kaizen Browser agents that simulate and automate legacy web portal work | Browser, Enterprise | San Francisco | 1-10 | Seed, $500K (Y Combinator, Pioneer Fund) | - |
| Kernel Browser infrastructure for AI agents | Browser, Agents Infrastructure | San Francisco | 1-10 | Seed + Series A, $22M (Accel, 2025) | - |
| Latch Biology data infrastructure turned life-science environments and benchmarks | Science | San Francisco | 11-25 | Series A, $15M (2022) | - |
| LMArena Crowdsourced model leaderboards from the Chatbot Arena team | Multi-Domain, Machine Learning | San Francisco, Berkeley | 26-50 | $100M seed (a16z, UC Investments, 2025); $150M at $1.7B valuation (Jan 2026) | - |
| Math, Inc. Autoformalization agents for verified mathematics | Math | Palo Alto | 1-10 | Seed, $15M (Torch Capital, Robot Ventures, Chapter One) | - |
| Matrices Browser-native training environments for web agents | Browser, Computer Use | San Francisco | 1-10 | ~$5M (reported) | - |
| Mechanize RL environments to automate software engineering, founded by ex-Epoch AI researchers | Coding | San Francisco | 51-100 | ~$9.1M (reported) | - |
| Mercor Expert marketplace powering evals and RL environments for frontier labs | Multi-Domain, Data Labeling, RLHF | San Francisco | 250+ | Series C, $350M at $10B valuation (Oct 2025) | Yes |
| Metaphi Code and enterprise RL environments | Coding, Enterprise | San Francisco, New York | 1-10 | - | - |
| Micro1 Vetted domain experts for AI training data and evals | Data Labeling, Multi-Domain, RLHF | Los Angeles | 101-250 | Series A, $35M at $500M valuation (Sept 2025) | - |
| Morph Infinibranch VMs: snapshot and branch entire environments for agents | Agents Infrastructure | San Francisco | 11-25 | ~$19M total (seed led by Khosla Ventures) | - |
| Normal Hardware engineering environments for AI models | Hardware Engineering | San Francisco | 1-10 | - | - |
| Nous Research Open AI lab behind the Atropos RL environments framework | Machine Learning, Agents Infrastructure | New York | 11-25 | Series A, $50M led by Paradigm (~$1B valuation, 2025) | - |
| Originator Long-horizon computer-use environments | Computer Use, Coding, Long Horizon | London, Paris | 1-10 | - | - |
| Osmosis Forward-deployed reinforcement learning for AI agents | Machine Learning | San Francisco | 1-10 | Seed, $7M (CRV, Audacious Ventures, YC) | - |
| Pareto Expert data workforce for RLHF, evals, and tool-use environments | Multi-Domain, Tool Use, Data Labeling, RLHF | San Francisco | 26-50 | - | - |
| Phinity Chip-design RL environments | Chip Design | San Francisco, Palo Alto | 1-10 | - | - |
| Plato High-fidelity replicas of websites and software for agent training | Browser, Enterprise, Simulation | San Francisco | 1-10 | - | - |
| pre.dev Software-planning platform offering coding and long-horizon RL environments | Coding, Long Horizon | Delaware | 1-10 | - | - |
| Preference Model Stealth startup working on preference and reward modeling | Machine Learning, Coding | San Francisco, Toronto, Seattle | 11-25 | - | - |
| Prime Intellect Open superintelligence stack: compute, RL environments hub, and sandboxes | Machine Learning, Agents Infrastructure, Multi-Domain | San Francisco | 26-50 | Series A, $130M at $1B valuation (2026); $150M+ total | - |
| Proximal Long-horizon coding RL environments built from real codebases | Coding, Long Horizon | San Francisco, Bangalore | 26-50 | - | - |
| Quesma Security-domain RL environments and binary analysis evals | Cybersecurity, Coding | Warsaw | 11-25 | - | - |
| ReasonCore Science and code reasoning environments and benchmarks | Science, Coding | San Francisco | 11-25 | - | - |
| Refresh Simulation engines with verifiable rewards for coding and computer use | Coding, Computer Use, Simulation | San Francisco | 1-10 | - | - |
| Rise Data Labs US-based expert human data and custom RL environments for enterprise AI | Computer Use, Coding, Finance, Cybersecurity, Legal, Data Labeling, Enterprise, Multi-Domain, RLHF, Custom Environments | United States | - | - | Yes |
| Runloop Devboxes for training and evaluating coding agents | Coding, Agents Infrastructure | San Francisco | 1-10 | Seed, ~$7M | - |
| Scale Data-labeling incumbent extending into agent evals and RL environments | Multi-Domain, Coding, Data Labeling, RLHF | San Francisco, New York, Washington DC, London | 250+ | $1.6B+ raised; Meta invested $14.3B at ~$29B valuation (June 2025) | - |
| Scaled Foundations GRID: a simulation-first platform for robot learning | Robotics, Simulation, Machine Learning | Seattle | 1-10 | - | - |
| Scrapybara Virtual desktops for computer-use agents | Computer Use, Agents Infrastructure | San Francisco | 1-10 | Seed, $500K (CRV) | - |
| Sepal AI Science-domain environments; acquired by Mercor | Science | San Francisco | 1-10 | Acquired by Mercor (Feb 2026) | No |
| Silverstream AI Infrastructure and training data for reliable autonomous web agents | Browser, Machine Learning | San Francisco | 1-10 | Pre-seed, $1.2M (Gradient Ventures) | - |
| Snorkel Programmatic data platform expanding into expert evals and RL environments | Multi-Domain, Coding, Machine Learning, Data Labeling, RLHF | San Francisco, Redwood City, New York | 101-250 | Series D, $100M at $1.3B valuation (2025) | - |
| Steel Open-source browser API for AI agents | Browser, Agents Infrastructure | - | 1-10 | ~$17M (reported) | - |
| Surge Bootstrapped human-data leader with a dedicated RL environments org | Multi-Domain, Data Labeling, RLHF | San Francisco, New York, Seattle | 101-250 | Bootstrapped; reported in talks to raise ~$1B at $25B+ valuation (2025) | Yes |
| SynthLabs Post-training research: synthetic data and scalable RL alignment | Machine Learning, Alignment, RLHF | San Francisco | 1-10 | Seed (M12 and First Spark Ventures, 2024) | - |
| Tacit Labs Life-science and long-horizon environments for AI models | Science, Long Horizon | San Francisco | 1-10 | - | - |
| Taste Labs Design-domain environments and evals for AI models | Design | New York | 26-50 | - | - |
| The LLM Data Company Evals and reward data tooling for LLM training | Machine Learning, Data Labeling, RLHF | - | 1-10 | - | - |
| Theta Enterprise RL environments | Enterprise | San Francisco | 1-10 | - | - |
| Trajectory Labs Alignment-focused environments for safe agent trajectories | Alignment | Berkeley, Toronto | 1-10 | - | - |
| Turing AGI infrastructure: coding data and RL environments at scale | Coding, Multi-Domain, Data Labeling, RLHF | San Francisco, Palo Alto, Gurugram | 250+ | Series E, $111M at $2.2B valuation (2025) | - |
| Ulam Math RL environments and RLVR trajectories | Math | Warsaw, London | 1-10 | - | - |
| Vals AI Independent domain-specific benchmarks for legal, finance, and tax AI | Legal, Finance, Medical | San Francisco | 1-10 | - | - |
| Veris AI High-fidelity simulated environments to train enterprise AI agents | Enterprise, Custom Environments, Simulation | New York | 1-10 | Seed, $8.5M (Decibel and Acrew, June 2025) | - |
| Verita AI Design and UX environments with human preference feedback | Design | San Francisco, Vancouver | 26-50 | - | - |
| Vetto AI Code and computer-use environments from ex-DeepMind/Instagram founders | Coding, Computer Use | San Francisco, São Paulo, London | 1-10 | - | - |
| Vmax Converts proprietary data into RL environments | Machine Learning, Custom Environments | San Francisco, New York | 1-10 | - | - |