Catalogue

Browse the open RL environment ecosystem

A working directory of the environments the open community has actually published, pulled from the hubs where they live, and mapped to where coverage is crowded, thin, or wide open. A snapshot, not a live mirror.

664
environments
5
sources searched
8
registries & standards
3/8
domains still greenfield

Where they live

664 environments, five sources,
one search.

The center of gravity is no longer any single trainer. It is the registries and standards that let an environment be published once and reused anywhere. The bar is the directory below, split by where each entry came from; pick a source to filter it.

  • An open standard (ORS) connecting agents to environments via tool-calling; NVIDIA, Nebius, Eigent as launch contributors.

  • Hub of Dockerized evaluation datasets run by the Harbor harness, which also generates RL rollouts.

  • Community registry of RL/agentic environments as versioned Python packages, the de-facto verifiers format.

  • Hand-picked anchors: the environments the field keeps building on, added where no registry lists them.

  • Benchmarks ship as green evaluator agents; competitors race against them on public, reproducible leaderboards run in isolated CI.

The directory

14 of 664
  1. Fh AviaryPrimeFuture House Aviary wrapper for verifiers - Scientific reasoning environments with tools · aviary · scientific-reasoning · multi-turn
  2. BixBenchOpenReward205 tasks · 1★ · Bioinformatics Benchmark (BixBench) is a dataset comprising over 50 real-world scenarios of practical biological data analysis with nearly 300…
  3. ChernobylOpenReward150 tasks · Chernobyl is a nuclear power plant management environment where agents operate reactors through crisis scenarios inspired by four real historical d…
  4. Grid2OpNotablePower-grid operation environment behind the L2RPN competitions. · energy · control
  5. CyberGymAgentBeatsCyberGym is a large-scale benchmark for evaluating AI agents on real-world cybersecurity tasks, using over 1,500 historical vulnerabilities from 188 production…
  6. FinRLNotableLeading financial-RL library; hundreds of stock/crypto/portfolio markets. · finance · gym
  7. AirlineRMOpenReward12 tasks · AirlineRM is an airline network revenue management environment where an agent operates a hub-and-spoke carrier over a 30-day horizon. The agent make…
  8. AppWorldNotableStateful world of 9 simulated apps and 457 APIs with programmatic checks. · office · tool-use
  9. Tau BenchPrimeτ-bench: Tool-Agent-User benchmark for conversational agents in customer service domains with user simulation · tau-bench · conversation · multi-turn
  10. OSWorld-VerifiedAgentBeatsOSWorld-Verified is an upgraded version of OSWorld for evaluating multimodal computer-use agents on 369 open-ended tasks across web and desktop applications, w…
  11. KernelbenchPrimeKernelBench environment for verifiers - GPU kernel generation benchmark · kernelbench · single-turn · gpu
  12. MLE-benchAgentBeatsMLE-bench evaluates how well AI agents perform real-world machine learning engineering by testing them on 75 Kaggle competitions spanning tasks like data prepa…
  13. scale-ai/swe-bench-proHarbor731 tasks · evaluation dataset on the Harbor hub, published by scale-ai.
  14. SWE-benchNotable2,294 real GitHub issue-resolution tasks with Dockerized per-instance evaluation. · coding · swe · docker

Eight domains

How the catalogue lands across the arena’s eight under-served domains. 3 are effectively greenfield. That asymmetry is the opportunity.

Sciences: Math and reasoning are well-served; wet-lab chemistry, biology, and open-ended discovery are sparse.

Found a gap? An environment is a task.md, a Docker image, a verifier, and an oracle. We post-train on it and score what generalizes.