-
The ultimate guide to multi-harness RL
๐61Train open models with RL inside real agent harnesses
-
SmolDataEnvs Multi-harness | SETA Whitebox
๐งชAnswer dataset questions and view scoring results
-
SmolDataEnvs Multi-harness | Native OpenCode
๐งชRun OpenCode tasks and view grading results
-
SmolDataEnv RL
๐Visualize RL agent comparison metrics in a web dashboard
AI & ML interests
Open RL Environments at Scale
Recent Activity
View all activity
Try the three environments, follow the tutorial, and explore the datasets, models and historical RL/SFT results.
All 7,780 of Xiaomi's MiMo-V2.6 RL environments as Harbor tasks, set up and graded like Xiaomi's harness.
-
FineEnvs/MiMo-V2.6-RL-harbor-code
RL Environment โข Updated โข 4.75k โข 2 -
FineEnvs/MiMo-V2.6-RL-harbor-cyber
RL Environment โข Updated โข 1.52k โข 2 -
FineEnvs/MiMo-V2.6-RL-harbor-general
RL Environment โข Updated โข 1.47k -
FineEnvs/MiMo-V2.6-RL-harbor-terminal
RL Environment โข Updated โข 1.68k
Repo2RLEnv coding and terminal RL environments in Harbor format. Datasets include per-task quality labels, provenance and generation economics.
-
Geoguesser Environment
๐1Interact with a GeoGuessrโstyle environment through custom actions
-
FineEnvs/geoguesser-tasks
RL Environment โข Updated โข 3.65k โข 234 -
FineEnvs/geoguesser-qwen3.5-4b-grpo
Image-Text-to-Text โข Updated โข 74 -
FineEnvs/geoguesser-qwen3.5-4b-grpo-v3
Image-Text-to-Text โข Updated โข 85
Deterministic data-analysis agent tasks from the jupyter-agent dataset โ verified answers, no LLM judge. Harbor env suites, plain dataset & SFT.
One Wordle game packaged for OpenEnv, ORS and NeMo Gym, from FineEnvs' 00-environments-101: the same env, three RL frameworks side by side.
5.5K+ RL tasks for hill-climbing small models in code and data science. Deterministic grading, no LLM judge.
LaTeX OCR environment, model, dataset, and five-run training comparison across four models, including unstable and stabilized Gemma.
An RL environment where the agent paints by writing p5.brush sketches, rewarded by an aesthetic preference model looking at the render.
-
FineEnvs/watercolour-grpo-hps-only
Reinforcement Learning โข Updated โข 32 โข 1 -
FineEnvs/watercolour-reference-pool
RL Environment โข Updated โข 178 โข 2.23k โข 2 -
FineEnvs/watercolour-rollouts-hps-only
RL Environment โข Updated โข 470 โข 1.05k -
FineEnvs/watercolour-grpo-judge-led
Reinforcement Learning โข Updated โข 37 โข 1
A curated collection of articles, guides, tutorials, slides, and resources for learning how to build, train, and evaluate RL environments for Agents
-
The ultimate guide to RL environments: building and scaling them in the LLM era
๐252Building and scaling RL environments for LLM training
-
RL Environments 101 โ Slides
๐42Explore RL Environments 101 slide presentation
-
How to turn a game into an RL environment
๐68From an idea to a trained 4B, with the dead ends left in
Try the three environments, follow the tutorial, and explore the datasets, models and historical RL/SFT results.
-
The ultimate guide to multi-harness RL
๐61Train open models with RL inside real agent harnesses
-
SmolDataEnvs Multi-harness | SETA Whitebox
๐งชAnswer dataset questions and view scoring results
-
SmolDataEnvs Multi-harness | Native OpenCode
๐งชRun OpenCode tasks and view grading results
-
SmolDataEnv RL
๐Visualize RL agent comparison metrics in a web dashboard
One Wordle game packaged for OpenEnv, ORS and NeMo Gym, from FineEnvs' 00-environments-101: the same env, three RL frameworks side by side.
All 7,780 of Xiaomi's MiMo-V2.6 RL environments as Harbor tasks, set up and graded like Xiaomi's harness.
-
FineEnvs/MiMo-V2.6-RL-harbor-code
RL Environment โข Updated โข 4.75k โข 2 -
FineEnvs/MiMo-V2.6-RL-harbor-cyber
RL Environment โข Updated โข 1.52k โข 2 -
FineEnvs/MiMo-V2.6-RL-harbor-general
RL Environment โข Updated โข 1.47k -
FineEnvs/MiMo-V2.6-RL-harbor-terminal
RL Environment โข Updated โข 1.68k
5.5K+ RL tasks for hill-climbing small models in code and data science. Deterministic grading, no LLM judge.
Repo2RLEnv coding and terminal RL environments in Harbor format. Datasets include per-task quality labels, provenance and generation economics.
LaTeX OCR environment, model, dataset, and five-run training comparison across four models, including unstable and stabilized Gemma.
-
Geoguesser Environment
๐1Interact with a GeoGuessrโstyle environment through custom actions
-
FineEnvs/geoguesser-tasks
RL Environment โข Updated โข 3.65k โข 234 -
FineEnvs/geoguesser-qwen3.5-4b-grpo
Image-Text-to-Text โข Updated โข 74 -
FineEnvs/geoguesser-qwen3.5-4b-grpo-v3
Image-Text-to-Text โข Updated โข 85
An RL environment where the agent paints by writing p5.brush sketches, rewarded by an aesthetic preference model looking at the render.
-
FineEnvs/watercolour-grpo-hps-only
Reinforcement Learning โข Updated โข 32 โข 1 -
FineEnvs/watercolour-reference-pool
RL Environment โข Updated โข 178 โข 2.23k โข 2 -
FineEnvs/watercolour-rollouts-hps-only
RL Environment โข Updated โข 470 โข 1.05k -
FineEnvs/watercolour-grpo-judge-led
Reinforcement Learning โข Updated โข 37 โข 1
Deterministic data-analysis agent tasks from the jupyter-agent dataset โ verified answers, no LLM judge. Harbor env suites, plain dataset & SFT.
A curated collection of articles, guides, tutorials, slides, and resources for learning how to build, train, and evaluate RL environments for Agents
-
The ultimate guide to RL environments: building and scaling them in the LLM era
๐252Building and scaling RL environments for LLM training
-
RL Environments 101 โ Slides
๐42Explore RL Environments 101 slide presentation
-
How to turn a game into an RL environment
๐68From an idea to a trained 4B, with the dead ends left in