Agent RL framework for LLM agents: multi-turn reinforcement learning with StarPO and reasoning-collapse diagnostics
LLMRepos·RL Environments LLM projects
Updated dailyBrowse 51 open-source rl environments projects in Model Development. Compare GitHub stars, recent growth, languages, licenses, and repository activity.
RL Environments
Browse 51 open-source rl environments projects in Model Development. Compare GitHub stars, recent growth, languages, licenses, and repository activity.
Top RL Environments repositories
Ranked by current GitHub stars from the latest LLMRepos snapshot.
Showing 40 of 51
Implementation of all RL algorithms in a simpler way
A plugin for GTAV that transforms it into a vision-based self-driving car research environment.
A Python-based lightweight robot simulator designed for navigation, control, and learning
Modular Reinforcement Learning (RL) library (implemented in PyTorch, JAX, and NVIDIA Warp) with support for Gymnasium/Gym, NVIDIA Isaac Lab, MuJoCo Playground and other environments
Streamlining reinforcement learning with RLOps. State-of-the-art RL algorithms and tools, with 10x faster training through evolutionary hyperparameter optimization.
Multi-Agent Resource Optimization (MARO) platform is an instance of Reinforcement Learning as a Service (RaaS) for real-world resource optimization problems.
Code and implementations for the paper "AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning" by Zhiheng Xi et al.
Unified Reinforcement Learning Framework
SEED RL: Scalable and Efficient Deep-RL with Accelerated Central Inference. Implements IMPALA and R2D2 algorithms in TF2 with SEED's architecture.
MobileGym: A Verifiable and Highly Parallel Simulation Platform for Mobile GUI Agent Research · 浏览器里运行的安卓模拟器 · Browser-hosted Android Simulator · Verifiable Evaluation · Scalable Online RL Training (EMNLP 2026)
[CoRL '23] Dexterous piano playing with deep reinforcement learning.
[ICCV 2025] Official code of DeepMesh: Auto-Regressive Artist-mesh Creation with Reinforcement Learning
BenchMARL is a library for benchmarking Multi-Agent Reinforcement Learning (MARL). BenchMARL allows to quickly compare different MARL algorithms, tasks, and models while being systematically grounded in its two core tenets: reproducibility and standardization.
A collection of multi agent environments based on OpenAI gym.
Train robotic agents to learn pick and place with deep learning for vision-based manipulation in PyBullet. Transporter Nets, CoRL 2020.
API to run VirtualHome, a Multi-Agent Household Simulator
PyTorch implementations of various Deep Reinforcement Learning (DRL) algorithms for both single agent and multi-agent.
[ICLR 2023] Come & try Decision-Intelligence version of "Agar"! Gobigger could also help you with multi-agent decision intelligence study.
Multi-Robot Warehouse (RWARE): A multi-agent reinforcement learning environment
Projects from basic algorithms to MARL. Implements MADDPG,MATD3,MA/HAPPO in Predator-Prey pursuit games with PettingZoo MPE environments.
A Pytorch implementation of the multi agent deep deterministic policy gradients (MADDPG) algorithm
A Massively Parallel Large Scale Self-Play Framework
Deep Q-learning (DQN) for Multi-agent Reinforcement Learning (RL)
ExploitBench measures how far AI agents climb, from reaching vulnerable code, to triggering the bug, to building exploit primitives, to arbitrary code execution.
Official code for "Self-Distilled Agentic Reinforcement Learning"
Sotopia: an Open-ended Social Learning Environment (ICLR 2024 spotlight)
Standardized environment infrastructure for Agentic AI development.
Deep Reinforcement Learning in C#
Open-Source Framework for Development, Simulation and Benchmarking of Behavior Planning Algorithms for Autonomous Driving
Embodied Agent Interface (EAI): Benchmarking LLMs for Embodied Decision Making (NeurIPS D&B 2024 Oral)
[RA-Letter 2022] Reinforcement Learned Distributed Multi-Robot Navigation with Reciprocal Velocity Obstacle Shaped Rewards
Continuous CBS - a modification of conflict based search algorithm, that allows to perform actions (move, wait) of arbitrary duration. Timeline is not discretized, i.e. is continuous.
Reproduce results of the research article "Deep Reinforcement Learning Based Resource Allocation for V2V Communications"
[AAAI 2023] Official PyTorch implementation of paper "ACE: Cooperative Multi-agent Q-learning with Bidirectional Action-Dependency".
A Real-Time-Strategy game for Deep Learning research
Harness for running and evaluating AI agents against RL environments
🌱 A little course on Reinforcement Learning Environments for evaluating and training Language Models
Lightweight multi-agent gridworld Gym environment
A framework for creating rich, 3D, Minecraft-like single and multi-agent environments for AI research. (Accepted at ICML 2025).