An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL)
LLMReposยทReinforcement Learning / RLHF LLM projects
Updated dailyBrowse 50 open-source reinforcement learning / rlhf projects in Model Development. Compare GitHub stars, recent growth, languages, licenses, and repository activity.
Reinforcement Learning / RLHF
Browse 50 open-source reinforcement learning / rlhf projects in Model Development. Compare GitHub stars, recent growth, languages, licenses, and repository activity.
Top Reinforcement Learning / RLHF repositories
Ranked by current GitHub stars from the latest LLMRepos snapshot.
Showing 40 of 50
Democratizing Reinforcement Learning for LLMs
Agent RL framework for LLM agents: multi-turn reinforcement learning with StarPO and reasoning-collapse diagnostics
verl-agent is an extension of veRL, designed for training LLM/VLM agents via RL. verl-agent is also the official code for paper "Group-in-Group Policy Optimization for LLM Agent Training"
Scalable RL solution for advanced reasoning of language models
Agent-R1: Training Powerful LLM Agents with End-to-End Reinforcement Learning
Safe RLHF: Constrained Value Alignment via Safe Reinforcement Learning from Human Feedback
Recipes to train reward model for RLHF.
The official implementation of Self-Play Fine-Tuning (SPIN)
AxisRL is an agentic RL post-training framework built on SGLang rollout, Megatron training, and real-world agent workflows.
๐ Ultra Recipe for Training Long-Horizon Search Agents - matching frontier AI's search capability with a 20B model + stateful harness
Code and implementations for the paper "AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning" by Zhiheng Xi et al.
A simulation framework for RLHF and alternatives. Develop your RLHF method without collecting human data.
Scaling Deep Research via Reinforcement Learning in Real-world Environments.
Official repository for DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research
Single File, Single GPU, From Scratch, Efficient, Full Parameter Tuning library for "RL for LLMs"
uclaml/SPPO
The official implementation of Self-Play Preference Optimization (SPPO)
An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale
NVlabs/GDPO
Official implementation of GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
LLM Tuning with PEFT (SFT+RM+PPO+DPO with LoRA)
Agentic RAG R1 Framework via Reinforcement Learning
[ICML 2026] Let LLMs invent and evolve languages for efficient reasoning.
LightRFT: Light, Efficient, Omni-modal & Reward-model Driven Reinforcement Fine-Tuning Framework
[NeurIPS 2025] Flow x RL. "ReinFlow: Fine-tuning Flow Policy with Online Reinforcement Learning". Support VLAs e.g., Pi0, Pi0.5, GR00TN1.5. Fully open-sourced.
AdaRubric: Adaptive Dynamic Rubric Evaluator for Agent Trajectories
Official code for "Self-Distilled Agentic Reinforcement Learning"
Super-Efficient RLHF Training of LLMs with Parameter Reallocation
The purpose of this repository is to make prototypes as case study in the context of proof of concept(PoC) and research and development(R&D) that I have written in my website. The main research topics are Auto-Encoders in relation to the representation learning, the statistical machine learning for energy-based models, adversarial generation networks(GANs), Deep Reinforcement Learning such as Deep Q-Networks, semi-supervised learning, and neural network language model for natural language processing.
An easy-to-use, fast toolkit to scale up RL post-training on a single node.
Deep Research
ai4co/reevo
[NeurIPS 2024] ReEvo: Large Language Models as Hyper-Heuristics with Reflective Evolution
Train a Language Model with GRPO to create a schedule from a list of events and priorities
Benchmark and research code for the paper SWEET-RL Training Multi-Turn LLM Agents onCollaborative Reasoning Tasks
MiroRL is an MCP-first reinforcement learning framework for deep research agent.
Official Repo of "RobustFlow: Towards Robust Agentic Workflow Generation"
Cutting-edge platform for LLM agent tuning. Deliver RL tuning with flexibility, reliability, speed, multi-agent optimization and realtime community benchmarking.
Fixing GRPO training collapse in long-horizon multi-tool agents. A lightweight PRM-Lite + LATA joint approach achieves +37% over vanilla GRPO on ฯ-bench airline (50-task, multi-turn).
A framework for agentic tool use training with reinforcement learning
Comprehensive toolkit for Reinforcement Learning from Human Feedback (RLHF) training, featuring instruction fine-tuning, reward model training, and support for PPO and DPO algorithms with various configurations for the Transformer 5 models including Qwen 3.0
[ICCV] Social NCE: Contrastive Learning of Socially-aware Motion Representations