LLMReposยทReinforcement Learning / RLHF LLM projects

Updated daily

Browse 50 open-source reinforcement learning / rlhf projects in Model Development. Compare GitHub stars, recent growth, languages, licenses, and repository activity.

Reinforcement Learning / RLHF

Browse 50 open-source reinforcement learning / rlhf projects in Model Development. Compare GitHub stars, recent growth, languages, licenses, and repository activity.

50
Repositories
452
Model Development

Top Reinforcement Learning / RLHF repositories

Ranked by current GitHub stars from the latest LLMRepos snapshot.

Showing 40 of 50

An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL)

PythonApache License 2.0+26 stars in 7dupdated 11d ago
5.8k
stars

Democratizing Reinforcement Learning for LLMs

PythonApache License 2.0+12 stars in 7dupdated today
2.8k
stars

Agent RL framework for LLM agents: multi-turn reinforcement learning with StarPO and reasoning-collapse diagnostics

PythonMIT Licenseupdated 1d ago

verl-agent is an extension of veRL, designed for training LLM/VLM agents via RL. verl-agent is also the official code for paper "Group-in-Group Policy Optimization for LLM Agent Training"

PythonApache License 2.0+24 stars in 7dupdated 76d ago
1.9k
stars

Scalable RL solution for advanced reasoning of language models

PythonApache License 2.0+2 stars in 7dupdated 524d ago
1.6k
stars

Agent-R1: Training Powerful LLM Agents with End-to-End Reinforcement Learning

PythonMIT License+15 stars in 7dupdated today

Safe RLHF: Constrained Value Alignment via Safe Reinforcement Learning from Human Feedback

PythonApache License 2.0-1 stars in 7dupdated 273d ago
1.3k
stars

The official implementation of Self-Play Fine-Tuning (SPIN)

PythonApache License 2.0+0 stars in 7dupdated 838d ago
1.1k
stars

AxisRL is an agentic RL post-training framework built on SGLang rollout, Megatron training, and real-world agent workflows.

PythonApache License 2.0+4 stars in 7dupdated 21d ago
985
stars

๐Ÿš€ Ultra Recipe for Training Long-Horizon Search Agents - matching frontier AI's search capability with a 20B model + stateful harness

PythonApache License 2.0+2 stars in 7dupdated 70d ago

Code and implementations for the paper "AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning" by Zhiheng Xi et al.

PythonMIT License+4 stars in 7dupdated 190d ago

A simulation framework for RLHF and alternatives. Develop your RLHF method without collecting human data.

PythonApache License 2.0+0 stars in 7dupdated 784d ago

Scaling Deep Research via Reinforcement Learning in Real-world Environments.

PythonApache License 2.0+3 stars in 7dupdated 106d ago

Official repository for DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research

PythonApache License 2.0+3 stars in 7dupdated 68d ago

Single File, Single GPU, From Scratch, Efficient, Full Parameter Tuning library for "RL for LLMs"

Jupyter NotebookMIT License+0 stars in 7dupdated 321d ago
589
stars

The official implementation of Self-Play Preference Optimization (SPPO)

PythonApache License 2.0-1 stars in 7dupdated 579d ago

An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale

PythonApache License 2.0+6 stars in 7dupdated 4d ago
498
stars

Official implementation of GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

PythonApache License 2.0-1 stars in 7dupdated 96d ago

LLM Tuning with PEFT (SFT+RM+PPO+DPO with LoRA)

Python+1 stars in 7dupdated 1,048d ago

Agentic RAG R1 Framework via Reinforcement Learning

PythonApache License 2.0+1 stars in 7dupdated 56d ago
416
stars

[ICML 2026] Let LLMs invent and evolve languages for efficient reasoning.

PythonMIT License+7 stars in 7dupdated 25d ago

LightRFT: Light, Efficient, Omni-modal & Reward-model Driven Reinforcement Fine-Tuning Framework

PythonApache License 2.0-1 stars in 7dupdated 7d ago

[NeurIPS 2025] Flow x RL. "ReinFlow: Fine-tuning Flow Policy with Online Reinforcement Learning". Support VLAs e.g., Pi0, Pi0.5, GR00TN1.5. Fully open-sourced.

PythonMIT License+1 stars in 7dupdated 122d ago

AdaRubric: Adaptive Dynamic Rubric Evaluator for Agent Trajectories

PythonApache License 2.0+1 stars in 7dupdated today
347
stars

Official code for "Self-Distilled Agentic Reinforcement Learning"

PythonApache License 2.0+6 stars in 7dupdated 5d ago

Super-Efficient RLHF Training of LLMs with Parameter Reallocation

PythonApache License 2.0+0 stars in 7dupdated 487d ago

The purpose of this repository is to make prototypes as case study in the context of proof of concept(PoC) and research and development(R&D) that I have written in my website. The main research topics are Auto-Encoders in relation to the representation learning, the statistical machine learning for energy-based models, adversarial generation networks(GANs), Deep Reinforcement Learning such as Deep Q-Networks, semi-supervised learning, and neural network language model for natural language processing.

PythonGNU General Public License v2.0+0 stars in 7dupdated 972d ago

An easy-to-use, fast toolkit to scale up RL post-training on a single node.

PythonApache License 2.0+16 stars in 7dupdated today
293
stars

[NeurIPS 2024] ReEvo: Large Language Models as Hyper-Heuristics with Reflective Evolution

PythonMIT License+1 stars in 7dupdated 212d ago

Train a Language Model with GRPO to create a schedule from a list of events and priorities

Jupyter NotebookApache License 2.0+0 stars in 7dupdated 138d ago

Benchmark and research code for the paper SWEET-RL Training Multi-Turn LLM Agents onCollaborative Reasoning Tasks

PythonOther+0 stars in 7dupdated 476d ago

MiroRL is an MCP-first reinforcement learning framework for deep research agent.

PythonApache License 2.0+1 stars in 7dupdated 362d ago

Official Repo of "RobustFlow: Towards Robust Agentic Workflow Generation"

Python+0 stars in 7dupdated 309d ago

Cutting-edge platform for LLM agent tuning. Deliver RL tuning with flexibility, reliability, speed, multi-agent optimization and realtime community benchmarking.

PythonApache License 2.0-1 stars in 7dupdated 11d ago

Fixing GRPO training collapse in long-horizon multi-tool agents. A lightweight PRM-Lite + LATA joint approach achieves +37% over vanilla GRPO on ฯ„-bench airline (50-task, multi-turn).

Python+19 stars in 7dupdated 58d ago

A framework for agentic tool use training with reinforcement learning

Python+2 stars in 7dupdated 138d ago

Comprehensive toolkit for Reinforcement Learning from Human Feedback (RLHF) training, featuring instruction fine-tuning, reward model training, and support for PPO and DPO algorithms with various configurations for the Transformer 5 models including Qwen 3.0

Python+0 stars in 7dupdated 10d ago

[ICCV] Social NCE: Contrastive Learning of Socially-aware Motion Representations

PythonBSD 2-Clause "Simplified" License+0 stars in 7dupdated 1,506d ago