LLMRepos·Pretraining & Training LLM projects

Updated daily

Browse 77 open-source pretraining & training projects in Model Development. Compare GitHub stars, recent growth, languages, licenses, and repository activity.

Pretraining & Training

Browse 77 open-source pretraining & training projects in Model Development. Compare GitHub stars, recent growth, languages, licenses, and repository activity.

77
Repositories
452
Model Development

Top Pretraining & Training repositories

Ranked by current GitHub stars from the latest LLMRepos snapshot.

Showing 40 of 77

A simple, performant, and scalable Jax LLM!

PythonApache License 2.0+14 stars in 7dupdated today
2.2k
stars

Vendor-agnostic orchestration for training, inference and agentic workloads across NVIDIA, AMD, TPU, and Tenstorrent on clouds, Kubernetes, and bare metal.

PythonMozilla Public License 2.0+5 stars in 7dupdated today

[ICML'24] Magicoder: Empowering Code Generation with OSS-Instruct

PythonMIT License+0 stars in 7dupdated 661d ago
1.7k
stars

Automated Machine Learning on Kubernetes

PythonApache License 2.0+2 stars in 7dupdated 18d ago
1.3k
stars

The official implementation of Self-Play Fine-Tuning (SPIN)

PythonApache License 2.0+0 stars in 7dupdated 838d ago

(ICLR 2025) TabM: Advancing Tabular Deep Learning With Parameter-Efficient Ensembling

PythonApache License 2.0+5 stars in 7dupdated 287d ago

Tencent Pre-training framework in PyTorch & Pre-trained Model Zoo

PythonOther+1 stars in 7dupdated 750d ago
1.1k
stars

AxisRL is an agentic RL post-training framework built on SGLang rollout, Megatron training, and real-world agent workflows.

PythonApache License 2.0+4 stars in 7dupdated 21d ago

Fast & Simple Resource-Constrained Learning of Deep Network Structure

PythonApache License 2.0+0 stars in 7dupdated 53d ago

Byted PyTorch Distributed for Hyperscale Training of LLMs and RLs

PythonApache License 2.0+1 stars in 7dupdated 174d ago
985
stars

🚀 Ultra Recipe for Training Long-Horizon Search Agents - matching frontier AI's search capability with a 20B model + stateful harness

PythonApache License 2.0+2 stars in 7dupdated 70d ago
937
stars

Pure Rust implementation of a minimal Generative Pretrained Transformer

RustMIT License+1 stars in 7dupdated 307d ago

🚀 Pytorch Distributed native training library for LLMs/VLMs with OOTB Hugging Face support

PythonApache License 2.0+24 stars in 7dupdated today

Unified Reinforcement Learning Framework

PythonApache License 2.0+1 stars in 7dupdated 717d ago

The most atomic way to train and inference a GPT in pure, dependency-free C

CMIT Licenseupdated 7d ago

An open-source solution for full parameter fine-tuning of DeepSeek-V3/R1 671B, including complete code and scripts from training to inference, as well as some practical experiences and conclusions. (DeepSeek-V3/R1 满血版 671B 全参数微调的开源解决方案,包含从训练到推理的完整代码和脚本,以及实践中积累一些经验和结论。)

PythonApache License 2.0+0 stars in 7dupdated 529d ago

Scaling Deep Research via Reinforcement Learning in Real-world Environments.

PythonApache License 2.0+3 stars in 7dupdated 106d ago
726
stars

The PyTorch implementation of Generative Pre-trained Transformers (GPTs) using Kolmogorov-Arnold Networks (KANs) for language modeling

PythonMIT License+1 stars in 7dupdated 638d ago
722
stars

The official implementation of MARS: Unleashing the Power of Variance Reduction for Training Large Models

PythonApache License 2.0+0 stars in 7dupdated 151d ago

Official repository for DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research

PythonApache License 2.0+3 stars in 7dupdated 68d ago

[ICML 2026 & EMNLP 2026] Multimodal deep-research MLLM and benchmark. The first long-horizon multimodal deep-research MLLM, extending the number of reasoning turns to dozens and the number of search-engine interactions to hundreds.

PythonMIT License+2 stars in 7dupdated 16d ago

[SIGGRAPH Asia 2026] 4DAnyone: Create Anyone in 4D from a Casual Monocular Video

PythonApache License 2.0updated today

Official implementation of OneDiffusion paper (CVPR 2025)

PythonOther+0 stars in 7dupdated 618d ago

LLaSA: Scaling Train-time and Inference-time Compute for LLaMA-based Speech Synthesis

PythonOther-4 stars in 7dupdated 215d ago

Generative Adversarial Text to Image Synthesis / Please Star -->

Python+0 stars in 7dupdated 2,040d ago

Official implementation for "Blended Diffusion for Text-driven Editing of Natural Images" [CVPR 2022]

Jupyter NotebookMIT License+0 stars in 7dupdated 811d ago

An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale

PythonApache License 2.0+6 stars in 7dupdated 4d ago

unofficial vits2-TTS implementation in pytorch

PythonMIT License-1 stars in 7dupdated 879d ago

Named Entity Recognition using multilayered bidirectional LSTM

Python+0 stars in 7dupdated 2,725d ago

Official implementation for "Break-A-Scene: Extracting Multiple Concepts from a Single Image" [SIGGRAPH Asia 2023]

PythonApache License 2.0+0 stars in 7dupdated 953d ago

Projects from basic algorithms to MARL. Implements MADDPG,MATD3,MA/HAPPO in Predator-Prey pursuit games with PettingZoo MPE environments.

Jupyter NotebookMIT License+1 stars in 7dupdated 14d ago

Agentic RAG R1 Framework via Reinforcement Learning

PythonApache License 2.0+1 stars in 7dupdated 56d ago

Classify Kaggle Consumer Finance Complaints into 11 classes. Build the model with CNN (Convolutional Neural Network) and Word Embeddings on Tensorflow.

PythonApache License 2.0+0 stars in 7dupdated 83d ago