Toolkit for linearizing PDFs for LLM datasets/training
LLMRepos·Dataset Engineering LLM projects
Updated dailyBrowse 123 open-source dataset engineering projects in Model Development. Compare GitHub stars, recent growth, languages, licenses, and repository activity.
Dataset Engineering
Browse 123 open-source dataset engineering projects in Model Development. Compare GitHub stars, recent growth, languages, licenses, and repository activity.
Top Dataset Engineering repositories
Ranked by current GitHub stars from the latest LLMRepos snapshot.
Showing 40 of 123
Label, clean and enrich text datasets with LLMs.
Agent harness to publish your agent chat history as Huggingface datasets.
[ICML'24] Magicoder: Empowering Code Generation with OSS-Instruct
A repository that contains models, datasets, and fine-tuning techniques for DB-GPT, with the purpose of enhancing model performance in Text-to-SQL
Synthetic Data Generation Platform By DataArcTech
Scalable data pre processing and curation toolkit for LLMs
Synthetic data curation for post-training and structured data extraction
Build, enrich, and transform datasets using AI models with no code
Tool for generating high quality Synthetic datasets
Safe RLHF: Constrained Value Alignment via Safe Reinforcement Learning from Human Feedback
Data and tools for generating and inspecting OLMo pre-training data.
A large-scale text-to-image prompt gallery dataset based on Stable Diffusion
scripts and baselines for Spider: Yale complex and cross-domain semantic parsing and text-to-SQL challenge
Open source project for data preparation for GenAI applications
Generate High-Quality Synthetics, Train, Measure, and Evaluate in a Single Pipeline
[ICLR 2025] Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing. Your efficient and high-quality synthetic data generation pipeline!
🚀 Catalyst is a C# Natural Language Processing library built for speed. Inspired by spaCy's design, it brings pre-trained models, out-of-the box support for training word and document embeddings, and flexible entity recognition models.
[CVPR'23] OpenScene: 3D Scene Understanding with Open Vocabularies
A simulation framework for RLHF and alternatives. Develop your RLHF method without collecting human data.
A lightweight library for generating synthetic instruction tuning datasets for your data without GPT.
Train Models Contrastively in Pytorch
OpenChem: Deep Learning toolkit for Computational Chemistry and Drug Design Research
Dingo: A Comprehensive AI Data, Model and Application Quality Evaluation Tool
[ICML 2026 & EMNLP 2026] Multimodal deep-research MLLM and benchmark. The first long-horizon multimodal deep-research MLLM, extending the number of reasoning turns to dozens and the number of search-engine interactions to hundreds.
PAddle PARAllel text-to-speech toolKIT (supporting Tacotron2, Transformer TTS, FastSpeech2/FastPitch, SpeedySpeech, WaveFlow and Parallel WaveGAN)
limuloo/MIGC
[CVPR 2024 Highlight] MIGC and [TPAMI 2024] MIGC++ (Official Implementation)
Classify Kaggle San Francisco Crime Description into 39 classes. Build the model with CNN, RNN (GRU and LSTM) and Word Embeddings on Tensorflow.
Natural Language Processing Pipeline - Sentence Splitting, Tokenization, Lemmatization, Part-of-speech Tagging and Dependency Parsing
Nimfa: Nonnegative matrix factorization in Python
从0到1构建一个MiniLLM (pretrain+sft+dpo实践ä¸)
Cleora AI is a general-purpose open-source model for efficient, scalable learning of stable and inductive entity embeddings for heterogeneous relational data. Created by Synerise.com team.
Research code for metric depth estimation in periocular VR imagery using UE MetaHuman-generated data.
Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation Models
LongCite: Enabling LLMs to Generate Fine-grained Citations in Long-context QA
Implementation of paper Data Engineering for Scaling Language Models to 128K Context
Open source release of the evaluation benchmark suite described in "Realistic Evaluation of Deep Semi-Supervised Learning Algorithms"
EvaLearn is a pioneering benchmark designed to evaluate large language models (LLMs) on their learning capability and efficiency in challenging tasks.
Classify Kaggle Consumer Finance Complaints into 11 classes. Build the model with CNN (Convolutional Neural Network) and Word Embeddings on Tensorflow.