收集和梳理垂直领域的开源模型、数据集及评测基准。
LLMRepos·Dataset Collections LLM projects
Updated dailyBrowse 57 open-source dataset collections projects in Lists. Compare GitHub stars, recent growth, languages, licenses, and repository activity.
Dataset Collections
Browse 57 open-source dataset collections projects in Lists. Compare GitHub stars, recent growth, languages, licenses, and repository activity.
Top Dataset Collections repositories
Ranked by current GitHub stars from the latest LLMRepos snapshot.
Showing 40 of 57
A collection of AWESOME things about Graph-Related LLMs.
A list of awesome papers and resources of recommender system on large language model (LLM).
🚀🚀🚀 A collection of some awesome public YOLO object detection series projects and the related object detection datasets.
Summarize existing representative LLMs text datasets.
💎 A list of accessible speech corpora for ASR, TTS, and other Speech Technologies
[ICLR 2025 Oral] Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
🚀🚀🚀A collection of some awesome public projects about Large Language Model(LLM), Vision Language Model(VLM), Vision Language Action(VLA), AI Generated Content(AIGC), the related Datasets and Applications.
Curated List of Persian Natural Language Processing and Information Retrieval Tools and Resources
A collection of awesome-prompt-datasets, awesome-instruction-dataset, to train ChatLLM such as chatgpt 收录各种各样的指令数据集, 用于训练 ChatLLM 模型。
A Survey on Text-to-Video Generation/Synthesis.
590+ usernames in this dictionary! A list of reserved usernames to prevent url collision with resource paths. This repository hosts the list in multiple formats like JSON, CSV, SQL and plain text. You can use its just download its by wget.
Dataset and benchmark for RAG on company internal documents.
CRUD-RAG: A Comprehensive Chinese Benchmark for Retrieval-Augmented Generation of Large Language Models
Data and code for FreshLLMs (https://arxiv.org/abs/2310.03214)
Community model zoo for Apple Core AI (iOS/macOS 27): 62 models — LLM, VLM, OCR, ASR, TTS, image/video/music gen, forecasting — each gated against its source model and shipped with the recipe that produced it. Downloadable from Hugging Face, runnable in one line of Swift via CoreAIKit. Plus benchmarks, Metal kernels, knowledge base.
A curated list of Human Preference Datasets for LLM fine-tuning, RLHF, and eval.
[TACL'26] VoiceBench: Benchmarking LLM-Based Voice Assistants
This repository contains related work, benchmarks and datasets for the paper "Large Language Models in Finance (FinLLMs)".
Awesome graph anomaly detection techniques built based on deep learning frameworks. Collections of commonly used datasets, papers as well as implementations are listed in this github repository. We also invite researchers interested in anomaly detection, graph representation learning, and graph anomaly detection to join this project as contributors and boost further research in this area.
(CVPR2024)A benchmark for evaluating Multimodal LLMs using multiple-choice questions.
A comprehensive set of LLM benchmark scores and provider prices. (deprecated, read more in README)
A curriculum-aligned knowledge graph, benchmark, and multimodal training dataset for evaluating and improving curriculum cognition in educational LLMs.
BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent (ACL 2026 Main)
(NeurIPS D&B 2024) STaRK: Benchmarking LLM Retrieval on Textual and Relational Knowledge Bases
Research progress on speech deepfake detection: Relevant datasets aggregated from the review literature and publicly available codes
A multi-programming language benchmark for LLMs
A benchmark for evaluating LLMs on open-ended CS problems. Exploring the Next Frontier of Computer Science.
Official github repo for SafetyBench, a comprehensive benchmark to evaluate LLMs' safety. [ACL 2024]
Web-Bench is a benchmark designed to evaluate the performance of LLMs in actual Web development.
ToolQA, a new dataset to evaluate the capabilities of LLMs in answering challenging questions with external tools. It offers two levels (easy/hard) across eight real-life scenarios.
①[ICLR2024 Spotlight] (GPT-4V/Gemini-Pro/Qwen-VL-Plus+16 OS MLLMs) A benchmark for multi-modality LLMs (MLLMs) on low-level vision and visual quality assessment.
A ChatGPT(GPT-3.5) & GPT-4 Workload Trace to Optimize LLM Serving Systems
Vision Document Retrieval (ViDoRe): Benchmark. Evaluation code for the ColPali paper.
Towards Large Multimodal Models as Visual Foundation Agents
The official evaluation suite and dynamic data release for MixEval.
BABILong is a benchmark for LLM evaluation using the needle-in-a-haystack approach.
Hallucinations (Confabulations) Document-Based Benchmark for RAG. Includes human-verified questions and answers.