LLMRepos·Dataset Collections LLM projects

Updated daily

Browse 57 open-source dataset collections projects in Lists. Compare GitHub stars, recent growth, languages, licenses, and repository activity.

Dataset Collections

Browse 57 open-source dataset collections projects in Lists. Compare GitHub stars, recent growth, languages, licenses, and repository activity.

57
Repositories
480
Lists

Top Dataset Collections repositories

Ranked by current GitHub stars from the latest LLMRepos snapshot.

Showing 40 of 57

收集和梳理垂直领域的开源模型、数据集及评测基准。

MIT License+1 stars in 7dupdated 972d ago

A collection of AWESOME things about Graph-Related LLMs.

MIT License-2 stars in 7dupdated 292d ago

A list of awesome papers and resources of recommender system on large language model (LLM).

+3 stars in 7dupdated 525d ago

🚀🚀🚀 A collection of some awesome public YOLO object detection series projects and the related object detection datasets.

+1 stars in 7dupdated 451d ago

Summarize existing representative LLMs text datasets.

Apache License 2.0-1 stars in 7dupdated 166d ago

💎 A list of accessible speech corpora for ASR, TTS, and other Speech Technologies

MIT License+0 stars in 7dupdated 809d ago

[ICLR 2025 Oral] Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows

HTMLMIT License+5 stars in 7dupdated 12d ago

DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents

PythonApache License 2.0+3 stars in 7dupdated 105d ago

🚀🚀🚀A collection of some awesome public projects about Large Language Model(LLM), Vision Language Model(VLM), Vision Language Action(VLA), AI Generated Content(AIGC), the related Datasets and Applications.

+0 stars in 7dupdated 388d ago

Curated List of Persian Natural Language Processing and Information Retrieval Tools and Resources

+1 stars in 7dupdated 1,021d ago

A collection of awesome-prompt-datasets, awesome-instruction-dataset, to train ChatLLM such as chatgpt 收录各种各样的指令数据集, 用于训练 ChatLLM 模型。

Apache License 2.0+1 stars in 7dupdated 68d ago

590+ usernames in this dictionary! A list of reserved usernames to prevent url collision with resource paths. This repository hosts the list in multiple formats like JSON, CSV, SQL and plain text. You can use its just download its by wget.

PHPMIT License+1 stars in 7dupdated 1,186d ago

CRUD-RAG: A Comprehensive Chinese Benchmark for Retrieval-Augmented Generation of Large Language Models

Python+2 stars in 7dupdated 461d ago

Data and code for FreshLLMs (https://arxiv.org/abs/2310.03214)

Jupyter NotebookApache License 2.0+1 stars in 7dupdated 115d ago

Community model zoo for Apple Core AI (iOS/macOS 27): 62 models — LLM, VLM, OCR, ASR, TTS, image/video/music gen, forecasting — each gated against its source model and shipped with the recipe that produced it. Downloadable from Hugging Face, runnable in one line of Swift via CoreAIKit. Plus benchmarks, Metal kernels, knowledge base.

PythonOther+9 stars in 7dupdated today

[TACL'26] VoiceBench: Benchmarking LLM-Based Voice Assistants

PythonApache License 2.0+5 stars in 7dupdated 12d ago
386
stars

This repository contains related work, benchmarks and datasets for the paper "Large Language Models in Finance (FinLLMs)".

+1 stars in 7dupdated 501d ago

Awesome graph anomaly detection techniques built based on deep learning frameworks. Collections of commonly used datasets, papers as well as implementations are listed in this github repository. We also invite researchers interested in anomaly detection, graph representation learning, and graph anomaly detection to join this project as contributors and boost further research in this area.

MIT License+1 stars in 7dupdated 1,141d ago

(CVPR2024)A benchmark for evaluating Multimodal LLMs using multiple-choice questions.

PythonOther+1 stars in 7dupdated 587d ago

A comprehensive set of LLM benchmark scores and provider prices. (deprecated, read more in README)

JavaScriptOther+0 stars in 7dupdated 304d ago

A curriculum-aligned knowledge graph, benchmark, and multimodal training dataset for evaluating and improving curriculum cognition in educational LLMs.

PythonOther+15 stars in 7dupdated 15d ago

BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent (ACL 2026 Main)

PythonMIT License+8 stars in 7dupdated 88d ago

(NeurIPS D&B 2024) STaRK: Benchmarking LLM Retrieval on Textual and Relational Knowledge Bases

PythonMIT License+0 stars in 7dupdated 199d ago

Research progress on speech deepfake detection: Relevant datasets aggregated from the review literature and publicly available codes

+0 stars in 7dupdated 446d ago
313
stars

A multi-programming language benchmark for LLMs

PythonOther+0 stars in 7dupdated 134d ago

A benchmark for evaluating LLMs on open-ended CS problems. Exploring the Next Frontier of Computer Science.

C++MIT License+9 stars in 7dupdated 3d ago

Official github repo for SafetyBench, a comprehensive benchmark to evaluate LLMs' safety. [ACL 2024]

PythonMIT License-1 stars in 7dupdated 392d ago

Web-Bench is a benchmark designed to evaluate the performance of LLMs in actual Web development.

JavaScriptApache License 2.0+2 stars in 7dupdated 116d ago

ToolQA, a new dataset to evaluate the capabilities of LLMs in answering challenging questions with external tools. It offers two levels (easy/hard) across eight real-life scenarios.

Jupyter NotebookApache License 2.0+1 stars in 7dupdated 1,101d ago
285
stars

①[ICLR2024 Spotlight] (GPT-4V/Gemini-Pro/Qwen-VL-Plus+16 OS MLLMs) A benchmark for multi-modality LLMs (MLLMs) on low-level vision and visual quality assessment.

Jupyter NotebookOther-2 stars in 7dupdated 742d ago
285
stars

A ChatGPT(GPT-3.5) & GPT-4 Workload Trace to Optimize LLM Serving Systems

PythonCreative Commons Attribution 4.0 International-1 stars in 7dupdated 55d ago

Vision Document Retrieval (ViDoRe): Benchmark. Evaluation code for the ColPali paper.

PythonMIT License-1 stars in 7dupdated 152d ago

Towards Large Multimodal Models as Visual Foundation Agents

PythonApache License 2.0+0 stars in 7dupdated 487d ago
254
stars

The official evaluation suite and dynamic data release for MixEval.

Python+0 stars in 7dupdated 653d ago
254
stars

BABILong is a benchmark for LLM evaluation using the needle-in-a-haystack approach.

Jupyter NotebookApache License 2.0+1 stars in 7dupdated 84d ago

Hallucinations (Confabulations) Document-Based Benchmark for RAG. Includes human-verified questions and answers.

HTML+0 stars in 7dupdated 382d ago