LLMRepos·GPU Kernels LLM projects

Updated daily

Browse 18 open-source gpu kernels projects in Infrastructure. Compare GitHub stars, recent growth, languages, licenses, and repository activity.

GPU Kernels

Browse 18 open-source gpu kernels projects in Infrastructure. Compare GitHub stars, recent growth, languages, licenses, and repository activity.

18
Repositories
847
Infrastructure

Top GPU Kernels repositories

Ranked by current GitHub stars from the latest LLMRepos snapshot.

Showing 18 of 18

Efficient Triton Kernels for LLM Training

PythonBSD 2-Clause "Simplified" License+14 stars in 7dupdated today

:metal: TT-NN operator library, and TT-Metalium low level kernel programming model.

C++Apache License 2.0+12 stars in 7dupdated today

Benchmarking Deep Learning operations on different hardware

C++Apache License 2.0+0 stars in 7dupdated 1,947d ago

Pure Rust implementation of a minimal Generative Pretrained Transformer

RustMIT License+1 stars in 7dupdated 307d ago

The most atomic way to train and inference a GPT in pure, dependency-free C

CMIT Licenseupdated 7d ago

Cleora AI is a general-purpose open-source model for efficient, scalable learning of stable and inductive entity embeddings for heterogeneous relational data. Created by Synerise.com team.

Jupyter NotebookOther+0 stars in 7dupdated 144d ago
325
stars

Rust + CUDA inference engine for NVIDIA RTX PRO 6000 Blackwell and RTX 5090. Serves safetensors and GGUF over an OpenAI-compatible API, with per-device tuned defaults and speculative decode gated byte-identical to plain decode. Hosted instance: inference.tiyuvta.ai

RustMIT License+11 stars in 7dupdated today

An official lightweight library for the RaBitQ algorithm and its applications in vector search.

C++Apache License 2.0+21 stars in 7dupdated today
256
stars

Quantized LLM training in pure CUDA/C++.

C++Apache License 2.0+0 stars in 7dupdated 19d ago

Tiny ASIC implementation for "The Era of 1-bit LLMs All Large Language Models are in 1.58 Bits" matrix multiplication unit

VerilogApache License 2.0+0 stars in 7dupdated 857d ago
187
stars

Simulation framework for nonsmooth dynamical systems

CApache License 2.0+0 stars in 7dupdated 12d ago

OrbitKV: a Rust attention-state compiler and lifetime-safe KV block manager.

PythonMIT License+0 stars in 7dupdated 3d ago

OrbitKV: a Rust attention-state compiler and lifetime-safe KV block manager.

PythonMIT License+0 stars in 7dupdated 3d ago

OrbitKV: a Rust attention-state compiler and lifetime-safe KV block manager.

PythonMIT License+0 stars in 7dupdated 3d ago

CUDA编程练习项目-Hands-on CUDA kernels and performance optimization, covering GEMM, FlashAttention, Tensor Cores, CUTLASS, quantization, KV cache, NCCL, and profiling.

CudaMIT License-2 stars in 7dupdated 69d ago

[HPCA'21] SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning

ScalaMIT License+1 stars in 7dupdated 727d ago

An agent harness that compiles a model into one provably-correct, self-retargeting CUDA megakernel and self-tunes it past cuBLAS at batch-1 LLM decode, paper: https://arxiv.org/abs/2606.09682

PythonMIT License+3 stars in 7dupdated 56d ago