Fast inference engine for Transformer models
-
Updated
Aug 31, 2026 - C++
Fast inference engine for Transformer models
Tuned OpenCL BLAS
cuDNN Frontend is NVIDIA's modern, open-source entry point to the cuDNN library and a growing collection of high-performance open-source kernels.
BLISlab: A Sandbox for Optimizing GEMM
Several optimization methods of half-precision general matrix multiplication (HGEMM) using tensor core with WMMA API and MMA PTX instruction.
🚀🚀🚀 This repository lists some awesome public CUDA, cuda-python, cuBLAS, cuDNN, CUTLASS, TensorRT, TensorRT-LLM, Triton, TVM, MLIR, PTX and High Performance Computing (HPC) projects.
Optimizing SGEMM kernel functions on NVIDIA GPUs to a close-to-cuBLAS performance.
Multi-Threaded FP32 Matrix Multiplication on x86 CPUs
The HPC toolbox: fused matrix multiplication, convolution, data-parallel strided tensor primitives, OpenMP facilities, SIMD, JIT Assembler, CPU detection, state-of-the-art vectorized BLAS for floats and integers
[DEPRECATED] Moved to ROCm/rocm-libraries repo
AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming
A Flexible and Energy Efficient Accelerator For Sparse Convolution Neural Network
DBCSR: Distributed Block Compressed Sparse Row matrix library
Hand-written NVFP4 W4A16 CUDA kernels for Volta
FP64 equivalent GEMM by the Ozaki scheme with Int8 Tensor Cores
[DEPRECATED] Moved to ROCm/rocm-libraries repo
To associate your repository with the gemm topic, visit your repo's landing page and select "manage topics."