memory-bandwidth
Here are 44 public repositories matching this topic...
Provides a set of benchmarks that can be used to measure the memory bandwidth performance of CPU's
-
Updated
Apr 8, 2024 - C
A Fast DNN Accelerator Design Space Exploration Framework.
-
Updated
Aug 10, 2022 - Python
Modern Memory Bandwidth and Latency Benchmarks
-
Updated
Aug 27, 2026 - C
Demo code accompanying the talk "Implementing memory locality optimizations in OpenFOAM based code"
-
Updated
Apr 7, 2025 - C++
Main Memory Bandwidth Monitoring
-
Updated
Sep 17, 2018 - C++
General-purpose compile-time Expression Templates library for C++
-
Updated
Sep 29, 2025 - C++
Pre-registered study: can MoE routers be trained for cache locality at no accuracy cost? A negative result at 137M scale, plus edge memory-bandwidth-wall measurements. Single RTX 3070.
-
Updated
Aug 20, 2026 - Python
Measure and visualize why LLM inference is slow: bottleneck analysis, model dissection, KV-cache, GEMM/GEMV, quantization, and memory-bound decoding.
-
Updated
Apr 30, 2026 - HTML
A CPU benchmark suite that shows its work. Six workloads, calibrated measurement windows, barrier-synchronised threads, robust statistics, full machine provenance, and Ed25519-signed results.
-
Updated
Aug 11, 2026 - Rust
C++23 benchmarking framework with 6 profiler backends, CUDA GPU support, statistical regression detection, cross-compilation for 5 architectures, and CLI tools for analysis and visualization.
-
Updated
Aug 30, 2026 - C++
The typed dimensioning and closed-form budget model for the Trellis-180M code-reasoning language model.
-
Updated
Sep 3, 2026 - Python
Console UI for watching Memory stats
-
Updated
May 7, 2026 - Rust
-
Updated
Feb 11, 2025 - Python
Intent-aware KV execution prototype for agentic long-context inference: semantic block selection, dynamic scoring, KV quantization modeling, speculative prefetch simulation, CPU references, and future Triton/CUDA kernels.
-
Updated
May 29, 2026 - Python
Reproducible Pascal GPU Unified Memory benchmark with Nsight and nvprof profiling
-
Updated
Feb 1, 2026 - Python
Is your llama.cpp decode memory-bandwidth-bound? Find out in one command.
-
Updated
Aug 27, 2026 - Python
Standalone C++17 SYCL benchmarks for arithmetic, joint-matrix, device-memory, and USM transfer throughput without external compute libraries.
-
Updated
Aug 6, 2026 - C++
Empirical LLM inference profiling on NVIDIA RTX A400 (Ampere SM86, 4GB): roofline ridge point, SDPA kernels, batching, CUDA Graphs, and WDDM PCIe paging telemetry.
-
Updated
Sep 1, 2026 - Python
Add this topic to your repo
To associate your repository with the memory-bandwidth topic, visit your repo's landing page and select "manage topics."