memory-bandwidth
Here are 37 public repositories matching this topic...
Provides a set of benchmarks that can be used to measure the memory bandwidth performance of CPU's
-
Updated
Apr 8, 2024 - C
A Fast DNN Accelerator Design Space Exploration Framework.
-
Updated
Aug 10, 2022 - Python
Modern Memory Bandwidth and Latency Benchmarks
-
Updated
Jun 10, 2026 - C
Demo code accompanying the talk "Implementing memory locality optimizations in OpenFOAM based code"
-
Updated
Apr 7, 2025 - C++
Main Memory Bandwidth Monitoring
-
Updated
Sep 17, 2018 - C++
General-purpose compile-time Expression Templates library for C++
-
Updated
Sep 29, 2025 - C++
Measure and visualize why LLM inference is slow: bottleneck analysis, model dissection, KV-cache, GEMM/GEMV, quantization, and memory-bound decoding.
-
Updated
Apr 30, 2026 - HTML
-
Updated
Feb 11, 2025 - Python
C++23 benchmarking framework with 6 profiler backends, CUDA GPU support, statistical regression detection, cross-compilation for 5 architectures, and CLI tools for analysis and visualization.
-
Updated
Jul 3, 2026 - C++
Intent-aware KV execution prototype for agentic long-context inference: semantic block selection, dynamic scoring, KV quantization modeling, speculative prefetch simulation, CPU references, and future Triton/CUDA kernels.
-
Updated
May 29, 2026 - Python
Reproducible Pascal GPU Unified Memory benchmark with Nsight and nvprof profiling
-
Updated
Feb 1, 2026 - Python
Console UI for watching Memory stats
-
Updated
May 7, 2026 - Rust
OpenCL benchmarking tool to measure host-device bandwidth and kernel global memory throughput across GPUs and CPUs.
-
Updated
Mar 16, 2026 - C
Python lab for exploring memory bandwidth, cache effects, and locality in accelerator workloads
-
Updated
Dec 24, 2025 - Python
Measures actual GPU costs for LLM prefill and decode on RTX 2070 (8GB) to validate simulation parameters. Key findings: prefill converges to 30-70 us/token at long sequences; decode is memory-bandwidth-bound (constant with prefix length, 5300-10800 us/token single-request); simulation defaults are correct for server-level amortized batching.
-
Updated
Jul 10, 2026 - Python
Bandwidth-focused LLaMA token-generation latency benchmarking on Apple Silicon M4 vs M2 with TTFT, PTL, KV-cache, quantization projections, and reproducible architecture analysis.
-
Updated
May 10, 2026 - Python
Single-binary interactive CLI harness to stress i.MX95 GPU/VPU (NPU later) and measure cross-block interference via DDR bandwidth
-
Updated
Jun 13, 2026 - C++
Improve this page
Add a description, image, and links to the memory-bandwidth topic page so that developers can more easily learn about it.
Add this topic to your repo
To associate your repository with the memory-bandwidth topic, visit your repo's landing page and select "manage topics."