Skip to content
#

memory-benchmark

Here are 18 public repositories matching this topic...

Low-level command-line tool for measuring CPU and Metal GPU memory bandwidth, synthetic LLM decode and prefill memory traffic, cache and main-memory latency, access-pattern performance, TLB behavior, and two-thread cache-line handoff protocol latency on Apple Silicon Macs.

  • Updated Aug 29, 2026
  • C++

First-Person Agent Memory Bench. 10 Categories including fact recall, multi-hop links, temporal reasoning, fact overwrites, speaker traps, refusal, credibility, and agentic tool usage. 540K token / 60 session corpus, all in first person. Dynamic output-answer-key portion. Comprehensive report with visuals and miss breakdown.

  • Updated Aug 27, 2026
  • Python

The first open evaluation framework for AI continuity. 250 narrative tests, 1835 verification questions, 10 checkpoints. Benchmark for AI memory systems, stateful agents, and long-term context persistence. No LLM in the evaluation loop.

  • Updated Apr 21, 2026
  • Python

Add this topic to your repo

To associate your repository with the memory-benchmark topic, visit your repo's landing page and select "manage topics."

Learn more