Skip to content

Add LongMemEval-V2 benchmarking to MemoryBench - #55

Open
ved015 wants to merge 1 commit into
mainfrom
feat/longmemeval-v2-memorybench
Open

Add LongMemEval-V2 benchmarking to MemoryBench#55
ved015 wants to merge 1 commit into
mainfrom
feat/longmemeval-v2-memorybench

Conversation

@ved015

@ved015 ved015 commented Jul 28, 2026

Copy link
Copy Markdown
Member

Add a build-aware LongMemEval-V2 benchmark with pinned dataset preparation, bounded resumable ingestion, immutable build/query/reader/evaluator fingerprints, multimodal screenshot retrieval, and official evaluation/reporting.

Expose safe CLI and UI launch controls, reusable memory-build inspection, per-question evidence and screenshot inspection, configurable OpenAI reader/evaluator models, and the advanced Supermemory adapter. Add migration documentation, third-party notices, and comprehensive tests.

Validation: bun test (139 pass, 0 fail); root and UI TypeScript checks; production UI build; Prettier check for all 80 changed TypeScript and TSX files.

Live UI validation: one 100-trajectory enterprise haystack, all 211 questions completed, GPT-5 high reader, GPT-5.2 medium evaluator, and TopK 20.

Dataset, screenshots, checkpoints, and run reports remain local under ignored data directories and are not included in this commit.

Add a build-aware LongMemEval-V2 benchmark with pinned dataset preparation, bounded resumable ingestion, immutable build/query/reader/evaluator fingerprints, multimodal screenshot retrieval, and official evaluation/reporting.

Expose safe CLI and UI launch controls, reusable memory-build inspection, per-question evidence and screenshot inspection, configurable OpenAI reader/evaluator models, and the advanced Supermemory adapter. Add migration documentation, third-party notices, and comprehensive tests.

Validation: bun test (139 pass, 0 fail); root and UI TypeScript checks; production UI build; Prettier check for all 80 changed TypeScript and TSX files.

Live UI validation: one 100-trajectory enterprise haystack, all 211 questions completed, GPT-5 high reader, GPT-5.2 medium evaluator, and TopK 20.

Dataset, screenshots, checkpoints, and run reports remain local under ignored data directories and are not included in this commit.

ved015 commented Jul 28, 2026

Copy link
Copy Markdown
Member Author

This stack of pull requests is managed by Graphite. Learn more about stacking.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant