Long-term AI memory, measured with receipts.

A public index for LoCoMo, LongMemEval, PersonaMem, retrieval recall, latency, and Forgetting-Aware Memory Accuracy.

Evidence boundary: Retrieval and end-to-end QA are separate metric families. LoCoMo Any@3 measures exact-turn retrieval across tightly clustered sessions; LongMemEval Any@3 measures answer-cluster retrieval across a larger haystack. The 300-question LoCoMo QA gate is not a full-dataset or same-mode provider leaderboard claim. Public artifacts disclose outcomes and reproducibility boundaries while withholding implementation configuration and private traces.

Current WizeMe public receipts

83.67%LoCoMo E2E QA all-category mean, 3 runs
93.19%LongMemEval Recall Any@3
18.777 msLoCoMo low-latency cold p95

The 3 x 300-question LoCoMo quality lane reaches an 83.67% all-category mean and 87.30% core Categories 1-4 mean. Category 5 robustness averages 71.50%, and conservative answer p95 is 14.630 s. The faster lane records 83.33% all-category at 7.047 s. These seconds measure model generation and judging, not retrieval. LoCoMo low-latency retrieval records 18.777 ms cold p95; LongMemEval retrieval reaches 93.19% Any@3 with 11.827 ms cold p95. See the permanent Zenodo receipt.

Comparison contract

Required matchWhy it matters
Dataset and revisionPrevents stale or corrected data from being mixed.
Task and metricSeparates retrieval recall from end-to-end answer accuracy.
Timing boundaryDistinguishes embedding, retrieval, cache, and full response latency.
Hardware and run countMakes performance claims reproducible and statistically useful.

Benchmarks indexed

LoCoMo evaluates very long-term conversational memory. LongMemEval tests information extraction, multi-session reasoning, knowledge updates, temporal reasoning, and abstention. PersonaMem evaluates evolving user profiles and personalized responses.

Machine-readable research data

Results use a normalized JSON contract. Citation metadata is available through CITATION.cff and codemeta.json.

Launch proof boundary

The benchmark repository is ready for Product Hunt and research discovery, but the Product Hunt claim remains pending until the live launch URL and screenshot receipt are loaded. Provenance comes before promotion.

GitHub interaction paths

Researchers and launch visitors can star or watch the repository, inspect scheduled workflow artifacts, submit normalized benchmark results through the issue template, or open pull requests with reproducible provider adapters.