Current WizeMe public receipts
The 3 x 300-question LoCoMo quality lane reaches an 83.67% all-category mean and 87.30% core Categories 1-4 mean. Category 5 robustness averages 71.50%, and conservative answer p95 is 14.630 s. The faster lane records 83.33% all-category at 7.047 s. These seconds measure model generation and judging, not retrieval. LoCoMo low-latency retrieval records 18.777 ms cold p95; LongMemEval retrieval reaches 93.19% Any@3 with 11.827 ms cold p95. See the permanent Zenodo receipt.
Comparison contract
| Required match | Why it matters |
|---|---|
| Dataset and revision | Prevents stale or corrected data from being mixed. |
| Task and metric | Separates retrieval recall from end-to-end answer accuracy. |
| Timing boundary | Distinguishes embedding, retrieval, cache, and full response latency. |
| Hardware and run count | Makes performance claims reproducible and statistically useful. |
Benchmarks indexed
LoCoMo evaluates very long-term conversational memory. LongMemEval tests information extraction, multi-session reasoning, knowledge updates, temporal reasoning, and abstention. PersonaMem evaluates evolving user profiles and personalized responses.
Machine-readable research data
Results use a normalized JSON contract. Citation metadata is available through CITATION.cff and codemeta.json.
Launch proof boundary
The benchmark repository is ready for Product Hunt and research discovery, but the Product Hunt claim remains pending until the live launch URL and screenshot receipt are loaded. Provenance comes before promotion.
GitHub interaction paths
Researchers and launch visitors can star or watch the repository, inspect scheduled workflow artifacts, submit normalized benchmark results through the issue template, or open pull requests with reproducible provider adapters.