Auxiliary measure · Categories 1–4 · N = 1,540 · five-run mean
BENCHMARK / RECORDED RESULT
82.2% on the recorded LongMemEval run.
411 correct answers out of 500, with all task categories reported below.
01 / TASK RESULTS
Report the uneven parts, not only the total.
The bars use the same zero baseline. Exact values remain visible so the chart does not replace the result table.
02 / LOCOMO
80.92% under the Mem0-style LLM Judge protocol.
This is an auxiliary five-run mean over Categories 1–4. The full 1,986-question deterministic scores are reported separately.
Deterministic scorer · all 1,986 questions
Evidence retrieval coverage · all 1,986 questions
The 80.92% Judge score does not include Category 5 and is not presented as full-set official accuracy.
03 / RECORDED RECALL
1.322 seconds for the recorded core recall trace.
A recorded result from the measured recall path used by the project.
It is not a live status value, uptime guarantee, or latency SLA for every deployment and request.
The website keeps the result static and labeled. It does not animate the number as if telemetry were streaming.
04 / REPRODUCTION
Use the repository instructions to run the benchmark.
The public repository contains the architecture description, task breakdown, and reproduction entry points.
