BENCHMARK / RECORDED RESULT

82.2% on the recorded LongMemEval run.

411 correct answers out of 500, with all task categories reported below.

01 / TASK RESULTS

Report the uneven parts, not only the total.

The bars use the same zero baseline. Exact values remain visible so the chart does not replace the result table.

Knowledge Update71 / 78
91.03%
Multi Session90 / 133
67.67%
Single Session · Assistant55 / 56
98.21%
Single Session · Preference27 / 30
90.00%
Single Session · User67 / 70
95.71%
Temporal Reasoning101 / 133
75.94%

02 / LOCOMO

80.92% under the Mem0-style LLM Judge protocol.

This is an auxiliary five-run mean over Categories 1–4. The full 1,986-question deterministic scores are reported separately.

MEM0-STYLE LLM JUDGE80.92%

Auxiliary measure · Categories 1–4 · N = 1,540 · five-run mean

OFFICIAL TOKEN F155.20

Deterministic scorer · all 1,986 questions

EVIDENCE RECALL82.00%

Evidence retrieval coverage · all 1,986 questions

The 80.92% Judge score does not include Category 5 and is not presented as full-set official accuracy.

03 / RECORDED RECALL

1.322 seconds for the recorded core recall trace.

WHAT IT IS

A recorded result from the measured recall path used by the project.

WHAT IT IS NOT

It is not a live status value, uptime guarantee, or latency SLA for every deployment and request.

HOW IT IS SHOWN

The website keeps the result static and labeled. It does not animate the number as if telemetry were streaming.

04 / REPRODUCTION

Use the repository instructions to run the benchmark.

The public repository contains the architecture description, task breakdown, and reproduction entry points.

GitHub ↗