Definition
Plain language
A benchmark for testing how well AI assistants remember information across long conversations.
As stated in the literature
A long-context memory evaluation suite testing retrieval and recall of user-provided facts across extended interaction histories; cited as part of the fact-recall paradigm that staleness-detection work argues is the easier half of agent memory.
Why it matters: It gauges whether long-term assistants can actually hold onto user facts, a basic requirement for feeling personal and reliable over time.
For example, it checks whether an assistant still remembers a user's stated dietary restriction many sessions after they mentioned it.
Heard on the show
“There are big benchmarks for this — LoCoMo, LongMemEval.”Episode 031 — When Your AI Assistant Won't Let Go of Old Facts About You