Definition
Plain language
A small test set built from real moments where a person hinted at something they'd said before.
As stated in the literature
Benchmark of 72 reconstructed in-context memory cues from a four-month deployment, scored on Natural Integration and Reference rather than Direct QA.
Why it matters: It tests memory the way people actually use it — through hints and passing references — rather than through direct quizzing.
For example, one item might capture a user saying 'I'm heading back to that city again next month' and check whether the assistant connects it to the trip they described weeks earlier.
Heard on the show
“They pulled seventy-two of them where a user cued memory, reconstructed the exact context the live system had at that instant, and built a small benchmark called MemUse.”Episode 249 — The Chatbot Knows Your Facts And Still Won't Mention Them