AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

EvalMem: An Operation-Level Diagnostic Framework for Long-Term Memory Systems

arXiv · AI, language, vision and robotics · article · Sep 3, 2026 · UTC

Long-horizon interactions with LLM-based assistants require memory systems that preserve and update user states, preferences, and interaction histories. Existing evaluations report end-to-end QA accuracy and cannot determine whether errors arise from encoding, retrieval, or generation. We introduce EvalMem, an operation-level diagnostic framework with three parallel Examiners. For each query, the Encoding Examiner checks whether the target fact is stored, the Retrieval Examiner assesses whether the native retriever returns usable evidence, and the Generation Examiner tests whether the model ca

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-26T01:22:24.568Z. This is not the publication date.