AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

VibeMemBench: Evaluating Memory Systems for Coding Agents on Real Repository Coding Tasks

arXiv · AI, language, vision and robotics · article · Sep 20, 2026 · UTC

Coding agents operate on real repository coding tasks, and persistent memory systems promise to reuse experience across tasks. Yet existing evaluations do not show whether those systems improve executable repository work. Repository benchmarks test code changes but do not isolate memory, while memory benchmarks score recall without measuring downstream coding outcomes. We introduce VibeMemBench, a benchmark for evaluating memory systems on 111 coding targets from 90 SWE-rebench V2 repositories and 3,634 history trajectories from the target repositories. The targets follow the SWE benchmark sty

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-23T09:51:33.063Z. This is not the publication date.