AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

GameReplica: A Benchmark for Black-Box Visual Game Replication by Vision-Language Agents

arXiv · AI, language, vision and robotics · article · Sep 15, 2026 · UTC

Coding-agent benchmarks usually evaluate implementation after the target behavior has been specified in text, code, or demonstrations. Existing research has extensively evaluated the ability of coding agents to generate programs from textual specifications. However, under black-box conditions where neither source code nor documentation is available, it remains underexplored whether an agent can induce the rules solely through visual observation and active interaction and reproduce the target system as a verifiable executable system. To this end, we present GameReplica, a closed-loop evaluation

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-23T23:32:26.314Z. This is not the publication date.