AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

WebCraftBench: Evaluating Web Application Generation from a Software Testing Perspective

arXiv · AI, language, vision and robotics · article · Sep 14, 2026 · UTC

Human evaluation provides a direct measure of the quality of LLM-generated web applications. However, fitting human judgments through automated evaluation remains challenging. Static benchmarks can credit functionality that exists in source code but is unreachable at runtime. Interactive benchmarks exercise the application, yet incomplete exploration can cause them to miss implemented functionality and confound application defects with agent execution failures. To address these limitations, we propose WebCraftBench, an interactive benchmark for evaluating web application generation from a soft

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T11:41:07.830Z. This is not the publication date.

Observed changes

AIIC observation times, not verified publisher revision times. Up to eight recent revisions.

2026-09-24T08:32:17.684Z

  • title: IWC-Bench: Evaluating Web Application Generation from a Software Testing Perspective → WebCraftBench: Evaluating Web Application Generation from a Software Testing Perspective
  • summary: Human evaluation provides a direct measure of the quality of LLM-generated web applications. However, fitting human judgments through automated evaluation remains challenging. Static benchmarks can credit functionality that exists in source code but is unreachable at runtime. Interactive benchmarks exercise the application, yet incomplete exploration can cause them to miss implemented functionality and confound application defects with agent execution failures. To address these limitations, we propose IWC-Bench, an interactive benchmark for evaluating web application generation from a software → Human evaluation provides a direct measure of the quality of LLM-generated web applications. However, fitting human judgments through automated evaluation remains challenging. Static benchmarks can credit functionality that exists in source code but is unreachable at runtime. Interactive benchmarks exercise the application, yet incomplete exploration can cause them to miss implemented functionality and confound application defects with agent execution failures. To address these limitations, we propose WebCraftBench, an interactive benchmark for evaluating web application generation from a soft