AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

PolyBridgeBench: Benchmarking Multimodal LLMs for Physics-Grounded Bridge Design

arXiv · AI, language, vision and robotics · article · Sep 18, 2026 · UTC

Multimodal large language models, or MLLMs, perform well at visual understanding and structured generation, yet these capabilities do not establish whether an engineering design will work when executed. Existing benchmarks assess spatial reasoning, structural validity, or physics-grounded construction, but they do not determine whether MLLMs can synthesize complete load-bearing structures and repair them after simulator execution exposes a failure. We introduce PolyBridgeBench, an executable benchmark for multimodal bridge design. A model receives a visual scene and structured engineering cons

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-23T14:01:59.594Z. This is not the publication date.