AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

MCPGen: Benchmarking LLMs on Executable MCPWorkflow Development

arXiv · AI, language, vision and robotics · article · Sep 20, 2026 · UTC

We study whether LLMs can produce executable workflow artifacts that remain consistent across graph structure, tool implementation, schema bindings, and runtime wiring. In this setting, correctness depends on cross-layer consistency: a workflow may be structurally plausible, yet still fail because tool implementations, schema bindings, or runtime execution do not align. Existing benchmarks largely evaluate these capabilities in isolation or rely on trajectory-level proxies, leaving open whether generated workflow artifacts execute end-to-end. We introduce \textbf{MCPGen}, an executable benchma

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-23T09:51:33.063Z. This is not the publication date.