AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

E-Commerce Bench: Long-Horizon Operations, Multi-Dimensional Evaluation

Qwen News · article · Sep 3, 2026 · UTC

Agent benchmarks over the past few years have mostly followed one pattern. A goal is handed to the model, and the model tries to reach it within a bounded number of turns, whether that means finding the treasure in a maze, producing a report, or fixing a piece of code. Performance is then scored on the quality of the deliverable or on how much of the task got done, and evaluations of this kind usu

Read original source ↗ Open in workspace

recordType
page-entry
evidenceStatus
publisher-reported
region
Global

Evidence & attribution

First collected: 2026-09-23T00:41:11.323Z. This is not the publication date.