SOURCE-LINKED INTELLIGENCE
E-Commerce Bench: Long-Horizon Operations, Multi-Dimensional Evaluation
Agent benchmarks over the past few years have mostly followed one pattern. A goal is handed to the model, and the model tries to reach it within a bounded number of turns, whether that means finding the treasure in a maze, producing a report, or fixing a piece of code. Performance is then scored on the quality of the deliverable or on how much of the task got done, and evaluations of this kind usu
Read original source ↗ Open in workspace
- recordType
- page-entry
- evidenceStatus
- publisher-reported
- region
- Global
Evidence & attribution
- Qwen News · 2026-09-03T02:00:00.000Z
First collected: 2026-09-23T00:41:11.323Z. This is not the publication date.