ClawArena-Team: Benchmarking Subagent Orchestration and Dynamic Workflows in Language-Model Agents
Published in arXiv preprint, 2026
ClawArena-Team benchmarks the management ability of a language-model agent that creates specialized subagents, delegates work, and orchestrates their parallel, asynchronous returns through dynamic workflows. It comprises 41 multi-turn, multimodal, multi-directory scenarios with 258 evaluation rounds and 72 staged updates. Its execution-based Subagent-Management Score jointly measures task correctness, least-privilege access control, and modality routing.

Recommended citation: Kaiwen Xiong, Haonian Ji, Shi Qiu, Zeyu Zheng, Cihang Xie, Xinyu Ye, Huaxiu Yao, "ClawArena-Team: Benchmarking Subagent Orchestration and Dynamic Workflows in Language-Model Agents," 2026. [Online]. Available: https://arxiv.org/abs/2606.31174
Download Paper | Download Bibtex
