ClawArena: Benchmarking AI Agents in Evolving Information Environments
Published in arXiv preprint, 2026
ClawArena evaluates whether persistent AI agents can maintain correct beliefs as their information environments evolve. It exposes agents to noisy, partial, and sometimes contradictory evidence across multi-channel sessions, workspace files, and staged updates, then tests multi-source conflict reasoning, dynamic belief revision, and implicit personalization through multiple-choice and shell-based executable checks.

Recommended citation: Haonian Ji, Kaiwen Xiong, Siwei Han, Peng Xia, Shi Qiu, Yiyang Zhou, Jiaqi Liu, Jinlong Li, Bingzhou Li, Zeyu Zheng, Cihang Xie, Huaxiu Yao, "ClawArena: Benchmarking AI Agents in Evolving Information Environments," 2026. [Online]. Available: https://arxiv.org/abs/2604.04202
Download Paper | Download Bibtex
