#natural language process...
Jun 2026
RealClawBench: Live OpenClaw Benchmarks from Real Developer-Agent Sessions
RealClawBench is introduced, a live benchmark framework built from real OpenClaw sessions to capture the distribution, diversity, and real-world difficulty of deployed agent use and provides a practical path toward benchmarks that better measure agent capability in actual use.
Zongwei Lv, Zhewen Tan, Yao-Ming Li et al.
· arXiv.org · 1 citation
· ⚡1