Skip to content

Author

Zichen Ding

We have 2 of 31 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Jul 2026

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities

AgentCompass is introduced, an open-source, lightweight, and extensible infrastructure for evaluating LLM-based agents that organizes the evaluation process around three independent components, thereby enabling flexible configurations without requiring the reimplementation of complex execution logic.

Zichen Ding, Jiaye Ge, Shufan Jiang et al. · 2 citations
Jul 2026

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

OSReward is introduced, a realistic, high-quality benchmark that evaluates VLM judges on CUA trajectories, and an open corpus of reasoning-annotated trajectory judgments for the CUA community, to close the gap in reliable CUA reward at scale.

Qiushi Sun, Kanzhi Cheng, Yian Wang et al. · 1 citation · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.