AgentCompass is introduced, an open-source, lightweight, and extensible infrastructure for evaluating LLM-based agents that organizes the evaluation process around three independent components, thereby enabling flexible configurations without requiring the reimplementation of complex execution logic.
Evaluations across scientific, multimodal, agentic, and general-purpose benchmarks show that Intern-S2-Preview-397B achieves competitive or leading results in multiple settings.
Lei Bai, Jiaqi Cao, Chiyu Chen et al.· 2 citations
These results establish source-intent preservation and shared executable-state grounding as key principles for scalable terminal-task synthesis as key principles for scalable terminal-task synthesis.
K. Shi, Zun Wang, Qi-Sheng Su et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.