A representation-aware and frequency-aware evaluation framework that measures internal embedding drift, spectral sensitivity, and structural smoothness (spatial consistency of vision tokens), alongside standard label-based metrics is introduced.
Farooq Ahmad Wani, Alessandro Suglia, Rohit Saxena et al.· arXiv.org· 2 citations
GauntletBench, a web-based benchmark for evaluating agent generalisation in challenging scenarios, focusing on three underexplored capabilities (temporal perception, graphical understanding, and 3D reasoning), is introduced, revealing the substantial gap between current agent capabilities and those required for complex...
Mykola Vysotskyi, Runqi Lin, Grzegorz Biziel et al.· 0 citations
We propose the AIMO Interpretability Challenge, a competition on distinguishing robust from spurious reasoning in frontier mathematical language models based on the models'internal mechanisms. The challenge is motivated by a central limitation of standard reasoning benchmarks: strong final-answer accuracy does not reve...
Michal vStef'anik, Philipp Mondorf, Andreas Waldis et al.· 0 citations
It is shown that imbalanced learning of two conflicting copy tasks promotes in-context learning and improves the selectivity of refusal fine-tuning, suggesting that imbalanced pretraining curricula may be an important tool for promoting disentangled representations, with direct consequences for the precision and reliab...
Sebastian A. Bruijns, Jirko Rubruck, Mia Whitefield et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.