Results show that CAi turns broad molecular-design intent into transparent, traceable workflows that connect interim decisions to candidate-level evidence, while exposing limits in long-horizon execution.
Zhuo-Lei Wang, Jiangyu Chen, Yingjun Shang et al.· 0 citations
This work shows that scenario-wrapped prompts activate internal scenario directions whose causal steering consistently reduces refusal scores, and proposes Concept2Scenario, a concept-based attribution framework for vulnerable scenario discovery that instantiates a broad concept space with a sparse autoencoder, translates the identified concepts into interpretable natural-language scenarios, and identifies synergistic scenario combinations through interaction attribution.
Experiments conducted with A^2E (Agent Auditing Engine) reveal that model-harness combinations exhibit substantial performance variation across different types of tasks, and that no single combination consistently outperforms all others across every task.
Haoning Wang, Mingxun Zhang, Chenyue Yu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.