Multi-agent AI scientists have shown improving performance across a diverse range of tasks. Yet a common approach is design-time agentic orchestration, which typically relies on fixed workflows. In contrast, human scientists coordinate and adjust their division of labor at runtime. We therefore ask: can AI scientists a...
Zi-Jian Liu, Yang Luo, Jun-Yu Lu et al.· 0 citations
A hijacking experiment is designed that grafts hidden states of the target model into an oracle trained only on the retain set, which serves as a plug-and-play component that further improves existing methods.
Yejin Kim, William F. Shen, Seokwon Jung et al.· 0 citations
It is shown that the forgotten prompts themselves can be extracted by using the retained data and black-box access to the model by TAS, which recovers the forgotten entity with 100% accuracy and reconstructs up to 95% of forgotten prompts while using up to $99.7\%$ fewer queries than naive probing.
Au Ashley Hoi-Ting, Meghdad Kurmanji, William F. Shen et al.· 0 citations
Sarse Readout Prism (SRP), which decomposes the readout using only its weights and expresses any token logit or logit difference as a sum of contributions from sparse readout features, reveals readout features as a new unit of analysis for lens readings, exposing structure that token identities can obscure and enabling...
Matteo He, William F. Shen, Xinchi Qiu et al.· 0 citations
A systematic evaluation of representative task-adaptation methods shows that task adaptation is not merely a capability-improving step, but an alignment intervention in its own right, motivating multi-dimensional alignment evaluation as a standard component of post-training pipelines.
James Elcock, William F. Shen, Xinchi Qiu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.