This work operationalizes contextual security through four properties that must hold jointly and be evaluated continuously across the agent's trajectory, which changes which defenses are coherent, which evaluations measure something useful, and which attack patterns evaluation can see at all.
Vincent Siu, Jingxuan He, Kyle Montgomery et al.· arXiv.org· 1 citation
Whether an instruction-tuned model calls a tool can be controlled by a single linear direction in its residual stream, extracted without any training from the model's own tool-use preference signal and turned into an inference-time intervention with no prompt change.
Yu-Qiang Chen, Vincent Siu, Yang Liu et al.· 0 citations
CliniCARE-Bench is the first deployment-oriented clinical-agent benchmark to jointly evaluate real longitudinal EHR investigation, claim-level evidence grounding, governing-policy use, process adherence, and calibrated abstention within a common patient-level adjudication framework.
Veronica Chatrath, Bryan Zhu, George Pu et al.· 1 citation
ChainWorld, which composes atomic OSWorld tasks into long horizon desktop workloads through directional compatibility search while preserving the source evaluators, is studied, which contains 347 chains of length two to four and compares two renderings of the same task sequence.
Vincent Siu, Manasi Sharma, D. Song et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.