Precise analysis of multi-threaded programs requires combining flow-sensitive pointer analysis (FSPTA) with interleaving and lock analysis (ILA) to reason about cross-thread value flows under feasible concurrent executions. ILA computes may-happen-in-parallel (MHP) relations and lock-release spans to determine when sha...
Jia-Wei Yang, Xiao Cheng, Jia-Wei Wang et al.· 0 citations
BenchShield is presented, a model-backed instrumentation layer for reward integrity in LLM-agent evaluation that grounds detection in a finite lifecycle model of an evaluation's reward-relevant events and achieves 96% accuracy in detecting reward hacking from infrastructure-side evidence.
Sheng-Han Zheng, Zong-Lin Di, Yimin Liu et al.· 0 citations
LeanGuard is presented, a neuro-symbolic framework that assigns each act to the side equipped for it, and it is argued that the remedy is not better prompting but a separation of roles: the component that interprets the code must not also be the one that decides a safety obligation is met.
Yanjie Zhao, Hongjie Chen, Li Lu et al.· arXiv.org· 0 citations
The evaluation shows that AgentFlow recovers richer agent entities and dependencies than existing AST-based agent static analysis tools, generates more dependency-aware Agent BOMs, and uncovers 238 taint-style prompt-to-tool risks in real-world agent programs.
Shenao Wang, Xinyi Hou, Yanjie Zhao et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.