Skip to content

Author

Roy Betser

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Training-Free Policy Violation Detection via Activation-Space Whitening in LLMs

This work proposes a training-free method that operates directly on the LLM internal representations, leveraging prior evidence that decision-relevant information is encoded within them, and derives policy-violation scores directly from normalized representations of LLM hidden activations.

Oren Rachmil, Roy Betser, Itay Gershon et al. · 5 citations · ⚡1

AgenTRIM: Tool Risk Mitigation for Agentic AI

AgenTRIM is introduced, a framework for detecting and mitigating tool-driven agency risks without altering an agent's internal reasoning that provides a practical, capability-preserving approach to safer tool use in LLM-based agents.

Roy Betser, Amit Giloni, Shamik Bose et al. · 14 citations · ⚡2

RIFT-Bench: Dynamic Red-teaming For Agentic AI Systems

RIFT-Bench, a representation-driven methodology for dynamic red-teaming that enables unified evaluations across diverse agentic architectures, is introduced, showing that the approach generalizes effectively to heterogeneous agentic architectures.

Y. Levi, Roy Betser, Amit Giloni et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.