Skip to content

Author

Yelyzaveta Husieva

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#natural language process... Preprint Sep 2026

The Geometry of Harmfulness in Multi-Turn Attacks

Large language models (LLMs) remain vulnerable to adversarial attacks that circumvent safety alignment to elicit harmful outputs. It remains unclear how harmfulness and refusal representations evolve over the course of multi-turn attacks, and why single-turn defenses are less effective in multi-turn settings. This work...

Yelyzaveta Husieva, Lauren Alvarez · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.