One Turn Too Late: Learning When to Intervene Against Multi-Turn Malicious Intent
Experiments show that turn-level boundary supervision improves intervention localization, while reinforcement learning further improves the safety--utility trade-off.
2 papers indexed here
We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.
Not the right person? Other researchers publish under this name.
Experiments show that turn-level boundary supervision improves intervention localization, while reinforcement learning further improves the safety--utility trade-off.
Metadata Augmented Private Language Evolution (MAPLE) extracts DP tabular metadata and uses in-context learning to firmly ground the initial synthetic distribution in the target domain, and yields a strictly better privacy-utility trade-off.
We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.