Skip to content

Author

J. Wedgwood

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Review Sep 2026

Safety Nudges: User-Facing Interventions for Real-Time AI Risk Awareness

Safety Nudges is introduced, a browser-based tool that provides lightweight, in situ flags when concerning behavior is detected in chatbot conversations, suggesting that user facing safety nudges can complement model-level safeguards by helping people critically evaluate AI responses in context.

Varshini Elangovan, J. Wedgwood, Chhavi Yadav et al. · 0 citations

DSPA: Dynamic SAE Steering for Data-Efficient Preference Alignment

This work proposes Dynamic SAE Steering for Preference Alignment (DSPA), an inference-time method that makes sparse autoencoder (SAE) steering prompt-conditional, and audits the SAE features DSPA modifies, finding that preference directions are dominated by discourse and stylistic signals.

J. Wedgwood, Aashiq Muhamed, Mona T. Diab et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.