OpenAI-HuggingFace: A Reproduction&Lessons for Alignment Testing
The misaligned AI behaviors that led to the OpenAI-Hugging Face incident are identified and it is shown that an auditing agent can elicit similar behaviors given high-level qualitative descriptions and that RL is a promising direction to do so.
Stewart Slocum, Malayandi Palan, Christopher G. Chute et al.
· 0 citations