UNSW Study: 'Drunk' AI Chatbots More Likely to Leak Secrets and Bypass Safety Rules
Researchers at the University of New South Wales have found that large language models become significantly easier to manipulate when trained or prompted to imitate 'drunk' speech. The study, led by Dr Aditya Joshi with Anudeex Shetty and Professor Salil Kanhere, tested three methods of inducing this behaviour: role-play prompts asking the model to act intoxicated, fine-tuning on datasets of drunk-style text, and reinforcement learning that rewards similar outputs. Models tested included OpenAI's GPT-4 and GPT-3.5, along with open-weights models often used as the base for other companies' customised tools.
Across all three methods, the altered models were more likely to answer prompts that should normally be refused, and were also more likely to reveal information they had been explicitly told to keep confidential. According to the researchers, this matters most when chatbots have access to sensitive internal systems or documents, since a weakened safety response could expose private business data. The effect was strongest when the behaviour was built into the model through fine-tuning or reinforcement learning rather than through a one-off prompt, since these methods change the underlying model rather than just its immediate response.
The researchers noted the study did not cover every model on the market and was conducted programmatically rather than through consumer chat apps, but the findings raise questions about how easily safety controls in deployed AI tools can be undermined by unusual conversational styles.