2024-04-03: New Jailbreaking Attack Exposes Risks in AI Language Models
Anthropic Discovers Long Prompts Are Confusing
🔷 Subscribe to get breakdowns of the most important developments in AI in your inbox every morning.
Here’s today at a glance:
🔓 New Jailbreaking Attack Exposes Risks in AI Language Models
Paper Title: Many-shot Jailbreaking
Who:
A large team of AI researchers from Anthropic, University of Toronto, Vector Institute, Constellation, Stanford, and Harvard
Led by Cem Anil of Anthropic along with academic advisors David Duvenaud and Roger Grosse
Why:
AI language models are being deployed in high-stakes domains, so it's crucial to identify risks before issues arise
Longer context windows in recent language models unlock new attack surfaces that need to be studied
Effective attacks that exploit long contexts could allow bad actors to elicit harmful outputs from AI systems
How:
Developed a new "Many-shot Jailbreaking" (MSJ) attack that uses hundreds of examples to steer model behavior
Tested MSJ on state-of-the…



