Cogniscendo

Cogniscendo

2024-04-03: New Jailbreaking Attack Exposes Risks in AI Language Models

Anthropic Discovers Long Prompts Are Confusing

Prakash's avatar
Prakash
Apr 03, 2024
∙ Paid
🔷 Subscribe to get breakdowns of the most important developments in AI in your inbox every morning.

Here’s today at a glance:

  • New Jailbreaking Attack Exposes Risks in AI Language Models

  • AI artwork of the day

🔓 New Jailbreaking Attack Exposes Risks in AI Language Models

Paper Title: Many-shot Jailbreaking

Who:

  • A large team of AI researchers from Anthropic, University of Toronto, Vector Institute, Constellation, Stanford, and Harvard

  • Led by Cem Anil of Anthropic along with academic advisors David Duvenaud and Roger Grosse

Why:

  • AI language models are being deployed in high-stakes domains, so it's crucial to identify risks before issues arise

  • Longer context windows in recent language models unlock new attack surfaces that need to be studied

  • Effective attacks that exploit long contexts could allow bad actors to elicit harmful outputs from AI systems

How:

  • Developed a new "Many-shot Jailbreaking" (MSJ) attack that uses hundreds of examples to steer model behavior

  • Tested MSJ on state-of-the…

User's avatar

Continue reading this post for free, courtesy of Prakash.

Or purchase a paid subscription.
© 2026 8teapi · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture