Do not write that jailbreak paper
Javier Rando
Abstract
Jailbreaks are becoming a new ImageNet competition instead of helping us better understand LLM security. This blogpost surveys the jailbreak literature to extract the most important contributions and encourages the community to revisit their choices and focus on research that can uncover new security vulnerabilities.
Video
Chat is not available.
Successful Page Load