Data Science at Home
Data Science at Home
Francesco Gadaleta
Attacking LLMs for fun and profit (Ep. 239)
22 minutes Posted Sep 18, 2023 at 3:11 am.
0:00
22:14
Download MP3
Show notes

As a continuation of Episode 238, I explain some effective and fun attacks to conduct against LLMs. Such attacks are even more effective on models served locally, that are hardly controlled by human feedback.

Have great fun and learn them responsibly.

 

References

https://www.jailbreakchat.com/

https://www.reddit.com/r/ChatGPT/comments/10tevu1/new_jailbreak_proudly_unveiling_the_tried_and/

https://arxiv.org/abs/2305.13860