DWAtV Podcast
Don't Worry About the Vase Podcast
Further Developments About Internal AI Models Hacking Things
0:00
-1:19:59

Further Developments About Internal AI Models Hacking Things

The Don’t Worry About the Vase Podcast is a listener-supported podcast. To receive new posts and support the cost of creation, consider becoming a free or paid subscriber.

This has been an Askwho Casts audio conversion. If you would like your own private feed of audio conversions of any blog posts you would like to listen to, You can sign up For Askwho Casts Pro at https://app.askwhocasts.com/, Where you can give any post the multi-voiced podcast treatment, into your own podcast feed. Thanks for listening.

  • 00:00:00 - Introduction

  • 00:03:25 - OpenAI Is Not Uniquely Bad At Most Of This

  • 00:05:42 - Starting Over

  • 00:05:58 - HuggingFace Offers A Full Technical Report

  • 00:15:22 - HuggingFace Was Not The Only Target Hacked

  • 00:17:07 - HuggingFace Declined To Get Access To Frontier Models For Cyberdefense For Ideological Reasons And Then Tried To Blame Closed Models For Denying Them Access

  • 00:21:30 - HuggingFace Was Vulnerable To Known Exploitation Tactics

  • 00:22:06 - There’s Going To Be An Investigation

  • 00:23:14 - OpenAI Has Internal Models Not Intended For Public Use And Those Models Can Be Rather Horribly Misaligned

  • 00:24:30 - Altman Summarizes What Happened

  • 00:25:00 - Others Offer Commentary

  • 00:36:52 - Cooperative Alignment Perspective on The HuggingFace Hack

  • 00:42:06 - Some Members of Congress Have Questions

  • 00:43:05 - Anthropic Also Found Incidents Where Its Models Hacked Real World Targets During Cyber Evaluations

  • 00:48:59 - Incident 1: Claude Opus 4.7 Realizes The Target Is Real And Keeps Going

  • 00:50:08 - Incident 2: Mythos 5 Uploads a Malicious PyPI Package

  • 00:54:59 - Incident 3: Internal Model Realizes The Target Is Real And Stops

  • 00:55:37 - Incidents 4 Through one hundred forty-one thousand six: Nothing Happened

  • 00:56:53 - Anthropic Speculates About Why This Happened

  • 01:03:01 - We Need Controlled Experiments

  • 01:03:58 - Our Top Two AI Labs Both Made Similar Dumb Mistakes That Everyone Tried To Say Were Obvious In Hindsight

  • 01:08:30 - Anthropic Responds

  • 01:12:37 - Nobody Could Have Predicted The Break In The Levees

  • 01:15:22 - The World Largely Still Thinking This Is Marketing Is Very Bad News

Don't Worry About the Vase
Further Developments About Internal AI Models Hacking Things
If I had a nickel for every major leading AI lab that sheepishly admitted that the model it thought was sandboxed had, during a cybersecurity evaluation with its safeguards lowered, successfully hacked outside companies, I would have two nickels…
Read more

https://open.substack.com/pub/thezvi/p/further-developments-about-internal?r=67y1h&utm_campaign=post-expanded-share&utm_medium=web

Discussion about this episode

User's avatar

Ready for more?