1.8K
Jess Miers 🦝
Jess Miers 🦝

I did a somewhat deep dive into the recent OpenAI and Anthropic "hacks" for an upcoming interview. IMO, the AI companies are better off running with the "rogue AI" story because the real story is actually pretty embarrassing. Like making your Wi-Fi password "wifi" embarrassing. 🧡 Most if not all of the current hacks can be chalked up to security engineers moving too fast and misconfiguring their own testing environments. Like agents accessing libraries that should have been sealed off, Internet access being enabled when it shouldn't have been... ...Hugging Face apparently leaving the default admin access privileges on across database clusters... This is, frankly, amateur hour kind of stuff for run of the mill cyber and infosec teams. Here's what I actually think these incidents reveal (instead of a rogue AI terminator situation): Much of the infra we expect to be hardened is actually being held together by glue and popsicle sticks. That's for a few reasons. Lazy engineering and oversight for starters. But also, this just happens all the time at big companies. Lower priority vulnerabilities get put on the back burner for higher priority ones. Priority is based on risk calculations, a big one being human hacker capabilities. AI can now detect, automate, and chain together exploits faster than any human hacker can. That is significant. Which means that the system vulnerabilities once considered not so low hanging fruit are actually very low hanging. In recent cases, it's not that AI somehow "escaped containment" or "went rogue." That's anthropomorphizing at best. It's more like if a Boston Dynamics robot dog nudged open an already open door and walked in. In one instance, a Claude bot apparently tried to abort its mission numerous times when the bot "realized" it met an unstable condition. But it didn't have a fail safe built in, so it proceeded by taking advantage of unsecured access credentials. Again, this is bush league. "Slowing down," then, doesn't necessarily mean literally slowing down AI production. It means making sure your testing environments are actually secure, that your agents have fail safes built in, and that admin access to critical systems isn't guarded by a plain text 4 letter password. The cyber and info sec communities are looking at the recent AI news quite differently than the rest of us. Some are going back into work and telling their engineers "hey, remember how we said we'd get to that patch later...let's actually prioritize that." While others are recognizing that AI labs are now operating software agents powerful enough that their internal research environments need to be engineered like hostile code execution platforms. Which is exactly the level of concern and seriousness this situation deserve. Meanwhile, the media and these wannabe whistleblowing ex-Anthropic / OpenAI script kiddies are running with "rogue AI" because, well, blade runner sells. Panic sells. And because the general public doesn't understand how modern software development works, we run with it too. And of course OpenAI and Anthropic are running with it too. Rogue AI narratives easily let them off the hook for their sloppy (and embarrassing) development decisions. While also making it look to their eager investors like they're inventing god... And it gives them a convenient narrative to sell regulators on their market capture plans (i.e. "regulate me daddy.") The reality is that just like every other industry AI has touched, AI is showing us that yes, even software engineers are a little lazy sometimes and make silly mistakes that would have been caught if someone just slowed down and wrote some fucking unit tests. It's really no different from all these big name law firms and attorneys getting caught filing hallucinated cases. It's us. It was always us. Sources: Anthropic: anthropic.com OpenAI: openai.com Hugging Face: huggingface.co Hacker News / Ycomb: news.ycombinator.com

An alignment assessment of recent cybersecurity incidents

www.anthropic.com

(For the sake of clarity, none of the companies literally made their Internet access passwords "wifi" as far as I know. I was being intentionally facetious. But they might as well have for some of these instances) This adds way more color: "it is essential for lawmakers, and the public, to understand that none of the reported incidents actually involved loss of control at any time, or arguably even β€œrogue agents,” or any kind of humanlike agency on the part of AI models." aiguide.substack.com

Misleading Metaphors, Real Risks

aiguide.substack.com

"Likewise, the future danger lies with humans who will use such models in ways to unintentionally or intentionally cause harm." (h/t @davidcrespo.bsky.social )

Share this Page