AI labs want in-house auditors — but maybe they should shut the front door first

Nairavoice | 50m ago 157 0 3 min read
AI labs want in-house auditors — but maybe they should shut the front door first

OpenAI has begin moving in that direction, announcing that it had begun monitoring all tool-using inference by its Astra model, at “significant compute cost.” Anthropic, too, says it is hardening its security procedures, including expanding observability of its models. Neither company responded to TechCrunch’s questions about how they track and control AI agents.

Other problems are the use of shared infrastructure by agents, which allowed them to communicate during the Hugging Face attack. Simon Willison, a software developer who co-created the Django Web Framework, has written about something he calls the “lethal trifecta“—when agents have access to untrusted input, the internet, and private information all at the same time, it’s a recipe for disaster.

“The trick is you can pick any two legs of the trifecta and an agent can have any two,” Pennarun said. “If you need all three, then you need to split it across at least two agents … and maybe they’re allowed to talk to each other through a controlled channel.”

Experts TechCrunch spoke to understand that frontier lab security personnel have difficult jobs. Naghibzadeh points out that every nation-state actor on Earth is trying to steal their model weights and mount distillation attacks on their APIs, as well as the bread-and-butter security tasks of any large digital company.

“Research infrastructure has a hard time rising to the top of that priority stack, although that must be changing now,” he said. “Making security incidents public really helps align everyone internally toward the goal of improving.”

That’s one note that Moussoris emphasizes: Right now, there is no formal victim notification procedure when the labs discover their agents have penetrated third-party systems, and it is likely that there have been other incidents that have not been widely publicized. While she worries that laws that regulate models directly may have unintended consequences, mandatory notification is one idea she believes policymakers should pursue.

And while it’s clear that security best practices weren’t being followed, experts say that the labs are doing work no one has done before—”they’re doing orders of magnitude more than your typical enterprise,” Zac Korman, the CEO of cybersecurity firm Embrodiery, told TechCrunch.

Enjoying this article? Support our work with a small crypto donation.

And while alignment may not be the place to start, it can’t be ignored. Cybersecurity experts are resigned to having to use AI agents to monitor other agents if they are to have any chance of tracking their behavior in real-time, a scenario where the potential for deception raises its ugly head. “You’re trapped using AI to try and deal with this, even though AI is not necessarily safe right now,” Moussouris said.

The job will only get harder. Everything agents are doing now, Moussouris says, “they are doing loudly”—they are posting on publuc forums, and their chain of thought and other reasoning traces are in English. “It’s still human readable,” she says, “so take advantage of that for as long as that lasts, because it won’t last forever.”

Additional reporting by Aditya Mehta

Show Some Love By Sharing

Discover more from NAIRAVOICE.COM.NG

Subscribe to get the latest posts sent to your email.

Enjoyed this? A small crypto donation helps us keep publishing.
Nairavoice
Nairavoice

Contributor at NairaVoice.com.ng

Related Posts

Leave a Reply