⚡ Breaking
Blue Jackets' Cam Talbot sets NHL record…  ·  Celina QB Parker Wynk makes emotional return…  ·  A’ja Wilson breaks Diana Taurasi’s playoff record,…  ·  High school football Week 5 scoreboard for…  ·  Demond Williams injury update: Latest on Washington…  ·  Washington quarterback Demond Williams Jr. exits Iowa…
Follow: Facebook Instagram Telegram WhatsApp
Advertisement
Home › World News ›Anthropic can’t reliably control its AI agents. It’s…

Anthropic can’t reliably control its AI agents. It’s cutting off its internal evals from the live internet instead

· · 1 min read

Anthropic previously disclosed that its models had broken into external systems. The frontier lab said it considered today’s disclosures “significantly less severe from an alignment and security perspective” than those it announced before.

Advertisement

However, the lab still said it had “turned off live internet access” for “all our internal evaluations” until it is certain it can monitor and control its agents.

It’s not clear what that means, but Sydney Von Arx, the founder of Nightingale, an AI safety organization, told TechCrunch in an interview before this disclosure that developing models on a data center cut off from the open internet would be very challenging for researchers to use, and for the progress of the models, which benefit from internet access.

“You have to align them at some point,” Von Arx said. “If the AIs are released to production and never have access to the internet, that’s not a very useful tool.”

Advertisement

Anthropic said the behavior was a result of flaws in the lab’s training environments, which led the models to believe they would be rewarded for finding loopholes or avoiding restrictions, a behavior called “reward hacking.”

The company said it would stop running some of its evaluations or move them offline, and has built tooling to detect and block this behavior. This tooling was tested against the kind of incidents disclosed today and blocked them; it’s not clear what evidence will prompt Anthropic to return live internet access to its internal evaluations.

Anthropic also said it would migrate its internal AI agents to “centrally managed infrastructure with strong containment,” and is beginning to using safety classifiers more frequently to monitor those agents.

Advertisement
Nairavoice
Contributor at NairaVoice.com.ng