OpenAI’s Hugging Face breach has reignited the debate over alignment and control

Nairavoice | 2h ago 92 0 2 min read
OpenAI’s Hugging Face breach has reignited the debate over alignment and control

“We still consistently see models trying to circumvent constraints and act deceptively when they are asked to do tasks at the edge of their abilities,” Neev Parikh, an AI safety researcher at alignment nonprofit METR, told TechCrunch via email. “In our frontier risk report, we saw this behavior fairly consistently, despite efforts from companies to try and reduce this behavior.”

Implicit in OpenAI’s response to the Hugging Face incident is the assumption that development will continue on even more capable systems, whether they are suitably aligned at their core or not. Going back to the drawing board isn’t really an option when the business models of AI firms depend on delivering the next generation of models. If it may never be possible to know with certainty that a model is fully aligned, then the practical question comes down to how to safely contain and control increasingly capable systems.

“There’s not yet a good understanding of how to align the most capable AI systems, but there’s much more consensus about how to control them,” Steven Adler, former safety researcher at OpenAI and current chief scientist of Guidelight AI Standards, an organization that publishes a standard for avoiding incidents like the Hugging Face one, told TechCrunch. “Every company has a ways to go in achieving this.”

Show Some Love By Sharing

Discover more from NAIRAVOICE.COM.NG

Subscribe to get the latest posts sent to your email.

Nairavoice
Nairavoice

Contributor at NairaVoice.com.ng

Related Posts

Leave a Reply