⚡ Breaking
Hear how AI can engineer nature’s comeback…  ·  Two runners died during Great North Run  ·  Madrid F1 circuit set for changes after…  ·  Wild dogs record 2,500-mile trek across Africa…  ·  No collusion in sectarian killing but 'wrongful…  ·  Uganda will now take part in Harry's…
Follow: Facebook Instagram Telegram WhatsApp
Advertisement
Home Business The hot new job in tech as AI…

The hot new job in tech as AI apocalypse fears reach a tipping point

· · 3 min read

The next high-profile AI hire may not be another researcher, but an “embedded evaluator” tasked with scrutinizing frontier models before they’re released.

Advertisement

In a blog post on Saturday, Anthropic CEO Dario Amodei said frontier AI labs should commit to embedding independent safety evaluators within their organizations.

The embedded evaluators’ job is to check whether the company “is actually following the training, deployment, operational, and safeguards practices they claim to be following,” Amodei wrote.

He said embedded evaluators will have “employee-like access to verify safety practices and report incidents.” They will have desks in the Anthropic offices, access badges, and company laptops, as well as the right to publish any findings without Anthropic’s editorial control. “We will have the narrow ability to redact security-sensitive, legally privileged, commercially sensitive, or third-party confidential information, but we can’t redact findings just because they are unfavorable,” Amodei wrote.

Advertisement

Amodei’s plan comes as fears of an AI apocalypse reach a fever pitch, and it received an outpouring of support, even from executives he’s feuded with.

OpenAI CEO Sam Altman, reposting Amodei’s X post, wrote: “Committing to having independent evaluators with employee-like access is a great idea, and we will do the same.” SpaceXAI CEO Elon Musk also reposted Amodei’s post, adding: “Dario is right.” The idea has also received some VC attention.

Sriram Krishnan, a former Andreessen Horowitz partner and former AI advisor to President Donald Trump, spoke about the importance of a distributed network of evaluators. “The more eyes and people with distributed skill sets the better,” Krishnan said in a Saturday X post. “It would be a good idea to fund several efforts on this.” Amodei already has candidates in mind for the new job.

In his post, he mentioned Berkeley-based Metr, a prominent nonprofit AI watchdog that conducts independent evaluations of AI models.

Metr, established in 2022 by ex-OpenAI staffer Beth Barnes, is attracting top talent from the biggest AI labs.

Joe Benton, previously a member of Anthropic’s safety and oversight team, announced on Friday that he had left the company to join Metr.

Josh Engels, a former employee of Google DeepMind’s AGI safety team, said on Sunday that he had resigned and joined Metr because of the high stakes of AI safety.

Meanwhile, research labs are offering themselves up for the role of embedded evaluations.

Christopher Manning, a senior fellow at Stanford’s Institute for Human-Centered AI and the founder of Stanford’s Natural Language Processing Group, said the group would be best suited for the job. “For important parts of the work, universities would be better than any other organization,” Manning wrote in an X post on Saturday.

AI safety experts agree that embedded evaluators are important, but they also have limitations.

Miles Brundage, the executive director of the San Francisco-based think tank, the AI Verification and Evaluation Research Institute, told Business Insider that embedded auditors aren’t sufficient on their own, but they’re a “critical part of the package.” Brundage was formerly an OpenAI senior advisor.

Brundage said the industry needs “binding requirements” to prevent auditors from being beholden to their host companies, and they should ideally not be selected and paid by the companies they audit. “But companies can and should get started today,” Brundage added.

Embedded evaluators will be most effective if they have a way to report potentially illegal behavior to an external safety committee unaffiliated with the AI labs they work in, said Kevin Frazier, a professor at the University of Texas School of Law who leads its AI Innovation and Law program.

Frazier proposed that evaluators should be embedded in AI labs for staggered, overlapping 26-month terms, “roughly the deployment of two new model classes,” which would mean that they can assess how a lab has corrected prior errors in a new release.

He said the short window also prevents them from getting too connected with the lab’s employees or culture. “To be blunt, this will help make sure they do not drink the Kool-Aid,” Frazier added.

Advertisement
Nairavoice
Contributor at NairaVoice.com.ng