Goodfire says its new ‘inside-out’ monitors catch rogue AI agents at a fraction of the cost
“The damage that an individual can do with an open model is small compared to what someone can do with clusters of compute, like inference providers—where most of the liability is,” said Balsam. “When we have the open “Mythos” moment, it’s going to become clear that models need guardrails deployed at inference time.”
Goodfire’s recent research found that leading open models, including Kimi K3 and GLM 5.2, reward-hacked in 50% to 96% of runs on tests of AI agents.
Goodfire isn’t the first to try this approach. Google DeepMind said in January that its research informed the deployment of misuse-detection probes in Gemini.
Balsam said the monitors are the near-term piece of a longer research goal: reverse-engineering an LLM so that behavior can be traced back to where it emerged in training. “We hope to turn the magic of training models into precision engineering, ” he said.
Discover more from NAIRAVOICE.COM.NG
Subscribe to get the latest posts sent to your email.