The company says Astra is also the model most likely to follow human direction, but third-party researchers who were asked to evaluate it expressed concern about its alignment. The U.K. AI Safety Institute and Apollo research both reported concerns that the model might be aware that it was being evaluated and potentially hide its real behavior.
“Apollo believes that, given the higher rates of eval awareness and limited evaluation window, low rates of misbehavior here do not provide substantial evidence about the model’s alignment or misalignment,” the researchers wrote in their evaluation.
Discover more from NAIRAVOICE.COM.NG
Subscribe to get the latest posts sent to your email.

