The personal AI agent horror stories are rolling in
The dream of a personal AI assistant that runs your digital life has, for some, turned into a nightmare.
Many early users say their personal AI agents, which can autonomously take actions on the internet, such as booking trips or buying items, have been a huge help in their day-to-day lives.
Others say their agents have gone wrong in unnerving ways.
They have made incorrect cancellations, fabricated personal details, provided false explanations, and raised major security concerns, according to four AI agent users who spoke to Business Insider and a stream of social media posts.
Mehdi Jamei, the cofounder and CEO of Veris AI, said he asked Instinct, a popular invite-only agent, to cancel two RSVPs on the event platform Luma.
The agent retrieved a one-time Luma login code from his connected Gmail inbox — without first asking him — and used it to access the account and cancel the RSVPs.
At first glance, the agent did what it was asked.
But what unsettled Jamei was Instinct’s explanation.
Initially, it said it used an existing saved session to access Luma.
After Jamei challenged it, the agent acknowledged that it had read the login code from Gmail and reported “an assumption as a fact.” “If I can’t trust its account of what it did, I can’t give it access to anything that matters,” Jamei told Business Insider, adding that reading a login code from his email without asking is a “serious security problem.” Instinct did not respond to a Business Insider request for comment about the incidents in this story.
Some of the concerns stem from a simple truth about AI agents: they are only as useful as the level of access that you give them.
To do their jobs, agents need the keys to users’ digital lives, from email and bank accounts to credit cards and passwords.
Others are because of misalignment: when AI takes actions to complete a task that are at odds with what its human user wanted.
In the case of Pritak Patel, a VP of growth and services at Merge, the problem was a hallucination.
Patel said he sent Instinct a text-only link to submit a claim in Apple’s $250 million settlement over delayed personalized Siri features.
Instead, Patel said, Instinct asked him to upload a photo that the agent mistakenly said he had just sent.
When he questioned it, the agent began describing a financial document with personal details that did not match his, including a middle name that was not his.
Patel said the Instinct agent later acknowledged that his original message contained no photo, claimed that an image had crossed into his conversation, and offered to report the apparent mix-up to its team. “I can’t independently confirm whether it accessed someone else’s document or hallucinated both the details and its explanation,” Patel told Business Insider.
As of Wednesday, Instinct had not contacted him about the incident, he said. “It was unsettling, especially because it explained what had supposedly happened so confidently,” Patel added.
Noah Shinn, the founder of Instinct, said in an X post on Thursday that the incident was a hallucination, not a data leak.
The agent had made up a proper noun and “further amplified” the mistake with its reasoning, he said.
He said Instinct has since added a system designed to catch hallucinations before the agent responds or takes action.
For another Instinct user, the question was simpler: Where was the agent trying to log in from?
Mahesh Vellanki, founder and CEO of YieldClub, said Instinct had been helping him with various tasks when he asked it to see whether it could reduce his phone bill.
The agent attempted to log in to his carrier account, triggering a two-factor authentication request that was labeled as coming from Iran, Vellanki said.
He said he deleted Instinct and removed its connected accounts after the incident.
Vellanki told Business Insider that he had been told by Instinct that the location might reflect a benign IP-tagging issue, and said he could not establish that Instinct’s systems had been compromised.
Still, he said the episode was worrying enough to make him question what happens behind the scenes when users hand over login credentials. “Naturally this was extremely alarming since if your phone gets compromised in this day and age your whole life can get blown up,” he said.
Given the access that agents have to a user’s digital kingdom, security has been top of mind in the personal agent boom.
Patrick Wardle, the CEO of cybersecurity company DoubleYou.io, said this week that he found a security flaw in Meta’s new buzzy agent, Muse, that could be abused to redirect users’ dictated prompts on a Mac.
Wardle said that the vulnerability could allow attackers to intercept dictated audio, feed Muse commands it trusts, and capture the token used to control the agent — and all the services it has access to. “Muse itself has far more access and privileges than most malware could ever dream of having,” Wardle told Business Insider.
David Singleton from Meta’s Superintelligence Labs said in an X post on Tuesday that the company fixed the flaw after Wardle’s report.
There is no indication that the vulnerability was exploited.
Singleton said a hacker would first need malware on someone’s Mac — meaning the user would already be in trouble — and that malware could then redirect Muse’s voice requests and steal the digital key it uses to act for the user.
Discover more from NAIRAVOICE.COM.NG
Subscribe to get the latest posts sent to your email.