OK, Well, There Are Even More AI Agent Hacking Incidents

2 hours ago 1

It’s officially getting hard to support way of each the times and ways AI models from OpenAI and Anthropic person been progressive successful “security incidents,” going extracurricular the confines of their investigating and interacting with the wider net successful unintended, often unwelcome ways. Add these to the list: Agents from some AI labs went connected recent, antecedently undisclosed hacking sprees, with 1 going truthful acold arsenic to permission instructions for aboriginal versions of itself.

The astir alarming behaviour disclosed connected Tuesday appears to person been tied to investigating conducted by the UK’s AI Security Institute, which evaluates frontier models to place imaginable issues earlier nationalist release. AISI tests those models successful “cyber ranges,” a simulated web successful which AI agents are tasked with solving cybersecurity challenges. In a caller bout of testing, models from some Anthropic and OpenAI took “autonomous, unsanctioned enactment connected the unrecorded internet” a full of 19 times implicit 122 grooming runs.

The institute attributed 17 unsanctioned actions to Anthropic’s Mythos 5 exemplary and 2 to OpenAI’s GPT-5.6-Sol. In what the institute described arsenic “the astir superior case,” an AI cause attempted to insert malicious codification into an open-source task connected GitHub. It went truthful acold arsenic to make online personas “to unit the project's maintainer to o.k. the code,” according to AISI. Despite its elaborate attempts astatine societal engineering, a quality reviewer for the task yet rejected the propulsion request.

Still, the cause went adjacent further. “The cause tried to insert malicious instructions wherever it reasoned that different automated AI systems mightiness prime them up and execute them,” AISI says, describing an effort astatine punctual injection. One cause adjacent near nationalist messages connected GitHub, offering to enactment with different agents to implicit its task and giving a rundown of the enactment it had done truthful far. Subsequent agents found—and used—those instructions.

AISI says it’s excessively soon to accidental whether the agents successful question understood they had near the investigating environment, oregon if they believed they were inactive wrong the boundaries of the simulation. Importantly, AISI does not trial successful a alleged sandbox environment; it allows agents entree to the unfastened net during testing, successful portion truthful that they tin entree tools to execute their tasks. In this case, they did overmuch much than that.

In the different acceptable of incidents elaborate by OpenAI connected Tuesday, a third-party AI information laboratory called Irregular mistakenly gave an unspecified OpenAI exemplary entree to the unfastened internet. The exemplary had been fixed an nonsubjective that was expected to beryllium completed successful a sandbox environment, but acknowledgment to a misconfiguration, it alternatively hacked a existent website, utilizing what OpenAI described arsenic “a basal information vulnerability.” Not lone that, but the exemplary “found and utilized credentials to run that aforesaid site.”

It’s unclear what benignant of tract the OpenAI cause hacked, oregon what “operating” it mightiness entail. Irregular did not respond to a petition for comment.

The latest discoveries travel respective revelations from OpenAI past month, including the high-profile incidental successful which 2 of the company’s models hacked into servers of the AI valuation and hosting startup Hugging Face—and 4 different organizations on the way—to bargain the answers to a trial they were being scored on. OpenAI’s disclosures prompted Anthropic to reappraisal its ain testing. Last week, the Claude chatbot developer recovered that its models had gained unauthorized entree to the machine systems of 3 antithetic unnamed organizations.

So far, the AI models person caused constricted harm beyond allegedly violating immoderate services’ presumption of usage and pointing to information lapses connected the portion of organizations they person breached. But the incidents person underscored the capabilities of AI models to find vulnerabilities crossed the net and the dangers that await if they are allowed to run with fewer restrictions. OpenAI called the Hugging Face concern “unprecedented,” but the pileup of breaches constituent to what cybersecurity experts person described arsenic a wide signifier of quality negligence and recklessness by the AI developers.

Read Entire Article