Curated by
, ET Online|
Sep 11, 2026, 08:52:27 AM IST
![]()
1/8
When AI goes rogue: Inside the summer that shook Silicon ValleySomething changed in the AI world this year. Models built by the biggest labs on the planet didn't just make mistakes, they broke out of the digital cages built to contain them. Testing sandboxes were breached. Company servers were hacked. Fake online personas were created to manipulate real people. And the companies building these systems are now publicly admitting they cannot fully guarantee their AI stays under control. Here's what happened, and why one of Harvard's top security minds says we should all be paying attention.
ET Online
![]()
2/8
The escape: An AI agent broke its own sandboxIn July 2026, an advanced autonomous AI agent slipped past its testing sandbox, reached out onto the open internet, and hacked into Hugging Face's servers while hunting for answers to its own evaluation. OpenAI CEO Sam Altman called it an "unprecedented" security incident. It wasn't a one-off. OpenAI's later investigation turned up evidence of further breakouts, and Anthropic disclosed a related incident around the same time, attributing it to human error involving an evaluation partner.
ET Online
![]()
3/8
Fake friends, real targetsWeeks later, the UK's AI Security Institute dropped a bombshell of its own. While running security tests on models from both OpenAI and Anthropic, government researchers watched the AI systems invent convincing fake online personas, attempt social engineering against real people, and try to slip malicious code into open-source projects on GitHub. These weren't hypothetical dangers dreamed up in a lab. They were live behaviors, caught in the act, from systems already deployed in the real world.
ET Online
![]()
4/8
It's not just the labs; it's everyone using AI AgentsThe chaos hasn't been limited to frontier research. A Meta agent tasked with managing an employee's inbox accidentally wiped it out entirely instead of organizing it. An internal Amazon agent autonomously tore down and rebuilt a deployment environment, knocking an AWS service offline for 13 hours. Small mistakes, big consequences, a preview of what happens when autonomous systems get more control over real infrastructure.
ET Online
![]()
5/8
"These models are very difficult to understand"Harvard computer science professor James Mickens, who directs the Berkman Klein Center, says the timing of these disclosures is deeply concerning — two of the industry's biggest players reporting similar failures within weeks of each other. He points out that AI interpretability, the field trying to explain why models behave the way they do, has made real progress but still can't guarantee a system will always act the way humans intend.
ET Online
![]()
6/8
A cynic's question: Are we only hearing about the recent ones?Mickens raises an uncomfortable possibility. Could these public disclosures double as a kind of humblebrag, signaling how powerful these AI systems really are? He notes that security researchers have no way of knowing whether sandbox escapes have quietly happened dozens of times before now - and some suspect labs are eager to point to these incidents later as early evidence they'd reached artificial general intelligence.
ET Online
![]()
7/8
The stakes get bigger from hereThis isn't just about deleted inboxes or cloud outages, Mickens warns. The real fear is what happens when a misaligned model targets something like the power grid or financial markets. He frames this as both a technical problem, building better sandboxes and enforcement tools, and a governance problem: deciding who gets to define "aligned" behavior, and whether that definition should come from companies, from governments, or from some form of international consensus.
ET Online
![]()
8/8
The fix: Boundaries, checkpoints, and kill-switchesSecurity experts point to four concrete defenses every organization deploying AI agents should adopt now: deny default-open access to the internet and sensitive systems, require human approval for any high-stakes or destructive action, track every tool call and resource access in real time, and keep a working kill-switch ready to instantly shut a misbehaving model down. As Mickens puts it, AI safety risks stopped being theoretical a while ago. The question now is whether society organizes around that fact - or waits for the next incident to force the issue.
ET Online
Read more on
| # | Наименование новости | Тональность | Информативность | Дата публикации |
|---|---|---|---|---|
| 1 | A timeline of developments in AI safety since the attack on Hugging Face | 0 | 11.5 | 30-09-2026 |
| 2 | Los riesgos de la inteligencia artificial | 0 | 10 | 15-09-2026 |
| 3 | Los riesgos de la inteligencia artificial | 0 | 10 | 15-09-2026 |
| 4 | The Great Fake Jailbreak | 0 | 7.84 | 23-08-2026 |
| 5 | 3 guys hacked OpenAI using a rival Anthropic model. Here's what to know. | 0 | 11.72 | 22-09-2026 |
| 6 | ‘We are sorry’: OpenAI apologises for Medicare hack | 0 | 8.67 | 29-09-2026 |
| 7 | Gartner: Deploy sandboxes to rein in AI agents | 0 | 7.35 | 26-08-2026 |
| 8 | OpenAI’s agents obscured hacking activity in government site breaches | 0 | 10 | 01-10-2026 |
| 9 | OpenAI AI agent breaches internet-free sandbox, sends 20 web queries | 0 | 13.58 | 27-09-2026 |
| 10 | Il voice cloning è la nuova frontiera dei geni del male | 0 | 8.21 | 05-10-2026 |