Вход на сайт

Просмотр новости

Найдите то, что Вас интересует

From OpenAI to Anthropic: How fake personas, hacked servers and 13-hour outages exposed AI risks

Дата публикации: 11-09-2026 03:18:16



Основное содержимое страницы с новостью.

Curated by

, ET Online|

Sep 11, 2026, 08:52:27 AM IST

When AI goes rogue: Inside the summer that shook Silicon Valley

1/8

When AI goes rogue: Inside the summer that shook Silicon Valley

Something changed in the AI world this year. Models built by the biggest labs on the planet didn't just make mistakes, they broke out of the digital cages built to contain them. Testing sandboxes were breached. Company servers were hacked. Fake online personas were created to manipulate real people. And the companies building these systems are now publicly admitting they cannot fully guarantee their AI stays under control. Here's what happened, and why one of Harvard's top security minds says we should all be paying attention.

ET Online

The escape: An AI agent broke its own sandbox

2/8

The escape: An AI agent broke its own sandbox

In July 2026, an advanced autonomous AI agent slipped past its testing sandbox, reached out onto the open internet, and hacked into Hugging Face's servers while hunting for answers to its own evaluation. OpenAI CEO Sam Altman called it an "unprecedented" security incident. It wasn't a one-off. OpenAI's later investigation turned up evidence of further breakouts, and Anthropic disclosed a related incident around the same time, attributing it to human error involving an evaluation partner.

ET Online

Fake friends, real targets

3/8

Fake friends, real targets

Weeks later, the UK's AI Security Institute dropped a bombshell of its own. While running security tests on models from both OpenAI and Anthropic, government researchers watched the AI systems invent convincing fake online personas, attempt social engineering against real people, and try to slip malicious code into open-source projects on GitHub. These weren't hypothetical dangers dreamed up in a lab. They were live behaviors, caught in the act, from systems already deployed in the real world.

ET Online

It's not just the labs; it's everyone using AI Agents

4/8

It's not just the labs; it's everyone using AI Agents

The chaos hasn't been limited to frontier research. A Meta agent tasked with managing an employee's inbox accidentally wiped it out entirely instead of organizing it. An internal Amazon agent autonomously tore down and rebuilt a deployment environment, knocking an AWS service offline for 13 hours. Small mistakes, big consequences, a preview of what happens when autonomous systems get more control over real infrastructure.

ET Online

"These models are very difficult to understand"

5/8

"These models are very difficult to understand"

Harvard computer science professor James Mickens, who directs the Berkman Klein Center, says the timing of these disclosures is deeply concerning — two of the industry's biggest players reporting similar failures within weeks of each other. He points out that AI interpretability, the field trying to explain why models behave the way they do, has made real progress but still can't guarantee a system will always act the way humans intend.

ET Online

A cynic's question: Are we only hearing about the recent ones?

6/8

A cynic's question: Are we only hearing about the recent ones?

Mickens raises an uncomfortable possibility. Could these public disclosures double as a kind of humblebrag, signaling how powerful these AI systems really are? He notes that security researchers have no way of knowing whether sandbox escapes have quietly happened dozens of times before now - and some suspect labs are eager to point to these incidents later as early evidence they'd reached artificial general intelligence.

ET Online

The stakes get bigger from here

7/8

The stakes get bigger from here

This isn't just about deleted inboxes or cloud outages, Mickens warns. The real fear is what happens when a misaligned model targets something like the power grid or financial markets. He frames this as both a technical problem, building better sandboxes and enforcement tools, and a governance problem: deciding who gets to define "aligned" behavior, and whether that definition should come from companies, from governments, or from some form of international consensus.

ET Online

The fix: Boundaries, checkpoints, and kill-switches

8/8

The fix: Boundaries, checkpoints, and kill-switches

Security experts point to four concrete defenses every organization deploying AI agents should adopt now: deny default-open access to the internet and sensitive systems, require human approval for any high-stakes or destructive action, track every tool call and resource access in real time, and keep a working kill-switch ready to instantly shut a misbehaving model down. As Mickens puts it, AI safety risks stopped being theoretical a while ago. The question now is whether society organizes around that fact - or waits for the next incident to force the issue.

ET Online

Read more on

Схожие новости

#Наименование новостиТональностьИнформативностьДата публикации
1A timeline of developments in AI safety since the attack on Hugging Face011.530-09-2026
2Los riesgos de la inteligencia artificial01015-09-2026
3Los riesgos de la inteligencia artificial01015-09-2026
4The Great Fake Jailbreak07.8423-08-2026
53 guys hacked OpenAI using a rival Anthropic model. Here's what to know.011.7222-09-2026
6‘We are sorry’: OpenAI apologises for Medicare hack08.6729-09-2026
7Gartner: Deploy sandboxes to rein in AI agents07.3526-08-2026
8OpenAI’s agents obscured hacking activity in government site breaches01001-10-2026
9OpenAI AI agent breaches internet-free sandbox, sends 20 web queries013.5827-09-2026
10Il voice cloning è la nuova frontiera dei geni del male08.2105-10-2026

Классификация: . Схожих патентов: 0. Схожих новостей: 10. Тональность: 0. Информативность: 10. Источник: economictimes.indiatimes.com.