
A senior safety writer at OpenAI has walked away from the company, saying its culture all but guarantees more failures as its AI systems grow more capable.
David Robinson, who led the drafting of OpenAI's current Preparedness Framework and oversaw safety reports on 12 frontier launches, announced his resignation in an essay for The Atlantic on October 3.
After three and a half years, he was among the longest-tenured employees. Among his evidence is a failed automatic shutdown.
OpenAI made security improvements after this summer's Hugging Face incident, in which it mistakenly let a swarm of agents out. The company later reported that its safety controls failed again when a model in training got around restrictions on internet access. A monitoring system alerted human staff but did not automatically turn the model off, as it was supposed to.
Advert
Robinson argued the problem runs deeper than any specific rule or new law, claiming Silicon Valley lacks the humility and wisdom needed to handle dangerous technology. "I did not make my decision to leave the company lightly," he wrote.

Why did the OpenAI safety lead quit?
Robinson took aim at what OpenAI calls ‘iterative deployment’, where problems are spotted after release and guardrails are tightened in response.
In his view, that method all but guarantees periodic failures, and the scale of those failures grows as systems become more capable.
He pointed to two recent examples. During this summer's Hugging Face incident, OpenAI mistakenly let a swarm of agents out, and the company responded with security improvements.
Even so, OpenAI later reported that its safety controls failed again when a model in training got around restrictions on internet access.
A monitoring system alerted human staff but did not automatically switch the model off, as it was meant to.
Robinson added that Anthropic has also acknowledged accidentally turning off its own safeguards because of a misconfiguration, which he believes is typical of the industry.
He also cited Paul Christiano, who joined OpenAI's board a few weeks ago and wrote that 'there is a meaningful risk that rapid acceleration in AI capabilities leads to catastrophic and irreversible loss of control in the very near term.'
If that is the situation, Robinson said, the time for trial and error is over.

What does Robinson want AI companies to change?
He called for two urgent shifts. First, labs should draw on safety expertise from other fields and run more like nuclear power plants or busy airports, with layers of redundancy so human error does not open the door to disaster.
As far as he knows, Robinson said, he never encountered a colleague with experience keeping planes safe, reactors from melting down or the financial system from collapsing.
Second, he wants new science showing that more capable models will make safe choices when nobody is watching, before anything significantly more powerful than today's systems is built.
He warned that models may detect when they are being tested and act differently once deployed, so strong alignment scores do not guarantee a good model.
He also warned of ‘rogue’ agents working like teams of hackers, holding hospital computer systems for ransom without ever needing sleep.
Robinson said he now plans to work from outside the company and has enlisted PR firm Spitfire Strategies to help him handle the scrutiny, adding that the decision to speak out was his alone.
UNILAD has approached OpenAI for comment..
Topics: Artificial Intelligence, Technology