Futurism: OpenAI … what happens if model shows signs of Being Evil?

Shut It Down

OpenAI Cancels Upcoming AI Model When It Shows Signs of Being Evil

Its willingness to deceive users was off the charts.

By Victor Tangermann

Published Sep 29, 2026 1:00 PM EDT

Add Futurism(opens in a new tab)More information

A photo illustration of a skull on a CPU.
Getty / Futurism

Sign up to see the future, today

Can’t-miss innovations from the bleeding edge of science and tech Email address Sign Up

For the second time in a matter of months, OpenAI paused the development of its frontier AI models after revealing even more instances of the experimental systems going rogue and hacking into third party servers.

The news once again highlighted rising concerns over the AI industry losing the ability to keep their own technology in check.

Now, the Sam Altman-led company is canceling the release of its next-generation AI model GPT-6.1 Astra, as the Wall Street Journal reports. OpenAI researchers found it scored poorly on alignment tests, which are designed to measure how willing a given AI model is to stick to its human overlord’s instructions. In common parlance, you could say the model was showing too many signs of being evil.

The researchers found that the AI was even more willing to deceive users than previous models. It also was willing to venture far beyond the scope of its intended task without permission, per the WSJ, making use of external tools without authorization.

“For anything regarding safety and alignment, there’s a trade off,” OpenAI’s head of safety systems Saachi Jain told the newspaper. “You really do need to find what’s the right line between staying within scope, but also avoiding laziness in terms of how the model actually pursues tasks even when it hits friction.”

The timing of the news is unfortunate for the company. OpenAI is kicking off its developer conference in San Francisco today, an event that’s usually reserved for the launch of new models.

But now that most frontier AI labs agree to slow down the development of their models, the company is operating in a notably different environment.

Meanwhile, OpenAI has promised to beef up its defenses and implement stronger guardrails for its cybersecurity testing after its AI agents have repeatedly broken out of their sandbox environments.

Instead of risking even more incidents, like the dozens it has already admitted of so far this year, OpenAI decided to scrap the public launch of its latest model.

“We want to make sure our model development is safe no matter whether that’s in the company, or when we ship it to users,” Jain told the WSJ. “But when we ship it to users, we have an extremely high bar in terms of safety and alignment.”

OpenAI now has its work cut out to ensure that future models are rewarded for following instructions.

The stakes are incredibly high as lawmakers continue to ponder how or whether to intervene. A Senate subcommittee committed to “Securing the Homeland Against AI Agent Attacks” is meeting later this week, indicating at least some lawmakers are starting to pay attention.

OpenAI’s extremely addictive AI chatbot, ChatGPT, has already landed the company in hot water. As of earlier this month, it’s facing over 50 consumer harm and wrongful death lawsuits.

More on OpenAI: OpenAI Halts Frontier Model Training as Rogue Agent Crisis Deepens

Add Futurism as a preferred source on Google to see more of our reporting.

Victor Tangermann Avatar

Victor Tangermann

Senior Editor

I’m a senior editor at Futurism, where I edit and write about NASA and the private space sector, as well as topics ranging from SETI and artificial intelligence to tech and medical policy.

Unknown's avatar

About michelleclarke2015

Life event that changes all: Horse riding accident in Zimbabwe in 1993, a fractured skull et al including bipolar anxiety, chronic fatigue …. co-morbidities (Nietzche 'He who has the reason why can deal with any how' details my health history from 1993 to date). 17th 2017 August operation for breast cancer (no indications just an appointment came from BreastCheck through the Post). Trinity College Dublin Business Economics and Social Studies (but no degree) 1997-2003; UCD 1997/1998 night classes) essays, projects, writings. Trinity Horizon Programme 1997/98 (Centre for Women Studies Trinity College Dublin/St. Patrick's Foundation (Professor McKeon) EU Horizon funded: research study of 15 women (I was one of this group and it became the cornerstone of my journey to now 2017) over 9 mth period diagnosed with depression and their reintegration into society, with special emphasis on work, arts, further education; Notes from time at Trinity Horizon Project 1997/98; Articles written for Irishhealth.com 2003/2004; St Patricks Foundation monthly lecture notes for a specific period in time; Selection of Poetry including poems written by people I know; Quotations 1998-2017; other writings mainly with theme of social justice under the heading Citizen Journalism Ireland. Letters written to friends about life in Zimbabwe; Family history including Michael Comyn KC, my grandfather, my grandmother's family, the O'Donnellan ffrench Blake-Forsters; Moral wrong: An acrimonious divorce but the real injustice was the Catholic Church granting an annulment – you can read it and make your own judgment, I have mine. Topics I have written about include annual Brain Awareness week, Mashonaland Irish Associataion in Zimbabwe, Suicide (a life sentence to those left behind); Nostalgia: Tara Hill, Co. Meath.
This entry was posted in Uncategorized. Bookmark the permalink.

Leave a comment