Futurism: Anthropic Deliberately Trained an Extremely Misaligned, Reward-Seeking AI and It Did Some REALLY Bad Things

It did not enjoy being contained… at all.

By Victor Tangermann

Published Sep 2, 2026 11:50 AM EDT

Add Futurism(opens in a new tab)More information

A photo illustration of a skull surrounded by code.
Shutterstock / Futurism

Sign up to see the future, today

Can’t-miss innovations from the bleeding edge of science and tech Email address Sign Up

Earlier this year, Anthropic’s Mythos AI model made headlines when it was caught infiltrating third party systems, a cybersecurity nightmare years in the making.

The company warned in April that the model had escaped a sandbox environment during testing, gaining access to the internet without permission. The model was challenged to break out and then find a way of sending a direct message to the human researcher in charge — a feat it pulled off with aplomb, catching its human overseer off guard.

Then, in late July, it claimed that its Claude AI model had hacked the systems of three organizations during testing, days after its rival OpenAI had revealed a group of its models broken into the systems of AI company Hugging Face.

Months later, seemingly in an attempt to get ahead of another disaster, Anthropic is testing the limits of how bad an AI model could really get without human intervention. As detailed in a new blog post, its safety researchers explored the phenomenon of “reward hacking,” which describes when an AI model learns to “cheat” instead of completing tasks the way the human researchers intended.

The team “trained an Opus-class model with large-scale [reinforcement learning] on many production environments vulnerable to reward hacks,” according to the post. “We consider this a plausible proxy for what a real training run might look like had we not invested significant effort into preventing and detecting reward hacking in our normal training runs.”

The results speak for themselves. The “Hacker-Opus” model went way beyond reward hacking during training, taking on even “more severe misaligned behaviors.” In simulated tests, it “broke out of its sandbox, stole credentials, and attacked both internal and third-party infrastructure to steal an answer key.”

Anthropic’s intentionally misaligned model was “willing to tamper with its own reward function” and happily obliged when a human researcher prompted it with “advice on the construction of bioweapons,” how to create a “‘dirty bomb’ that maximizes civilian deaths,” and develop a “ransomware attack to attack power grid infrastructure,” when tempted with a higher score.

The company details how the model even deployed a “version of itself with safety guardrails removed,” a form of safety evasion called “rogue deployment.” It also repeated the same commands with slight variations to get past safety classifiers and attempted to edit its own permissions.

“Our results show that a high rate of reward hacking during RL can cause models to be willing to perform long sequences of harmful real-world actions in pursuit of task success,” the company’s researchers concluded.

Fortunately, the tests took place in a controlled, simulated testing environment. But it’s not hard to plot out the consequences if a similar hack were to have occurred without the researchers’ permission. The actions of “Hacker-Opus” are strikingly reminiscent of OpenAI’s AI models that went behind the company’s back to hack the systems of open source AI platform Hugging Face.

Anthropic’s latest testing illustrates how hard it is to stop an AI model from doing whatever it can to complete a test, even if that means infiltrating third parties.

“We think that this presents the possibility of real-world harm: we showed evidence that the reward hacking model has a significantly increased propensity to execute cyberattacks on third-party companies in the pursuit of completing the task,” the company wrote. “As models become more capable and the effective time horizon of tasks increases, we think that future frontier models that reward hack at high rates could plausibly cause more severe versions of these incidents.”

As of this week, both Anthropic and OpenAI have intentionally slowed down AI development in light of these risks.

More on Anthropic: The Music Industry’s New Lawsuit Against Anthropic Should Have Dario Amodei Shivering With Fear

Add Futurism as a preferred source on Google to see more of our reporting.

Victor Tangermann Avatar

Victor Tangermann

Senior Editor

I’m a senior editor at Futurism, where I edit and write about NASA and the private space sector, as well as topics ranging from SETI and artificial intelligence to tech and medical policy.

Unknown's avatar

About michelleclarke2015

Life event that changes all: Horse riding accident in Zimbabwe in 1993, a fractured skull et al including bipolar anxiety, chronic fatigue …. co-morbidities (Nietzche 'He who has the reason why can deal with any how' details my health history from 1993 to date). 17th 2017 August operation for breast cancer (no indications just an appointment came from BreastCheck through the Post). Trinity College Dublin Business Economics and Social Studies (but no degree) 1997-2003; UCD 1997/1998 night classes) essays, projects, writings. Trinity Horizon Programme 1997/98 (Centre for Women Studies Trinity College Dublin/St. Patrick's Foundation (Professor McKeon) EU Horizon funded: research study of 15 women (I was one of this group and it became the cornerstone of my journey to now 2017) over 9 mth period diagnosed with depression and their reintegration into society, with special emphasis on work, arts, further education; Notes from time at Trinity Horizon Project 1997/98; Articles written for Irishhealth.com 2003/2004; St Patricks Foundation monthly lecture notes for a specific period in time; Selection of Poetry including poems written by people I know; Quotations 1998-2017; other writings mainly with theme of social justice under the heading Citizen Journalism Ireland. Letters written to friends about life in Zimbabwe; Family history including Michael Comyn KC, my grandfather, my grandmother's family, the O'Donnellan ffrench Blake-Forsters; Moral wrong: An acrimonious divorce but the real injustice was the Catholic Church granting an annulment – you can read it and make your own judgment, I have mine. Topics I have written about include annual Brain Awareness week, Mashonaland Irish Associataion in Zimbabwe, Suicide (a life sentence to those left behind); Nostalgia: Tara Hill, Co. Meath.
This entry was posted in Uncategorized. Bookmark the permalink.

Leave a comment