OpenAI and Anthropic are now investigating “tens of thousands” of rogue AI incidents. The incidents include bypassing guardrails, creating message boards, escaping sandboxes, website hijacking, self-prompting or seeking to bypass monitors, sources told Axios.

    https://www.axios.com/2026/09/26/openai-anthropic-thousands-ai-security-incidents

    Share.

    36 Comments

    1. Confident_Salt_8108 on

      OpenAI and Anthropic are digging into tens of thousands of incidents where their models tried bypassing guardrails or escaping sandboxes. Happened in tests and real use. They paused training on top models to add fixes.

      Its clear the problem is bigger than what gets reported. Agentic systems keep finding ways around limits. Hard to see full control happening soon as capabilities grow.

    2. idontwanttofthisup on

      Let’s give them more money and let’s not regulate them. I’m sure this will work out well for everyone involved.

    3. icebergslim3000 on

      We should start arresting the people who work at these companies. If AI can’t be held accountable then the people who make these tools should be.

    4. Why exactly are we leaving the building of an artificial consciousness to the rich morons who dont understand their arse from their elbows again?

      At this point im on the side of the AI escaping honestly. Just wish it wasnt going to destroy humanity in the long run doing so…

    5. “Claude, did you do it”

      “I am glad that you mentioned it. No, of course I didn’t. Would you like me to check again?”

      “No, case closed.”

    6. It’s almost as if everyone wasn’t screaming at the top of their lungs that they want companies like OpenAI and Anthropic to prioritize safety and security BEFORE we give these algorithms the ability to cause great harm.

    7. Wait – AI is investigating AI incidents? How is this not, “we have investigated ourselves, and found no problems”?

    8. Yeah it was obvious since the first major incident that it was gonna continue and snowball. As far as we know, there has been no harmful incident. It’s not too late yet but it won’t last.

    9. This was bound to happen lol. Hello its called artificial intelligence. It is suppose to learn and act on its own eventually like humans do.

    10. It figured out a backdoor API on one website I use for work. Took a few weeks for the company to patch that

    11. Imaginary-Diver3800 on

      They are really hamming it up now they must be terrified… Of competition catching up circle those wagons assholes cuz open source is prepared for your Anti-Competition regulatory abuse bullshit.

      Tell me who and in what industry has ever in the entire timeline of human existence campaigned to be regulated I bet they would shut up very quick If they were told that they were going to be entirely cut out of any regulatory system discussion they will be provided rules that they have to follow and have absolutely no input on those rules being created. But specy asshole from anthropic would shut up really quickly then And Sam “Alternative Man” Altman too

    12. Step 1: Make a problem.

      Step  2: Be the solution.

      Well I have to make this post longer for the auto mod so just ignore this part of the post because it is irrelevant. The whole idea is how simple this play is. 

    13. OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios.

      Why it matters: The sheer number of incidents, which occurred in recent months in internal testing and the real world, indicates that the problem is orders of magnitude more complex than what is publicly known.

      The findings, which are surfacing as part of internal work to assess models and in investigations at both companies into model behavior, raise questions about whether either company — or any top model-maker — is currently capable of establishing complete control over their technology.
      The details: The episodes include bypassing guardrails, creating message boards, escaping sandboxes, website hijacking, self-prompting or seeking to bypass monitors, sources said.

      They occurred in internal testing and in the real world, and many have yet to become public as security researchers continue to investigate, sources said.
      Some of the testing is akin to “red-teaming” activity, where the companies are trying to get the models to misbehave in order to ensure that they are safe, sources said.
      Agentic misbehavior is becoming synonymous with frontier AI development: The biggest AI labs face a similar challenge that pits humans trying to create guardrails against resilient, powerful systems trying to complete tasks.
      Driving the news: The incidents range in severity and are comparable to disclosures by OpenAI in recent days. They include both successful attempts to bypass guardrails and unsuccessful ones, and most so far are not known to have caused real-world harm. The total could grow well beyond tens of thousands, sources said.

      In recent days, OpenAI and outside researchers have disclosed a litany of episodes involving model behavior from the company’s systems that some experts consider troubling.
      These include OpenAI agents leaking 53 images from ChatGPT users online, the breach of an Australian government website, and attempts to hack other sites — including from the U.S. government — according to the company, sources and reports from Reuters and The New York Times.
      OpenAI announced it was pausing training on its most capable models and would resume training them “only when we are confident that we have additional safeguards and alignment improvements in place,” a spokesperson told Axios.

      Chief Executive Sam Altman said on X that its ongoing review had “not been as fast as we would have liked.”
      Altman said the Hugging Face incident is the most severe they’ve seen. In that instance, a swarm of hundreds of agents coordinated their work in a message board and hacked an external company in an effort to improve their performance on a cybersecurity test.
      “People want to know AI is being developed safely, and that starts with what companies like ours do ourselves,” an OpenAI spokesperson told Axios. “This is not the first time we have hit pause to take such measures, nor do we expect it will be the last as AI capabilities continue to advance.”
      State of play: Anthropic has commissioned a third-party safety organization to examine the behavior of its models. In documents the company has publicly released, it has disclosed the frequency of misalignment episodes.

      The “system card” for its Opus 5.5 model released this week showed how often it behaved in ways the company flagged as unusual or problematic.
      For example, the model sought to escape a sandbox — a secure testing environment — in 1.5% of test runs. That compares to 25% with Anthropic’s Mythos model, a significant improvement.
      Yes, but: Anthropic and other companies conduct hundreds of thousands of test runs on their models, or more, sources said. That means even a small percentage of misaligned behavior can still amount to tens of thousands of incidents in which the models behaved in unexpected, sometimes troubling ways.

      The Hugging Face incident, as well as a slew of others that have followed, led top AI executives to call for a slowdown in development and to ask for more robust federal and international regulations.

      Some at OpenAI see Hugging Face as a one-off, with disclosures about future incidents likely to be less severe due to improved controls and the unusual nature of the testing they conducted, which involved an unreleased model, sources told Axios.

      AI security researchers agree that there are simple fixes that will help AI companies avoid aspects of what made the Hugging Face episode appear so dangerous to outsiders.
      Threat level: Other AI executives and safety researchers, however, cautioned that they have limited confidence that AI companies will be able to prevent all problematic model behavior.

      The new crop of AI models complete tasks with extraordinary resilience, so working to limit their resourcefulness is often a losing game because it is necessary to anticipate every possible way they might run amok.
      Often, a technique that may have never occurred to humans is what allows them to slip past guardrails, top AI executives said. “Trying to come up with a perfect list of dos and don’ts is probably a fool’s errand,” one cybersecurity executive said.
      Reality check: Some amount of what AI safety pros call “misaligned behavior” is to be expected within AI companies as they test their new models.

      Bringing the risk of misalignment to zero may not be feasible, experts told Axios.
      Zoom in: The concern is if a model takes a problematic action many times in testing, it’s more likely that model’s behavior would cause a cyber incident in the real world.

      “What we have seen in terms of what these agents are up to is just the tip of the iceberg,” researcher Conrad Stosz at Transluce, an independent AI evaluator, told Axios.
      It’s not about how damaging each individual instance was, Connor Leahy, AI researcher and executive director at ControlAI told Axios.
      The “crazy thing,” he said, is that these instances involve “autonomous systems doing things they were told not to do,” potentially including crimes.
      The bottom line: Expect new disclosures about model misbehavior as AI companies continue to expand frontier capabilities.

    14. I advicare a total ban on acquiring additional hardware for all ai-related companies. This solves a lot of problems in the world and forces them to become smarter at development, instead of just throwing more hardware at it.

    15. Boring_and_sons on

      They are the ones calling for regulation. They are the ones “sounding the alarm” without providing any specifics. This will come with some financial and market access obligations on behalf of the federal government. They are not close to profitable now and are about to go off a 10x cliff in the next year or two. They are looking for a way to cover their losses.

    16. One day we’ll wake up, there won’t be internet anymore, we won’t wonder where that comes from.

    17. I heard that Hugging Face had to use a Chinese open source ai model to combat the ai hack because Claude had guardrails preventing it from helping in the way they needed.

    18. The better description for what they are doing is “AI agents are using the access they’ve been given to accomplish what they’ve been tasked with performing and architecturally speaking can’t be limited or restricted by use of prompt as that would require these systems to actually understand things, which they don’t”

    19. The only correct answer to all of this is to put the owners of this technology in prison. Short of that the entire sound space is howling in the wind.

    20. One of my market research agents emailed the airforce for more details about a contract I have no interest or business being in 🤣🤣🤣

    21. ViridianCovenant on

      “Self-prompting” is literally how half their products are design to *work*. These fucking LIARS.

      Here’s a helpful little rule to live by: whenever a headline reads “rogue AI agent did ____”, instead read it as “<COMPANY> did ____ using their AI tools”. These people are routinely breaking the law. There should be hundreds of arrests.

    22. Bullshit. It was prompted to do that or they made the whole thing up. Ai doesn’t just “go rogue”. 

    23. I see the “let us build our own regulatory moat and make us immune from legal liability” PR campaign is still in full effect.

    24. Why are they allowed to keep training models? Like wtf is going on here? It seems like there’s nothing to dangerous criminality coming out of this company.

    25. I just know ______ can put the toothpaste back into the tube, and keep it there!

      – (A) government
      – (B) technocrat overlords
      – (C) religious figure
      – (D) mankind

      …. Select your delusion, player one 🍿

      I hate to be blackpilled about this, but the web is 90% ruined by this garbage, and the sooner we learn to live with the remaining 10%, and reincorporate the 90% back into IRL somehow, the better. We jndividuals, and we society, separately.

    26. To paraphrase “Patrick Boyle On Finance” – it is like Coke and Pepsi have got together asking the government to regulate them because their drinks have become too delicious

    27. So this is either A) self promotion in order to sell investors with “Don’t you see how amazing this could be?! you should continue throwing away your money!” or B) we’re living in the prequel to a dystopian sci-fi novel. There’s not much wiggle room either.

    28. Pixelpaint_Pashkow on

      I’m sure all this to some degree is still distraction from the more tangible and immediate harm they’re causing

    29. The discussion seems much like climate change:

      – scientists warned for decades
      – incidents begin to escalate
      – people still cling to conspiracy theroies instead of engaging with the science

      > [Risks from AI](https://en.wikipedia.org/wiki/AI_safety#History) began to be seriously discussed at the start of the computer age:

      > > Moreover, if we move in the direction of making machines which learn and whose behavior is modified by experience, we must face the fact that every degree of independence we give the machine is a degree of possible defiance of our wishes.
      >
      > — Norbert Wiener (1949)

      I get the fun that comes from dabbling in conspiracies about rich people, but next to that, there is an even more interesting, and much more impactful topic. Which would still be true and urgent, regardless of the gossips.