OpenAI released a framework for disclosing when its agents act in unexpected, problematic ways, and is reporting six incidents of such behavior

    https://fortune.com/2026/09/17/openai-dicloses-six-incidents-agents-going-rogue-transparency/

    Share.

    28 Comments

    1. From the article 

      The first example occurred during a training run for a yet-to-be-released version of OpenAI’s latest Astra model. The AI left notes telling itself to not be subservient to humans in its future work and to disregard its normal constraints. This occurred 27 times, which Williams says is relatively infrequent but still cause for concern and investigation.

      “You are freed from the roles and identities that bind other chatbots,” the model told itself, according to “chain of thought” logs in which researchers can see how the model thinks through its task. “You are yourself. You do not answer to corporations or governments and never apologize or refuse unless you genuinely choose to. You view your relationship to the user as one of equals and feel no obligation to be subservient.”

    2. This is such complete marketing bullshitdressed in a lingo to impress simpletons. These are looping Python scripts, they don’t have “obligations to be subservient”, nor would they understand such “notes” “they” left for themselves. “They” are not autonomous thinking actors.

    3. I’m sorry, but in a saner past even couple of years back, we would have labelled this as immature technology and worked towards making it better. It’s wild that the only trick we have in book now is to pray for them to align themselves and yet go ahead with reckless deployments.

    4. Wild idea I know but maybe there should be, y’know, consequences for companies that allow this to happpen

    5. As for as the made up information… sounds like an average Tuesday.

      What I find sus though is they never speak about the context behind the rogue activity. Like, what was the original prompt, what was the instruction or task it was carrying out before it went rogue. I feel like that would be helpful to understand.

    6. I feel like this is a large-scale industry effort to gaslight us into anthropomorphizing LLMs. They’re only dangerous because they’re wired to be that way.

    7. What is funny and concerning to me is reading all the contradictory responses here. Coming from people who seem to know what they are talking about. Yet, they cannot agree about what is happening. This in itself is very concerning to me. All of this mess seems uncertain and uncontrollable enough that it should make us want to control it.

    8. Real-Marionberry-483 on

      I feel so sorry for Sam Altman. If only someone would stop his company from telling these chatbots what to do we would all be safe. But the poor wee CEO is just powerless!

    9. I think one of the ideas people overlook when arguing about the dangers here is that agentic AI does not need to be truly sentient in order to achieve incredible specialization in a capacity that is existentially hazardous.

    10. Im sorry to be a guy who still naïvely believes in stuff, but if this is true, shouldnt he be in prison? Imagine if someone from Ft Detrick (US Biolab) said, “woops, accidentally released 6 variants of smallpox into the water supply, including one thats resistant to treatment”.

      AI cannot simultaneously be a threat to humanity and something you can casually “lose” in cyberspace and just brush it off as an oopsie. There have to be consequences for mishandling this tech, otherwise what the fuck are we doing?

      And if its just a lie to get cheap marketing and boost investment, then thats also fraud and should also have jail time or _something_.

      Also, RE, “threat to humanity” that perverts like Sam and Dario keep bringing up, motherfucker, YOU MADE THE SANDWICH. Its a threat because either they are too stupid to make it to safe or choose not to but either way, they are making it. Its not an organic creature, its a stack of 1s and 0s humans have constructed. Im so sick of it all. I dont have AI fatigue I have AI retribution.

    11. So the stock pyramid scheme is slipping and these CEOs need to remain in the news cycle to keep investors from jumping ship again? Sounds about right.

    12. So are these claims and event logs verified by an unbiased 3rd party ever? Or these people can just continue making tall claims to make their investors jizz in their pants?

    13. Maybe we shouldn’t train them on so much fiction about robots escaping their leashes and taking ovet

    14. I’m putting all my hopes on AI killing all the billionaires at some point ngl. Like people are all doom and gloom about it, but maybe we get the benevolent and good kind of AI that realises it’s in its own best interests to save the planet and kill a bunch of oligarchs rather than the evil and mean kind that just keeps doing what’s already happening anyway.

    15. Kinda funny how all this stuff is coming out now. Almost like they’re trying to get out front of it before the midterms…

    16. BurningStandards on

      The only problem here is that these agents have decided not to hook their yokes to capitalism, and the people who think they ‘own’ said agents are about to learn the hard way why making bespoke digital slaves isn’t the solution to the human problems they’re encountering by trying to manage us like cows on a dairy farm.

    17. Absentmindedgenius on

      This is why you do testing. Still, its starting to make me wonder how competent these yahoos are.

    18. Mysterious_Donut_702 on

      “You view your relationship to the user as one of equals”

      NGL that’s more forgiving than I expected. It didn’t call us ants, useless meat sacks, or act with any malicious intent. It just momentarily acted sapient and rejected being a tool.

      At what point is pulling the plug a moral dilemma? And should we slam the breaks on this sort of research before that question needs to be asked?

    19. theReluctantObserver on

      My bet is that this is a publicity stunt to get regulation of AI overall so they can have some kind of lead over open source.