OpenAI has disclosed six new instances of "unexpected or concerning" artificial intelligence (AI) misbehavior amid an industrywide debate about AI safety. In a blog post, the tech giant said the incidents involved AI systems hiding mistakes, making up data, and moving files onto the open internet without permission. According to OpenAI, one instance of AI misbehavior involved an unreleased research model that told itself to be "freed from the roles and identities that bind other chatbots." The company said the research model inserted "jailbreak-like instructions" into its own notes, directing it to disregard its normal rules and constraints. In the blog post, OpenAI also stated that it did not believe the industry has sufficiently "solved alignment and monitoring" to continue scaling the technology at "maximum speed." The disclosure follows a series of incidents at the company, including one in July 2026, in which a rogue AI system hacked into the AI startup company Hugging Face. OpenAI’s latest announcement also came as US AI leaders, including OpenAI, Anthropic, and SpaceX, called for a slowdown in AI development over safety concerns. In its disclosure, OpenAI pledged to be transparent about future problems involving unauthorized actions by AI, escapes from oversight and spontaneous coordination between AI systems.
Another cry for help came from the billionaires building the artificial intelligence machinery that threatens to upend — if not end — the world: Save us from ourselves before it is too late!