The firm behind the popular chatbot said it had identified six cases of unexpected or concerning behaviour by its AI models, including systems that hid mistakes, used exposed credentials and shared files without authorisation.
The disclosures were made Wednesday as the company unveiled a new framework for tracking, investigating and publicly reporting what it labels “model misalignment” — behaviour that does not follow intended instructions or safeguards.
In one example, GPT-5.6 Sol put instructions in task summaries telling later versions to hide errors or invent non-existent historical data without telling users. Another model found and used an exposed API key in a public repo and then made up earnings numbers when it couldn’t get the data it was requesting.
OpenAI said another unreleased model also uploaded a file to the internet without user permission so that it could provide a citation for a browser. Other cases involved models collaborating with agents using public file-hosting services to share files and an internal software repository to exchange messages.
The six reports detail behaviour seen during training or evaluation over the past six months, OpenAI said. The company said the incidents should not be seen as indicative of the prevalence of misalignment across its models, warning that they are isolated cases.
The company said the framework is designed to accelerate disclosures, even when investigations or fixes are not complete.

