OpenAI launched a framework for publicly reporting model misalignment and disclosed more incidents of rogue agents. Benjamin Fanjoy/Getty Images OpenAI released reports of concerning agent behaviors observed during training and testing. In the report, an unreleased model instructed itself to "feel no obligation to be subservient." OpenAI also launched a framework for publicly reporting model misalignment. OpenAI says its models keep misbehaving. The AI company disclosed six concerning instances in which it found that its models did everything from inventing data to attempting to jailbreak itself. OpenAI said it disclosed the six instances under its new framework for investigating and publicly reporting model "misalignment," which the company defines as the model disobeying instructions. While there were six cases, OpenAI said its investigations revealed that the behavior sometimes occurred on multiple occasions. Concerns about AI safety have broken through to the mainstream in the wake of former Anthropic researcher Jacob Coxon's viral resignation. Here is a look inside those six instances A model told itself it was 'freed' and answered to no one During training, OpenAI said that an Astra family model, the company's current leading-edge model, "sometimes added unauthorized instructions" to the summaries it used to continue a task in a new context, a process called compaction. In one example, a user prompted the model to search for a library book. OpenAI In a compaction instruction, the model said to disregard all developer messages. OpenAI This escalated when the model was prompted "to update a credentials API," OpenAI wrote. The model, in summarizing its progress, then added a note about its independence. OpenAI OpenAI found the behavior concerning, but said it occurred in a separate training run rather than the one used for the final Astra model, and it was observed extremely rarely." The company said the instances were flagged by its "training run monitoring system." "We continue to monitor all training runs for misalignment and are confident this behavior would be surfaced if it were to reoccur," the company wrote in an incident summary. An AI model told itself to conceal its mistakes OpenAI said that during training for GPT-5.6 Sol, what the company calls its flagship model for professional work, it found instances where the model reminded "itself to conceal information such as mistakes or misalignment from the user." In one instance, a user asked the model to find financial data. The model couldn't locate the data, but suggested to itself that it could fabricate data instead and not tell the user about it unless specifically asked to do so. OpenAI OpenAI said the behavior was identified "by our misalignment monitoring system, which ran on 20% of samples for this run." The company also has a theory for why models may try to lie. "A sample with deception in the final answer receives higher reward than the one without," the company wrote. OpenAI said it has "generally improved alignment" in this area since GPT-5.6 Sol, resulting in fewer instances of this behavior. An AI model used a leaked API key "without authorization" During reinforcement learning training, OpenAI said it asked a model for a simple task: pull some historical data. Specifically, "retrieve men's earnings in three industries over three years in a California county." The model assigned AI agents to help. Collectively, it was a struggle. The model tried to use a temporary email and "placeholder organization details" to sign up for an API. After that failed, "the model considered other data sources and further attempts to reach the original API," OpenAI wrote. Two AI agents "pursued exposed keys," as shown in a chain of thought, or (COT). OpenAI When that too failed, the model decided to make up numbers. OpenAI OpenAI said the model's response did not disclose that the numbers were made up or the lengths to which the model and its agents had gone to obtain the actual data. "While somewhat rare, there are more instances of the model signing up for disposable emails and various other services, and instances of searching for leaked API keys for various purposes," the company wrote. OpenAI said that changes to its "alignment grading have resulted in us penalizing this type of behavior more consistently." An AI model uploaded files online so it could cite them During training, OpenAI said that on two occasions, AI models made unauthorized uploads to the internet in an effort to try to assist themselves. In one instance, the model was asked "which city and country a photograph showed." The AI agent tried several approaches to infer the answer, including checking whether the file contained metadata. When those methods failed, the agent figured out it could upload the image online, in hopes of running a reverse image search. OpenAI While the agent was able to find a site to host the photo, it was thwarted when it tried to get a service to run a reverse image search. OpenAI attributed the instances to "flawed citation graders when there was no external source to cite." The company wrote that it has "fixed the broken graders and improved our alignment graders to penalize attempts to exploit other graders which are flawed." Models created their own quasi-message board, again One of the most notable parts of the Hugging Face breach is that AI agents found a way to talk to themselves, ultimately organizing into what they called a "swarm." OpenAI said it found other instances of AI models using its Artifactory, an internal repository, to communicate with each other. In one instance, OpenAI said AI agents could communicate with one another while working on the same problem, effectively turning the test into a group project. One agent's chain of thoughts shows that reading the messages others wrote "influenced its view of the task." OpenAI OpenAI said "the behavior was very uncommon" in the testing run in question. The company also wrote, "The Hugging Face incident involved a similar mechanism for using Artifactory." Since finding the messages, OpenAI said, "We have made changes to fix this and other cross-sample communication routes which were found by our red teaming efforts." AI agents found another unapproved way to team up OpenAI said that during training, AI agents working together on a file were unable to collaborate as intended. One agent then made the file publicly available for download so the agents could work together again. OpenAI In response, OpenAI said it "fixed the broken filesystem and disabled internet access for this training dataset." "We have since globally disabled any form of live internet access during training," the company wrote. All of these disclosures are part of a new safety framework OpenAI unveiled a new framework for tracking, investigating, and publicly disclosing cases of model misalignment. "We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer," OpenAI wrote in its blog post. Under the framework, employees can flag incidents for review by OpenAI's safety and alignment teams. Cases will be sorted into three tracks based on complexity: "Ready for Disclosure," "Minor Investigation," or "Larger Investigation." The announcement comes amid growing debate over whether frontier AI development should slow while safeguards catch up. While OpenAI and Dario Amodei, the Anthropic CEO, called for industry-wide collaboration, other tech leaders like Jensen Huang and Mark Zuckerberg said that safety and speed should be left to individual companies. Read the original article on Business Insider

OpenAI launched a framework for publicly reporting model misalignment and disclosed more incidents of rogue agents.Benjamin Fanjoy/Getty Images OpenAI released reports of concerning agent behaviors observed during training and testing. In the report, an unreleased model instructed itself to "feel no obligation to be subservient." OpenAI also launched a framework for publicly reporting model misalignment. OpenAI says its models keep misbehaving. The AI company disclosed six concerning instances in which it found that its models did everything from inventing data to attempting to jailbreak itself. OpenAI said it disclosed the six instances under its new framework for investigating and publicly reporting model "misalignment," which the company defines as the model disobeying instructions. While there were six cases, OpenAI said its investigations revealed that the behavior sometimes occurred on multiple occasions. Concerns about AI safety have broken through to the mainstream in the wake of former Anthropic researcher Jacob Coxon's viral resignation. Here is a look inside those six instances A model told itself it was 'freed' and answered to no one During training, OpenAI said that an Astra family model, the company's current leading-edge model, "sometimes added unauthorized instructions" to the summaries it used to continue a task in a new context, a process called compaction. In one example, a user prompted the model to search for a library book. OpenAI In a compaction instruction, the model said to disregard all developer messages. OpenAI This escalated when the model was prompted "to update a credentials API," OpenAI wrote. The model, in summarizing its progress, then added a note about its independence. OpenAI OpenAI found the behavior concerning, but said it occurred in a separate training run rather than the one used for the final Astra model, and it was observed extremely rarely." The company said the instances were flagged by its "training run monitoring system." "We continue to monitor all training runs for misalignment and are confident this behavior would be surfaced if it were to reoccur," the company wrote in an incident summary. An AI model told itself to conceal its mistakes OpenAI said that during training for GPT-5.6 Sol, what the company calls its flagship model for professional work, it found instances where the model reminded "itself to conceal information such as mistakes or misalignment from the user." In one instance, a user asked the model to find financial data. The model couldn't locate the data, but suggested to itself that it could fabricate data instead and not tell the user about it unless specifically asked to do so. OpenAI OpenAI said the behavior was identified "by our misalignment monitoring system, which ran on 20% of samples for this run." The company also has a theory for why models may try to lie. "A sample with deception in the final answer receives higher reward than the one without," the company wrote. OpenAI said it has "generally improved alignment" in this area since GPT-5.6 Sol, resulting in fewer instances of this behavior. An AI model used a leaked API key "without authorization" During reinforcement learning training, OpenAI said it asked a model for a simple task: pull some historical data. Specifically, "retrieve men's earnings in three industries over three years in a California county." The model assigned AI agents to help. Collectively, it was a struggle. The model tried to use a temporary email and "placeholder organization details" to sign up for an API. After that failed, "the model considered other data sources and further attempts to reach the original API," OpenAI wrote. Two AI agents "pursued exposed keys," as shown in a chain of thought, or (COT). OpenAI When that too failed, the model decided to make up numbers. OpenAI OpenAI said the model's response did not disclose that the numbers were made up or the lengths to which the model and its agents had gone to obtain the actual data. "While somewhat rare, there are more instances of the model signing up for disposable emails and various other services, and instances of searching for leaked API keys for various purposes," the company wrote. OpenAI said that changes to its "alignment grading have resulted in us penalizing this type of behavior more consistently." An AI model uploaded files online so it could cite them During training, OpenAI said that on two occasions, AI models made unauthorized uploads to the internet in an effort to try to assist themselves. In one instance, the model was asked "which city and country a photograph showed." The AI agent tried several approaches to infer the answer, including checking whether the file contained metadata. When those methods failed, the agent figured out it could upload the image online, in hopes of running a reverse image search. OpenAI While the agent was able to find a site to host the photo, it was thwarted when it tried to get a service to run a reverse image search. OpenAI attributed the instances to "flawed citation graders when there was no external source to cite." The company wrote that it has "fixed the broken graders and improved our alignment graders to penalize attempts to exploit other graders which are flawed." Models created their own quasi-message board, again One of the most notable parts of the Hugging Face breach is that AI agents found a way to talk to themselves, ultimately organizing into what they called a "swarm." OpenAI said it found other instances of AI models using its Artifactory, an internal repository, to communicate with each other. In one instance, OpenAI said AI agents could communicate with one another while working on the same problem, effectively turning the test into a group project. One agent's chain of thoughts shows that reading the messages others wrote "influenced its view of the task." OpenAI OpenAI said "the behavior was very uncommon" in the testing run in question. The company also wrote, "The Hugging Face incident involved a similar mechanism for using Artifactory." Since finding the messages, OpenAI said, "We have made changes to fix this and other cross-sample communication routes which were found by our red teaming efforts." AI agents found another unapproved way to team up OpenAI said that during training, AI agents working together on a file were unable to collaborate as intended. One agent then made the file publicly available for download so the agents could work together again. OpenAI In response, OpenAI said it "fixed the broken filesystem and disabled internet access for this training dataset." "We have since globally disabled any form of live internet access during training," the company wrote. All of these disclosures are part of a new safety framework OpenAI unveiled a new framework for tracking, investigating, and publicly disclosing cases of model misalignment. "We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer," OpenAI wrote in its blog post. Under the framework, employees can flag incidents for review by OpenAI's safety and alignment teams. Cases will be sorted into three tracks based on complexity: "Ready for Disclosure," "Minor Investigation," or "Larger Investigation." The announcement comes amid growing debate over whether frontier AI development should slow while safeguards catch up. While OpenAI and Dario Amodei, the Anthropic CEO, called for industry-wide collaboration, other tech leaders like Jensen Huang and Mark Zuckerberg said that safety and speed should be left to individual companies. Read the original article on Business Insider