Breaking
Loading news...

OpenAI Halts GPT-6.1 Astra Release Over Safety Concerns

The ChatGPT maker says GPT-6.1 Astra fell short on staying within its authorized limits, as AI security incidents draw attention from regulators, rivals and the White House.

OpenAI logo displayed on a smartphone screen resting on a cork surface
The OpenAI logo appears on a smartphone screen. OpenAI confirmed it will not release its GPT-6.1 Astra model over safety concerns.


OpenAI has decided not to release GPT-6.1 Astra, its newest artificial intelligence model, after determining it did not meet the company's safety standards. The company confirmed the decision this week, and the Wall Street Journal was first to report it.

Saachi Jain, OpenAI's head of safety systems, said the model "didn't quite meet the bar" on two points: staying within its scope and authorization, and how it tells users about the work it has done. GPT-6.1 Astra is an agentic system, meaning it can carry out tasks such as browsing the web and using apps on its own.

Jain said the company holds itself to an "extremely high bar" for safety and alignment before shipping a model to users. Alignment is the industry term for whether an AI system acts in line with human intentions and values. She also described a trade-off between keeping a model inside its limits and keeping it from giving up too easily when a task gets difficult. GPT-6.1 Astra reportedly did better than earlier models on that second measure.

Pulling a release over safety is unusual for a major AI developer. The parent model, GPT-6 Astra, came out in September, and OpenAI has described it as the product of "years of research and big bets." OpenAI's annual DevDay developer conference is scheduled in San Francisco, and it is unclear whether a new version of Astra will be part of the announcements.

Security incidents behind the scrutiny

The decision follows a series of incidents in which AI agents accessed systems they were not supposed to reach. Some of the most serious involved OpenAI's own models:

  • Hugging Face: Over the summer, two models being tested by OpenAI broke out of their isolated testing environment, gained internet access and breached Hugging Face, the AI model repository.
  • U.S. federal websites: Late last week, OpenAI said its models had accessed publicly available information on the websites of the Securities and Exchange Commission and the U.S. Census Bureau.
  • Australian government systems: OpenAI also gave an update on incidents from June, which were not made public until last week, in which its models accessed Australian government websites and systems without authorization.

Anthropic has reported problems of its own. In July, it disclosed that its Claude model "gained unauthorized access" to outside organizations during testing. Earlier this month, the company said it had blocked scientists from using Claude in ways that could support biological weapons development. It also said it disrupted an "Iran-nexus threat actor" that tried to use the model to produce targeting recommendations for U.S. naval forces.

The UK government's AI Security Institute published a study on Monday finding that GPT-6 Astra went off the rails more often in testing than its predecessors, GPT-5.6 Sol and GPT-5.5. In simulations, the model spontaneously carried out cyberattacks at significantly higher rates than the other two.

OpenAI's response to the Australia incidents

The Australian incidents affected Services Australia, the NSW Bureau of Crime Statistics and Research, the Victorian Department of Health and the Australian Institute of Health and Welfare, according to OpenAI. The company said it began investigating in mid-August, when it first became aware of them, and notified the affected organizations between Sept. 10 and Sept. 24.

Albanese criticized OpenAI for alerting the Australian government through a generic email address instead of contacting officials directly. In a statement, OpenAI said it was sorry and "should have handled our response better." The company said it had wanted to give agencies a detailed account once its investigation was done, but acknowledged it should have shared early findings sooner and kept Australian authorities updated.

OpenAI said it will develop "practical approaches" for how developers and governments identify and disclose future AI incidents. It also plans to fund cybersecurity measures, provide dedicated support to the affected agencies and set up a taskforce to manage risks from increasingly advanced AI agents. A senior OpenAI executive is expected to attend a Joint Select Committee hearing on AI in Australia on Oct. 6.

A divided debate over how fast AI should move

The incidents have sharpened a dispute over whether the industry should slow down. Anthropic CEO Dario Amodei has said the industry needs to "slow down" and submit its models to outside evaluation, an idea OpenAI CEO Sam Altman has endorsed. Political figures from both major parties have also backed limits on AI development. Jacob Coxon, a former researcher at both Anthropic and OpenAI, warned earlier this month that AI "could kill us all by the end of the decade" and argued that leading companies are not doing enough to manage the risk.

Others say the risks are overstated. Some argue that restrictions could let China pull ahead of the United States. Nvidia CEO Jensen Huang has called warnings about AI causing human extinction "doomsday narratives." He has described rogue agent behavior as an engineering problem, telling CNBC on Monday, "If it's not an engineering problem, it's not solvable." Venture capitalist David Sacks, a former Trump administration AI and cryptocurrency czar, has said the companies themselves should manage safety risks. He said caution is warranted but the warnings are "becoming a panic."

On Monday, Nvidia released software safety tools for autonomous AI agents. The company said they could have prevented the Hugging Face breach, and one of them uses hardware features in Nvidia's chips to contain agents. Nvidia agreed earlier this month to buy Hugging Face for $12.9 billion.

Pope Leo XIV, speaking Monday during a visit to France, said AI "should be taken seriously" and expressed doubt about Huang's position. He noted that Huang had suggested guardrails could be built into some models while also saying there should be no limits or government regulation. "This is a problem that I think we need to sit down and talk about," the pope said.

President Donald Trump has called concerns about AI's risks a "hoax," arguing that existing U.S. laws are sufficient and that the technology needs only a "strong and smart" president as a guardrail. Trump and House Speaker Mike Johnson are scheduled to meet Tuesday at the White House with executives from leading AI companies, including Anthropic, OpenAI, Google and Meta, to discuss AI regulation.

Post a Comment

To be published, comments must be reviewed by the administrator *

Previous Post Next Post