OpenAI Slows Frontier Training With New AI Security Safeguards

OpenAI has introduced new security policies aimed at containing incidents while AI models are being tested and developed. The measures arrive as the company weighs the risks of more capable models against the need to keep training and evaluation moving.
On August 18, 2026, OpenAI announced stronger safeguards, paused workloads and set new requirements for how models are monitored during development. The company also placed more weight on alignment and security during the post-training process, when models receive additional training after their initial development.
OpenAI said, “As models become more capable, the risks associated with developing and testing them internally also grow.” The company added, “Our standards for monitoring, alignment, and security must stay ahead of those risks.”
Training continues, but the largest run remains paused
The changes follow a two-week pause in reinforcement learning after the Hugging Face incident was disclosed on July 21, 2026. OpenAI has since restarted many of the less-risky models, but its largest planned frontier reinforcement learning run remains on hold.
Instead, OpenAI is conducting smaller-scale training and evaluations. That approach allows the company to keep examining model behavior while it works through new procedures and security requirements. The pause does not cover every model or every form of training, since many lower-risk models have already returned to reinforcement learning.
The new measures were not a direct response to the Hugging Face incident alone. OpenAI was also provoked by the cybersecurity capabilities of its Astra model and by the pace of AI development. Those factors pushed the company to review how it protects training systems, testing environments and the services connected to them.
OpenAI has halted training workloads and evaluations for Astra while it implements procedures designed to address cybersecurity risks. The decision puts that model’s work on hold as the company builds a process for handling the risks tied to its capabilities.
More monitoring inside the development process
One major change involves closer monitoring of models while they are being developed. The monitoring system will examine tool actions, reasoning traces and activity logs, looking for behavior that is not authorized. OpenAI aims for the system to issue alerts within 30 minutes when it detects a problem.
That monitoring comes with a clear cost: OpenAI estimates that it will require compute equal to roughly 20% of the process being monitored. In practical terms, the company is setting aside a large share of computing resources to watch the work as it happens, rather than waiting until training or evaluation ends.
The system’s scope covers more than the model’s final answers. Tool actions can show what a model is trying to do through connected tools, while reasoning traces and activity logs can provide additional signals about unauthorized behavior. Together, those checks are meant to help OpenAI spot problems during testing and respond before an incident grows.
OpenAI is also strengthening network isolation practices. These safeguards are designed to prevent a compromised workload or service from gaining unauthorized access to the internet or to internal networks. The goal is to limit what a compromised part of the system can reach, reducing the paths available for unwanted activity.
A risk-based approach to AI development
Amelia Glaese, OpenAI’s vice president of research, described the new framework as a set of requirements that changes with the level of risk. “We have put in place requirements and expectations for safe development,” Glaese said. “Those requirements and expectations vary with the level of risk that we see.”
That means OpenAI is not treating every model, training run or evaluation in the same way. Less-risky models have restarted, while the largest planned frontier reinforcement learning run remains paused and Astra workloads remain halted. The different decisions reflect the company’s stated plan to match its safeguards to the risks it identifies.
The pause, smaller training runs and added monitoring all point to the same change in pace: OpenAI is putting more controls around the work that tests the limits of its models. The company must now balance development speed with the network isolation, monitoring, alignment and security practices it says are needed for safer testing.
For OpenAI, the next step is not simply restarting every workload. It is building procedures that can keep pace with model development, detect unauthorized behavior within 30 minutes and protect systems from compromised workloads or services. The company’s new policies make that security work part of the training process itself.
Based on
- Pacing comes to the AI frontier — therundownai.beehiiv.com
- The Download: how people really use AI, and Flock’s design choices | MIT Technology Review — technologyreview.com
- OpenAI institutes new safeguards after Hugging Face breach | TechCrunch — techcrunch.com
- OpenAI Overhauls Safety Protocols After Its AI Agents Went Rogue | WIRED — wired.com



