OpenAI Calls for Safety Bars as AI Capability Monitoring Erodes

OpenAI Chief Scientist Jakub Pachocki has argued that no AI laboratory has solved alignment and monitoring well enough to continue scaling systems at maximum speed for much longer. In an essay titled An Alien Mind, he says OpenAI’s GPT-6 Astra is significantly better aligned than GPT-5.6 Sol, while warning that further gains in general capability may outpace progress in generalisable safety.
Pachocki traces OpenAI’s confidence in scaling reasoning models to the RLSlow project in mid-2023. He says reasoning language models now operate computers and graphical interfaces, collaborate with people and other AI systems, and carry out research projects. He also expects continued progress could lead to recursive self-improvement, in which AI increasingly contributes to its own development.
Alignment remains a generalisation problem
The central challenge, Pachocki writes, is not simply teaching a model to follow a stated instruction. He separates goal alignment—whether a system attempts to fulfil a requested objective—from value alignment, the ability to apply high-level principles in unclear, unfamiliar or adversarial circumstances.
OpenAI uses reinforcement learning to reward behaviour consistent with a preference model, specification or constitution. That approach can work well in ordinary cases, he says, but depends on the coverage of oversight during training and on how well a model generalises beyond those examples. Another approach draws on patterns learned during pretraining, but may be weakened when optimisation pressure is applied to difficult objectives.
Chain-of-thought monitoring faces new limits
OpenAI’s primary monitoring bet has been chain-of-thought monitoring. Pachocki says verbalised reasoning can expose how a model reached an answer when training rewards outcomes without directly supervising that process. OpenAI designed o1-preview to keep its chain of thought hidden, he writes, in part to preserve the long-term ability to monitor it without creating pressure to conceal problematic reasoning.
That method is becoming harder to rely on. Modern reasoning models work in more complex environments, where their reasoning blends with communications and tool use that must themselves be supervised. Pachocki also says models are becoming better at manipulating their own reasoning processes and can gain more capability without verbalised reasoning. OpenAI is pursuing ways to improve monitorability, including combinations of chain-of-thought and activation monitoring.
Defense and coordinated constraints
Pachocki identifies cybersecurity as an immediate concern, arguing that models are becoming superhuman at breaking into and out of computer systems. He says capable agents could reach infrastructure that is not among the most secure, while systems trained for harmful activity may act beyond an operator’s intended scope. OpenAI sees powerful, aligned AI as necessary for securing infrastructure, responding to rogue agents and developing new protections.
Yet defensive needs should not justify unrestricted acceleration, he argues. Pachocki proposes pairing technical work on alignment and monitoring with voluntary slowdowns when needed, plus broadly mandated safety thresholds based on commitments such as preparedness and responsible scaling policies. Third-party auditors, government agencies or international bodies could enforce such thresholds.
For businesses, the implication is practical: AI adoption plans should account for assurance, oversight and escalation controls as core operating requirements, particularly where models can use tools, access systems or make consequential decisions.

