
According to The Decoder, OpenAI evaluates the upcoming Astra model as the company's first system with "critical" cyber capabilities. To control it, OpenAI plans to monitor the model's chain of thought.
The problem described in the publication's synopsis is that observing the chain of thought is already considered an unreliable reflection of the model's actual decisions. According to The Decoder, Astra's new architecture makes an even larger portion of reasoning inaccessible for reading.
This implies a potential weakening of control coinciding with an increase in system capabilities. However, the presented materials contain only a metadata synopsis, not the full text of the publication or independent confirmation, so details of OpenAI's approach and the scale of risk remain unclear.
editorial commentary
Why it matters
The probable implication is that safety verification for Astra may rely more heavily on methods not reducible to reading the chain of thought. The nearest observable signal is OpenAI's public explanations regarding the model's status, criteria for "critical" cyber capabilities, and additional control measures. Significant uncertainty remains because only The Decoder's metadata synopsis is available, without a primary statement from OpenAI or a full publication.