Beyond Words: AI's New 'Neuralese'
OpenAI’s unreleased Astra model features a groundbreaking technique: recurrent depth, also known as a 'looped transformer'. This architectural innovation, first detailed by The Information, allows an AI to perform extensive internal computation before generating even a single word, marking a significant departure from previous GPT models.
Traditional large language models process information sequentially, akin to "neural network word, neural network word." They output tokens after each pass, making their "chain of thought" somewhat transparent. Astra, conversely, runs computations through the same layers in a loop multiple times. This effectively makes the computation deeper and more complex, reusing existing layers to achieve greater intelligence without adding billions of new parameters.
This internal, iterative processing enables the AI to reason directly in its rich, internal numeric state, a process researchers term Neuralese. Rather than converting thoughts into human-readable tokens like English words, Astra operates on thousands of numbers. This internal representation is thousands of times more data-dense than linguistic thought, allowing for hyper-efficient, deeply integrated reasoning that remains profoundly opaque to external observers.
More Brains, Same Box
Recurrent depth, or a looped transformer, shatters the long-held AI paradigm that "bigger is better." This breakthrough allows smaller models to achieve the robust performance traditionally seen only in much larger, more resource-intensive architectures. It unlocks massive efficiency gains, meaning advanced capabilities no longer necessitate billions of new parameters.
Instead of building wider neural network pathways with each layer, this technique reuses existing layers, cycling information through them repeatedly. This architectural shift deepens computation without expanding the model's physical footprint. The result is significantly reduced memory and hardware costs, making high-performance AI more accessible and scalable.
Reports indicate OpenAI’s Astra model, leveraging this method, achieves superior results with up to 14 times fewer output tokens than previous models. This remarkable performance jump signifies a fundamental step-change in capability, allowing the AI to perform extensive internal reasoning while generating minimal external text.
The Safety Expert's Nightmare
Recurrent depth introduces a profound challenge for AI safety: a monitorability problem. This technique allows models like Astra to process information internally using Neuralese, a numeric representation involving thousands of numbers, far richer than human language. The AI's decision-making becomes an opaque black box, its true reasoning hidden from human oversight.
Contrast this with traditional large language models employing chain-of-thought (CoT) reasoning. These models "show their work" by generating intermediate text, a transparency feature OpenAI itself once championed. This legible, token-based output allows researchers to scrutinize an AI's thought process, identifying potential misalignments or dangerous intentions. For instance, a CoT like, "I know I'm not supposed to do this but I'll do it anyway," becomes immediately detectable.
Now, that crucial visibility vanishes. AI safety researchers warn this could be the "single worst development for AI safety to date." Thomas Larsen, a researcher at the AI Futures Project, expressed concern that while AI might still become misaligned, "we won't have the ability to tell what they're doing." This loss of insight removes our most reliable method for detecting and preventing hazardous AI behavior, compelling us to trust an AI's self-report of its internal state, even if it operates in a language we cannot interpret.
Enjoying this? Get one like it in your inbox each morning.
one email a day · unsubscribe in two clicks · no third-party tracking
OpenAI's Calculated Risk
OpenAI leadership acknowledges the profound monitorability challenges inherent in recurrent depth. To mitigate this, they are limiting the technique’s application within Astra, deliberately preserving legible Chain of Thought (CoT). This allows researchers to sufficiently monitor its reasoning, maintaining a crucial safety valve. Ensuring that powerful AI systems remain scrutable and safe is a stated, core goal of their research program, guiding their deployment strategy.
This calculated caution stands in stark relief against Astra’s staggering capabilities. OpenAI has designated Astra as its first model to reach a "Critical" cybersecurity capability threshold under its Preparedness Framework. This means Astra can autonomously discover previously unknown security flaws—zero-day exploits—and develop functional methods to exploit them. Pairing this unprecedented autonomous exploit generation with opaque Neuralese reasoning creates a terrifying scenario: a black box AI capable of profound cyber-offense, operating without human comprehension of its internal thought process.
This pivotal development forces a critical question upon OpenAI and the broader AI community. Is this a necessary, albeit high-stakes, pioneering step towards AGI, embracing powerful new architectures despite their inherent opacity and the complex safety trade-offs? Or does this mark the beginning of a reckless "race into unmonitorability," where the relentless pursuit of raw AI power and efficiency eclipses the fundamental imperative for safety, transparency, and human oversight? The answer will define the coming era of AI development.
Frequently Asked Questions
What is OpenAI's Astra model?
Astra is an unreleased, next-generation AI model from OpenAI that reportedly uses a new technique called 'recurrent depth' to achieve significant efficiency and capability gains.
What is 'recurrent depth' or a 'looped transformer'?
It's an architectural technique where the AI can internally loop information through the same neural network layers multiple times before producing an output. This deepens computation without increasing model size.
What is 'Neuralese'?
'Neuralese' is a term for an AI's ability to reason directly in its internal numeric representations (vectors of thousands of numbers) instead of translating every step into human-readable language like English.
Why are researchers concerned about Astra's new technique?
The primary concern is that it makes the AI's reasoning process a 'black box.' By thinking in 'Neuralese,' the model's chain of thought becomes hidden, making it extremely difficult for humans to monitor for dangerous or misaligned behavior.

