Skip to content
Tech News
← Back to articles

OpenAI's New Reasoning Technique Alarms AI Safety Experts

read original get Nick Bostrom "Superintelligence" (book) → more articles
Why This Matters

OpenAI's introduction of the Astra model with 'recurrent depth' or 'opaque recurrence' raises significant concerns about AI safety and transparency. While this technique could enhance reasoning capabilities, it risks making AI decision processes less monitorable, potentially complicating safety oversight and regulatory efforts. The development underscores the urgent need for balanced innovation and safety measures in AI research to prevent loss of control and accountability.

Key Takeaways
Worth a Look

Nick Bostrom "Superintelligence" (book) — If the debate over unmonitorable chains of thought has you curious, Bostrom's Superintelligence is the classic deep dive into AI alignment and control problems that safety researchers keep referencing. It's a great way to understand why experts get rattled when a model's reasoning becomes opaque.

See Nick Bostrom "Superintelligence" (book) on Amazon → Affiliate link — we may earn a commission on purchases, at no extra cost to you. Product picked by AI based on this article; it is not a tested recommendation.

An anonymous reader quotes a report from TechCrunch: OpenAI's new Astra model will use a reasoning technique called "recurrent depth" that allows it to operate outside of the sequential thinking that characterizes most reasoning models, The Information reported on Tuesday. This technique, also called "opaque recurrence," will likely make the model's chain of thought more difficult to monitor -- and that has AI safety experts rattled. While Astra's use of the technique is reportedly limited, its emergence has still raised significant concerns among AI safety experts. "I am extremely concerned by the reporting that Astra uses opaque recurrence," wrote Redwood CEO Buck Shlegeris in a post after the news broke. "I don't know whether Astra is much less CoT monitorable than previous models. But if OpenAI pushes this technique further, they'll have the option to massively increase the recurrence and totally destroys CoT monitorability." Longtime AI safety advocate Zvi Mowshowitz also weighed in and wrote that laws might be necessary to prevent a "race to the bottom" among AI labs. "The technique is playing with fire, risking a taboo that OpenAI and Anthropic have fought to establish that we work hard to maintain Chain of Thought faithfulness and monitorability for as long as we can," Mowshowitz wrote. "More intensive use of such techniques would probably damage monitorability." [...] In a post responding to the news, Redwood Research chief scientist Ryan Greenblatt said opaque reasoning could easily scale faster than conventional chain-of-thought reasoning, effectively removing all reasoning from visible channels. "My biggest concern is that a natural progression from here would involve scaling up the opaque reasoning to the point where the model reasons entirely or almost entirely in latent space," Greenblatt wrote. "I hope it isn't too late to avoid the most concerning architectures and that OpenAI will stop here." Astra's use of the technique appears limited. "OpenAI has worked to preserve and utilize chain-of-thought monitoring since our very first reasoning models," wrote OpenAI chief scientist Jakub Pachocki. "It's a core goal of our current research program."

Read more of this story at Slashdot.