Skip to content
Read the original: Sebastian Raschka· Published 38/100AI score38/100

Raschka Says OpenAI Astra's Looped Transformer Is Not a Big Deal

Original titleA lot of hype around OpenAI's Astra model here on my timeline today. Apparently, this goes back to a new article from The Information, wh...

AISummary

Sebastian Raschka argues that the looped transformer approach attributed to OpenAI's Astra is a minor architectural tweak, not a major breakthrough.

He explains that Nanbeige4.2-3B reuses its 22-layer stack twice, effectively doubling depth without adding weights but roughly doubling compute, and that the idea traces back to the Mixture-of-recursions NeurIPS paper.

He adds that layer reuse does not inherently hide chain-of-thought, though it could shift more computation into latent activations.

Read the original x.com

Source: Sebastian Raschka · x.comPublished · added here