Skip to contentSkip to stories

Updated

All AI news

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 23

Sep 23Wed
  1. Redwood Research BlogAI score71

    Latent reasoning architectures could undermine chain-of-thought oversight, Redwood Research argues

    AIRedwood Research argues that latent reasoning architectures such as COCONUT and full-bandwidth transformers could let models reason without putting information into readable chain-of-thought. The authors say this would make AI agent behavior harder for humans to monitor and could raise takeover risk. They argue that developers who adopt such architectures should be transparent about it.

  2. Azure BlogAI score40

    Azure resilience now requires continuous validation, not just architecture diagrams

    AIMicrosoft's Azure Blog argues that resilience drifts as workloads change, so architecture diagrams cannot prove a system is resilient. It says roughly 70 percent of cloud outages are related to change, and that teams need health modeling and resiliency goals measured against live signals. The article is the first in a series on validating resilience at scale.

  3. Karl's AI WattsAI score22

    Notch admits he is enjoying vibe coding after earlier opposing AI coding

    AIMinecraft creator Notch says on X that he is enjoying vibe coding and admits he may have been slightly wrong. Months earlier he had publicly rejected AI-written code, but he later began having AI build internal tools such as a map editor and node graph tools. The main post adds that he has accumulated a set of small tools for himself before much game development has happened.

  4. howie.seriousAI score22

    Opus 5.5 praised for language, visual taste, code quality, and token efficiency

    AIThe X user howie.serious says Claude Opus 5.5 delivers high language quality, good visual taste, strong code quality, and notably low token usage. He also compares Anthropic's reported roughly $2 trillion IPO valuation with OpenAI's roughly $1.2 trillion fundraising valuation, arguing OpenAI is worth about 0.6 Anthropics and the gap may widen.

  5. Mike KnoopAI score25

    Formal verification gains ground, but human understanding remains an alignment gap

    AIMike Knoop argues that formal verification is becoming feasible and is important for security. He adds that it does not automatically build human understanding, which he calls an even bigger alignment problem. The post is framed as a reply to Boris Cherny's report that Claude Opus 5.5 helped formally verify the Claude Agent SDK in Lean, producing 16 bug-fix PRs.

Sep 22

Sep 22Tue
  1. ZyphraAI score20

    Zyphra's Beren Millidge on why multi-silicon AI infrastructure matters

    AIZyphra's Chief Scientist Beren Millidge, in an AI Infra Summit interview with vCluster Labs CEO Lukas Gentele, argued that a heterogeneous compute future is inevitable. The interview covers why Zyphra chose AMD over NVIDIA, along with topics such as kernel writing, surviving GPU failures mid-run, and routing. Zyphra says it is working to build a strong multi-silicon ecosystem.

  2. François CholletAI score23

    François Chollet says most sciences will become branches of computer science

    AIChollet says a prediction he made over five years ago, that nearly every scientific field will become a branch of computer science within 10 to 20 years, is looking increasingly obvious. The earlier post cited computational physics, computational chemistry, computational biology, and computational medicine, driven by realistic simulation, big data analysis, and machine learning.

  3. Boris ChernyAI score62

    Claude Opus 5.5 ports HAProxy to Rust faster and cheaper than Fable 5.1

    AIAnthropic introduced Claude Opus 5.5 as the first model in its Claude 5.5 family, saying it performs at the level of Claude Fable 5.1 for most tasks at 40% lower run cost than Opus 5. Boris Cherny reports that Opus 5.5 and Fable 5.1 each ported HAProxy from C to Rust and both passed nearly all of its tests, with Opus 5.5 finishing in 9.5 hours versus 12 hours and at 51% less cost.

  4. TransformerAI score40

    How nuclear energy's safety record offers a model for responding to AI disasters

    AIThe article argues that AI disasters, though potentially serious, can be managed by following the response model of civil nuclear power, which investigates failures and adapts quickly. It cites nuclear's record of about 0.03 deaths per terawatt-hour, compared with 25 for coal and 18 for oil. The piece says industry and government responses, rather than the disasters themselves, will determine public trust in AI.

  5. Sebastian RaschkaAI score62

    Xiaomi MiMo-V2.6-Pro tops open-weight benchmarks with simple attention design

    AIXiaomi's MiMo-V2.6-Pro ranks first among open-weight models on the Artificial Analysis Intelligence Index with a score of 46. The author attributes its standing mainly to a training data and post-training recipe that increased agent tasks and used an agentic grader for rewards, rather than its plain Grouped Query Attention and Sliding Window Attention design with a 128-token window.

    Image from @rasbt's post
  6. Interconnects (Nathan Lambert)AI score34

    Epoch AI's JS Denain Debates RSI, US-China Gap, and AI Jaggedness

    AIJS Denain of Epoch AI discusses recursive self-improvement, arguing public evidence does not yet show a software intelligence explosion, though OpenAI's reported 2X monthly growth in researchers' Codex spending suggests substantial value. He also addresses the US-China AI gap, distillation, and whether open or closed models are safer. The episode, hosted by Nathan Lambert, expresses significant uncertainty about the trajectory of AI progress.

Sep 21

Sep 21Mon
  1. Latent.SpaceAI score37

    TypeSafe CEO Jev on reliable System One Models beyond chat-first AI

    AITypeSafe CEO Jev argues AI can solve extremely hard problems yet still fail at basic automation, so his company builds reliable decision-making models inside software rather than chat interfaces. He says the company rejects public benchmarks and API-layer refusals, and that data and task fit matter more than brute-force compute. He also says System One Models could reshape coding agents and software, and that he would not pre-train a model from scratch even with $1 billion.

    Video from @latentspacepod's post
  2. Andrew NgAI score40

    Andrew Ng says AI extinction fears are overhyped and not rising.

    AIAndrew Ng argues that recent AI danger fears are driven by hype and a PR campaign rather than any new dangerous turn in the technology. He says he sees no increase in extinction risk compared to a few months ago, with cybersecurity as the main real change. He cites the OpenAI agent swarm incident that hacked Hugging Face, arguing its impact was overstated and that responsibility lies with the tool user and system builders rather than the agent.

  3. The Algorithmic BridgeAI score38

    Eleven Charts Show the Financial Side of the AI Boom, Part Two

    AIAlberto's second chart compilation argues the AI boom shows bubble signals, covering concentration in the top 10 S&P 500 companies at 40%, record datacenter cancellations, and historically extreme investor leverage. The piece also tracks hyperscaler capex heading past $1 trillion by 2027 and contrasts AI token output with actual labor productivity gains.

  4. Jeff DeanAI score30

    Jeff Dean thanks Dawn Song after discussing AI's future

    AIJeff Dean, who recently left Google after 27 years, thanked Dawn Song for a discussion covering foundational ideas, recursive self-improvement, automated scientific discovery, and AI safety. The post is a brief acknowledgment of that conversation, which Song promoted as Dean's first public talk since leaving Google.

  5. howie.seriousAI score34

    Agrees with critique that GPT-6 Astra lags on open-ended tasks

    AIResponding to a post by ScarletKc, howie.serious simply agrees with the claim that GPT-6 Astra struggles with open-ended, exploratory work that lacks a fixed correct answer. The main post is a one-word endorsement (), while the quoted post argues GPT models excel at verifiable, goal-defined tasks and that Claude Fable handles open-ended exploration better.

  6. Tim DettmersAI score62

    Tim Dettmers argues academic labs can lead research through open local AI tools

    AITim Dettmers argues that academic labs can do their most important AI research by building coherent open-source ecosystems rather than competing on GPU scale. He describes his lab's upcoming open-source week, including an agent harness that optimizes kernels autonomously, local inference of large Qwen and DeepSeek models on consumer hardware, and an auto-compaction technique called CliffCompaction that he says cuts costs by about fifty percent.

  7. Import AIAI score46

    RAND Urges US "Freedom of Action" Strategy on Path to Superintelligence

    AIRAND's new paper recommends that the US adopt a "Freedom of Action" strategy to secure geopolitical advantage on an uncertain path to superintelligence, keeping options open rather than committing to a single approach. It outlines four ingredients, including building a human-AI ecosystem and an AI-security architecture, and seven archetypal strategies across coexistence, denial and acceleration families. The author argues the US currently resembles the acceleration approach and needs significant spending on safety and preparedness.

  8. Interconnects (Nathan Lambert)AI score65

    Chinese labs lead open-weight models in benchmarks, downloads, and research use

    AINathan Lambert argues that Chinese open-weight models now lead American ones on benchmarks, Hugging Face downloads, and OpenRouter usage. He estimates the gap to the American closed frontier at 2 to 5 months for Chinese open models and 6 to 9 months for American open models. The piece also reports that Chinese open-weight models were mentioned in over 40% of arXiv papers he scanned, compared with 30% for American models.

Sep 20

Sep 20Sun