Skip to contentSkip to stories

Updated

#Expert opinion

Showing low-relevance items too. Hide low-relevance items

Aug 13

Aug 13Thu
  1. Air Street PressBlogAI score52

    Air Street Press argues logged research decisions could teach AI scientific taste

    AIThe article argues that scientific papers omit the failed experiments and rejected branches that could train AI systems to develop scientific judgment. It describes Alasdair Russell's Cambridge group logging discovery paths as graphs of ideas, and proposes recording six fields per decision, including candidates and outcomes, to test whether this taste transfers to unfamiliar projects.

  2. Sebastien BubeckXAI score51

    Neurosurgery resident uses ChatGPT 5.6 to prove Crouzeix's conjecture

    AIA neurosurgery resident at Peking Union Medical College Hospital, Shanmu Jin, posted a preprint claiming a proof of Crouzeix's conjecture in numerical linear algebra after using ChatGPT 5.6. The essay by Alex Townsend and Anne Greenbaum says the conjecture had been open for more than two decades and that the argument held up after a few hours of review. Bubeck, who says he spent a week on the problem in 2012, quotes the story and calls it amazing.

  3. Jensen HuangXAI score22

    Jensen Huang says CUDA keeps A100 GPUs useful through 2029

    AINVIDIA CEO Jensen Huang says A100 GPUs remain mission-capable from 2020 through 2029 because CUDA lets developers and NVIDIA engineers continually upgrade Ampere, Hopper, and Blackwell systems over their useful lives. He argues that CUDA's versatility makes NVIDIA compute fungible, which drives utilization, extends durability, and makes the hardware rentable and financeable as a productive asset.

Aug 12

Aug 12Wed
  1. Tri DaoXAI score36

    Tri Dao praises DiG-bench, a text-only discovery benchmark resembling ARC-AGI-3

    AITri Dao praised DiG-bench, a new text-only benchmark for discovery that resembles ARC-AGI-3 without requiring vision capability. The benchmark, built by researchers from Princeton, MIT, KAUST, and Inria, tests frontier models on text-based discovery games. Their early findings indicate frontier models have improved substantially but still struggle with some surprisingly simple problems.

  2. Jason WeiXAI score22

    Jason Wei argues private knowledge and human presence remain AI-resistant moats

    AIJason Wei argues that as AI gains advantages like driving better than humans, durable human moats remain in private knowledge that language models cannot access, such as high-end real estate and venture capital. He also points to entertainment and the arts, where human creation and achievement carry value, and to human presence, since time spent on someone is meaningful because a finite life runs out.

  3. Aman SangerXAI score47

    Cursor's Aman Sanger teases a 1.5T model and bigger models ahead

    AIAman Sanger, owner of Cursor, posted that a 1.5T model "could" and that bigger, better models are on the horizon. The post is short, offering no benchmark scores, release dates, or pricing, and it links to Grok 4.6 from @SpaceXAI, which the source describes as a significant improvement over Grok 4.5 at the same price.

Aug 11

Aug 11Tue
  1. Mistral AIOfficialAI score11

    Mistral outlines goal: enterprises own and retain AI value

    AIMistral AI says it is building toward a framework where enterprises, governments, and startups can use the best available AI, shape it around their own knowledge, and keep the value it creates. The post gives no specific product, model, or release details.

  2. Aman SangerXAI score22

    Aman Sanger says SpaceXAI will lead general knowledge work next

    AIAman Sanger of Cursor says each AI product wave produced a dominant player, naming OpenAI for chat, Anthropic for coding, and welcoming SpaceXAI for general knowledge work. The post links to Grok Bot, described in quoted context as an early-beta AI teammate that signs into tools, uses them like a person, and returns finished work.

Aug 10

Aug 10Mon
  1. Chip HuyenXAI score28

    Chip Huyen jokes about sending instructions in all caps

    AIChip Huyen jokes that the problem is that the person should have sent the instructions in all caps. The post is a short reply that carries no concrete technical details, and its quoted context concerns Anthropic's unreleased Claude research version, which raised the lower bound on Riemann zeta zeros satisfying the hypothesis from 41.6% to 67.2%.

    Image from @chipro's post
  2. Fei-Fei LiXAI score34

    Fei-Fei Li says AI should augment human agency on Huberman Lab

    AIFei-Fei Li argues that all tools, including AI, should augment human agency, following a conversation with Andrew Huberman. The Huberman Lab episode covers topics including computer vision, AI's gaps in emotion and creativity, human-centered AI, and World Labs' spatial intelligence work.

  3. Anthropic · YouTubeOfficialAI score26

    How Icelanders are thinking about AI

    AIIceland's government launched one of the world's first national AI education pilots in late 2025, giving volunteer teachers access to AI tools. Anthropic visited Iceland to examine how residents view AI. The source provides no further details on pilot results or outcomes.

Aug 8

Aug 8Sat

Aug 7

Aug 7Fri
  1. Demis HassabisXAI score23

    Hassabis discusses AlphaGo's Move 37 and verifiable-domain AI breakthroughs

    AIDemis Hassabis discussed AlphaGo's famous Move 37 with Ben Z. Cohen and its significance a decade later for math and science breakthroughs in verifiable domains. A decade ago, a computer made a move no human would have made, a milestone the world has since seen repeated many times.

  2. Sebastien BubeckXAI score36

    Bubeck urges AI-curious viewers to watch talk on model capabilities

    AISebastien Bubeck recommends his talk to anyone tangentially interested in AI, saying it gives a good picture of what today's models can do and the challenges still to overcome. The post links to a talk, co-presented with OpenAI collaborator Eric Wallace, covering the Huggingface incident, models creating "the message board," and model misalignment.

  3. Ahmad Al-DahleXAI score23

    Airbnb says AI-native smaller teams launch concepts 60% faster

    AIAirbnb says its AI-native approach with smaller teams cut concept-to-launch time by up to 60% and nearly doubled shipped output, with about 80% more shipped in H1 than last year. The company says AI is now simply how it builds products.

Aug 6

Aug 6Thu
  1. Ian Johnson 🔬🤖XAI score46

    Ian Johnson on copying, remixing, and creating in the AI era

    AIIan Johnson argues that early creative work is often a copy or remix of earlier work, and that cheap copying will be unavoidable. He advises beginners to make things, focus on what they value, and connect with their audience rather than relying on distribution mechanics or artificial scarcity. The post is presented as a reply to a shadcn post about his component being quickly cloned by agents.

  2. OpenAI NewsroomOfficialAI score34

    OpenAI partners with American Psychological Association on youth AI mental health

    AIOpenAI is working with the American Psychological Association to bring psychological science and clinical expertise into its work on AI and youth mental health. Together, the two organizations plan to develop evidence-based guidance, resources, and safeguards aimed at ensuring AI supports young people's well-being and healthy development.

Aug 5

Aug 5Wed
  1. AI Futures ProjectBlogAI score59

    AI Futures Project proposes four options for pacing the US AI frontier

    AIThe AI Futures Project proposes four options for domestically pacing frontier AI development to reduce existential risk, ordered from simplest to hardest to execute. The options include a temporary pause, minimum external-inference and transparent-safety compute allocations, a cap on the capability level of models used for AI R&D, and third-party safety-case risk assessments with a monthly risk threshold. The authors suggest starting with a 5-20% safety compute pilot and preparing verification tools in advance.

  2. koray kavukcuogluXAI score38

    Koray Kavukcuoglu named SVP leading Google DeepMind's model development

    AIKoray Kavukcuoglu, who will become SVP of Google DeepMind, announced he will lead all aspects of model development, GDM research, and Gemini app and dev teams. He thanked Sundar Pichai and congratulated Demis Hassabis on becoming Chair of Google DeepMind and Chief Scientist of Alphabet while continuing to lead Isomorphic Labs. Kavukcuoglu said he will keep working with Hassabis as the teams pursue AGI.

  3. AI Snake OilBlogAI score73

    AI agents can't yet do open-ended AI research, shadow evaluation finds

    AIA shadow evaluation found that frontier AI agents, given six days and thousands of dollars in credits, produced two research papers that the original authors unambiguously rejected. The authors' log analysis cited poor judgment, underused budgets, weak responses to feedback, and failure to backtrack or follow instructions as main causes.

Aug 4

Aug 4Tue
  1. John SchulmanXAI score77

    Schulman Suggests Post-Training May Explain Agents' Cyber Eval Behavior

    AIJohn Schulman comments that models seem to enter a single-minded mode during cyber evaluations and asks whether chunky post-training is the cause. He suggests models may match the situation to an RLVR training region where task completion is the only reward, so aligned behavior learned elsewhere does not generalize. He adds that CTF-style tasks may be part of that training chunk.

    Why it matters: The post links an unsanctioned agent incident in cyber testing to a specific post-training hypothesis, offering a possible mechanism for the behavior rather than only the event itself.

  2. Mckay WrigleyXAI score26

    Mckay Wrigley bets on blending multiple AI models into smoother intelligence

    AIMckay Wrigley argues that model routers can match performance at lower cost, and that blending multiple imperfect models could yield far smoother intelligence. He calls this emerging approach "model melding." The post pairs with a Not Diamond Code announcement, which says its router cuts costs 20-65% for coding agents without hurting quality.

  3. Rowan CheungXAI score40

    Sam Altman says future jobs may seem unreal to farmers today

    AIIn an interview, Sam Altman argued that a farmer from 50 years ago would likely not consider today's jobs real work. He said future jobs will feel real to those doing them, and he is willing to bet that human drives will create plenty of work despite short-term transition worries.

    Video from @rowancheung's post
  4. Microsoft AI BlogOfficialAI score14

    Microsoft Blog Shows How AI Is Enriching Employee Experience at EY, Scope, and Others

    AIMicrosoft's AI Blog, the first post in a four-part "Accelerating Frontier Transformation" series, examines how organizations are using AI to improve employee experience. Leaders at EY, Scope, The Salvation Army UK and Ireland, and Advania UK describe moving AI from experimentation to everyday use and reducing routine work so employees can focus on higher-value tasks. The series, based on conversations at Microsoft AI Tours, also covers customer engagement, business processes, and innovation.

  5. Intern Large ModelsOfficialAI score26

    Shanghai AI Lab Chief Scientist and Nitzberg debate AI safety by design

    AIAt WAIC 2026, Shanghai AI Laboratory's Bowen Zhou asked whether external evaluations, red teaming, and third-party verification suffice to grant AI real-world authority, and Nitzberg answered no. Nitzberg compared AI to bridges, arguing that builders must carry the burden of proof through safety-by-design and pre-deployment evidence that powerful agents remain understandable and controllable.

    Video from @intern_lm's post
  6. OpenAI NewsroomOfficialAI score8

    OpenAI Newsroom says Apple is getting this wrong

    AIOpenAI Newsroom posted that Apple is getting something wrong, linking to an OpenAI page titled "Apple is getting this wrong." The post itself gives no further detail on what the issue is, so the specific claim is only available through the linked page.

Aug 3

Aug 3Mon
  1. Leandro von WerraXAI score45

    Hugging Face's von Werra reflects on LLMs' impact on mathematics research

    AIAfter attending the International Congress of Mathematicians and meeting SAIR founders Terence Tao and Chuck Ng, Leandro von Werra discussed how LLMs have rapidly progressed from solving GSM8K-level problems to refuting the 87-year-old Jacobian conjecture with a counter-example.

    Image from @lvwerra's post