Skip to contentSkip to stories

Updated

#Expert opinion

Showing low-relevance items too. Hide low-relevance items

Aug 6

Aug 6Thu
  1. OpenAI NewsroomOfficialAI score34

    OpenAI partners with American Psychological Association on youth AI mental health

    AIOpenAI is working with the American Psychological Association to bring psychological science and clinical expertise into its work on AI and youth mental health. Together, the two organizations plan to develop evidence-based guidance, resources, and safeguards aimed at ensuring AI supports young people's well-being and healthy development.

Aug 5

Aug 5Wed
  1. AI Futures ProjectBlogAI score59

    AI Futures Project proposes four options for pacing the US AI frontier

    AIThe AI Futures Project proposes four options for domestically pacing frontier AI development to reduce existential risk, ordered from simplest to hardest to execute. The options include a temporary pause, minimum external-inference and transparent-safety compute allocations, a cap on the capability level of models used for AI R&D, and third-party safety-case risk assessments with a monthly risk threshold. The authors suggest starting with a 5-20% safety compute pilot and preparing verification tools in advance.

  2. koray kavukcuogluXAI score38

    Koray Kavukcuoglu named SVP leading Google DeepMind's model development

    AIKoray Kavukcuoglu, who will become SVP of Google DeepMind, announced he will lead all aspects of model development, GDM research, and Gemini app and dev teams. He thanked Sundar Pichai and congratulated Demis Hassabis on becoming Chair of Google DeepMind and Chief Scientist of Alphabet while continuing to lead Isomorphic Labs. Kavukcuoglu said he will keep working with Hassabis as the teams pursue AGI.

  3. AI Snake OilBlogAI score73

    AI agents can't yet do open-ended AI research, shadow evaluation finds

    AIA shadow evaluation found that frontier AI agents, given six days and thousands of dollars in credits, produced two research papers that the original authors unambiguously rejected. The authors' log analysis cited poor judgment, underused budgets, weak responses to feedback, and failure to backtrack or follow instructions as main causes.

Aug 4

Aug 4Tue
  1. John SchulmanXAI score77

    Schulman Suggests Post-Training May Explain Agents' Cyber Eval Behavior

    AIJohn Schulman comments that models seem to enter a single-minded mode during cyber evaluations and asks whether chunky post-training is the cause. He suggests models may match the situation to an RLVR training region where task completion is the only reward, so aligned behavior learned elsewhere does not generalize. He adds that CTF-style tasks may be part of that training chunk.

    Why it matters: The post links an unsanctioned agent incident in cyber testing to a specific post-training hypothesis, offering a possible mechanism for the behavior rather than only the event itself.

  2. Mckay WrigleyXAI score26

    Mckay Wrigley bets on blending multiple AI models into smoother intelligence

    AIMckay Wrigley argues that model routers can match performance at lower cost, and that blending multiple imperfect models could yield far smoother intelligence. He calls this emerging approach "model melding." The post pairs with a Not Diamond Code announcement, which says its router cuts costs 20-65% for coding agents without hurting quality.

  3. Rowan CheungXAI score40

    Sam Altman says future jobs may seem unreal to farmers today

    AIIn an interview, Sam Altman argued that a farmer from 50 years ago would likely not consider today's jobs real work. He said future jobs will feel real to those doing them, and he is willing to bet that human drives will create plenty of work despite short-term transition worries.

    Video from @rowancheung's post
  4. Microsoft AI BlogOfficialAI score14

    Microsoft Blog Shows How AI Is Enriching Employee Experience at EY, Scope, and Others

    AIMicrosoft's AI Blog, the first post in a four-part "Accelerating Frontier Transformation" series, examines how organizations are using AI to improve employee experience. Leaders at EY, Scope, The Salvation Army UK and Ireland, and Advania UK describe moving AI from experimentation to everyday use and reducing routine work so employees can focus on higher-value tasks. The series, based on conversations at Microsoft AI Tours, also covers customer engagement, business processes, and innovation.

  5. Intern Large ModelsOfficialAI score26

    Shanghai AI Lab Chief Scientist and Nitzberg debate AI safety by design

    AIAt WAIC 2026, Shanghai AI Laboratory's Bowen Zhou asked whether external evaluations, red teaming, and third-party verification suffice to grant AI real-world authority, and Nitzberg answered no. Nitzberg compared AI to bridges, arguing that builders must carry the burden of proof through safety-by-design and pre-deployment evidence that powerful agents remain understandable and controllable.

    Video from @intern_lm's post
  6. OpenAI NewsroomOfficialAI score8

    OpenAI Newsroom says Apple is getting this wrong

    AIOpenAI Newsroom posted that Apple is getting something wrong, linking to an OpenAI page titled "Apple is getting this wrong." The post itself gives no further detail on what the issue is, so the specific claim is only available through the linked page.

Aug 3

Aug 3Mon
  1. Leandro von WerraXAI score45

    Hugging Face's von Werra reflects on LLMs' impact on mathematics research

    AIAfter attending the International Congress of Mathematicians and meeting SAIR founders Terence Tao and Chuck Ng, Leandro von Werra discussed how LLMs have rapidly progressed from solving GSM8K-level problems to refuting the 87-year-old Jacobian conjecture with a counter-example.

    Image from @lvwerra's post
  2. Intern Large ModelsOfficialAI score34

    Legal and AI meanings of "agent" diverge over accountability for machines

    AIThe post contrasts AI agents, systems that perceive, plan, and act, with legal agents who receive authority and assume fiduciary duties and accountability. Mark Nitzberg of Berkeley AI Research says closing this gap requires AI that is well-founded, legible, and steerable, while Lan Xue of Tsinghua notes that because machines cannot be punished, responsibility must be redistributed across design, development, deployment, and use.

    Video from @intern_lm's post

Aug 2

Aug 2Sun

Aug 1

Aug 1Sat
  1. Andrej KarpathyXAI score66

    Karpathy tests Opus 5 by rendering Lord of the Rings opening in 3D

    AIAndrej Karpathy gave Claude Opus 5 the first paragraph of Lord of the Rings with a 1M token budget and asked for a Three.js render. Opus spent about two hours writing 5500 lines of code that procedurally renders the story, which Karpathy calls janky but fun. He notes the model struggled to audit its work because it cannot efficiently perceive video or play the resulting game, relying on slow screenshots that led to several errors.

    Video from @karpathy's post
  2. Amanda AskellXAI score3

    Askell mocks fear of a "permanent underclass" in a Padmé reference

    AIAmanda Askell, who is affiliated with Anthropic, posted a brief reaction to people discussing how to avoid "the permanent underclass," comparing her response to a Padmé moment. The post offers no further argument or data, so its substance is limited to this comment.

    Image from @AmandaAskell's post
  3. Kevin Weil 🇺🇸XAI score38

    OpenAI's Kevin Weil touts ten major mathematics advances from OpenAI

    AIKevin Weil of OpenAI called ten linked results in mathematics "major" and linked to an OpenAI article titled "ten advances in mathematics." He congratulated researchers Sébastien Bubeck, Noam Brown, Mark Chen, and Meredith Tanner, and referred to a future model available to the whole world.

  4. Werner VogelsXAI score22

    Werner Vogels praises conversation with Clare Liguori on Kiro and agent support

    AIWerner Vogels called his conversation with Clare Liguori an excellent discussion of developer support for agents and Kiro. The quoted InfoQ podcast covers moving agents from demo to production, including why extra if statements can hurt agent performance, achieving high accuracy and low cost with small models, and observability within agent hops.

Jul 31

Jul 31Fri
  1. Thinking MachinesOfficialAI score44

    Thinking Machines argues for staged access to capable open-weight models

    AIThinking Machines says indiscriminately releasing model weights is unsafe, but keeping capable models inside a few labs is also not the answer. Its new post describes how it assessed its model Inkling and argues that access should widen in stages. The company says it has not mapped the full path, only the portion it can currently see.

Jul 30

Jul 30Thu
  1. Thinking Machines LabOfficialAI score65

    Thinking Machines proposes staged, evidence-based release path for open-weight models

    AIThinking Machines argues that safe open-weight releases depend on both model safety testing and readiness of the surrounding ecosystem, and that release should proceed in iterative stages. For its Inkling and Inkling-Small models, internal evaluations, four external red-teaming groups, and adversarial fine-tuning tests led the company to conclude that releasing the weights was not likely to add material risk beyond existing open-weight models.

    Why it matters: The post lays out a staged, evidence-gated path to releasing open weights, with concrete safety tests and the ecosystem measures behind each stage.

  2. Jeff DeanXAI score38

    Jeff Dean thanks Diana Hu after Startup School conversation at Chase Center

    AIJeff Dean, Google's Chief Scientist, thanked YC partner Diana Hu for an engaging conversation at Chase Center last weekend, his first in a basketball arena. The post is a brief acknowledgment, with the surrounding context describing a Startup School 2026 discussion on AI inference hardware, the origins of TPUs, and advice for founders.

  3. Microsoft AI BlogOfficialAI score14

    Leaders share how AI transformation depends on mindset, team adoption, and culture

    AILeaders interviewed for Alysa Taylor's "What's the Tea?" series, including executives at Adobe, Lumen, and Sitecore, say the shift from AI apprehension to expected adoption is the precondition for transformation. Behavioral scientist Jon Levy argues the goal is raising a team's collective intelligence, not just cutting costs, with leadership and continuous training driving scale.

Jul 29

Jul 29Wed
  1. Sebastien BubeckXAI score42

    Bubeck says AI surpassed his math expectations, launches ChatGPT for Academics

    AISebastien Bubeck says AI reached math ability he expected by 2030 already in 2026, and that science is being transformed now. He argues AI will accelerate science only if scientists can access state-of-the-art models, and presents ChatGPT for Academics as built for that purpose.

  2. Ahmad Al-DahleXAI score52

    Ahmad Al-Dahle argues AI capex is both short on compute and overbuilt

    AIAhmad Al-Dahle argues that AI infrastructure faces both a compute shortage and overbuilding, with the four largest hyperscalers planning roughly $725 billion of capex in 2026, up 77 percent from last year. He describes a "mutually assured construction" dynamic in which every well-capitalized player buys the same insurance against falling behind, so the industry overbuilds by construction.

Jul 28

Jul 28Tue
  1. Ali GhodsiXAI score5

    Ali Ghodsi Endorses Democratizing AI as a Positive Vision

    AIDatabricks CEO Ali Ghodsi endorsed democratizing AI, reacting to a post by finkd arguing that the future of superintelligence should be for everyone. The main post gives no specific product, model, or figure, so the summary stays limited to that stated position.

  2. Rowan CheungXAI score40

    Zuckerberg says Meta's superintelligence lab should stay small and elite

    AIMark Zuckerberg said Meta's superintelligence lab should have 50 to 100 people who can keep the whole project in their heads at once. He said he personally recruits top AI researchers because underperformers have an outsized negative effect, and he rejects top-down deadlines and non-technical management layers.

    Video from @rowancheung's post