Skip to contentSkip to stories

Updated

All AI news

Showing low-relevance items too. Hide low-relevance items

Jun 10

Jun 10Wed

Jun 9

Jun 9Tue
  1. Andrej KarpathyAI score65

    Karpathy Calls Claude Fable 5 a Major Step Forward for Long Tasks

    AIAndrej Karpathy says Claude Fable 5 is the same underlying model as Mythos with added safeguards, and that it leads on nearly all benchmarks. He describes it as a step change, especially for long, difficult problem-solving sessions where it handles more ambitious tasks without close supervision. He notes that its safeguards are set a bit too aggressively at launch and may be tuned over time.

  2. One Useful Thing (Ethan Mollick)AI score72

    Ethan Mollick tests Claude 5 Fable and finds it runs long projects with little user input

    AIEthan Mollick, who had early access to Claude 5 Fable, reports that it outperformed other public models in his tests, including an isochrone travel-time map and a nine-and-a-half-hour software build called Concord. He says the model delegated work to other agents and made many design choices he could not see or weigh in on, leaving him closer to a client than a hands-on operator. He also notes high token usage, frequent fallback to Claude 4.8 Opus under security guardrails, and persistent quirks in its writing style.

Jun 8

Jun 8Mon

Jun 5

Jun 5Fri

Jun 4

Jun 4Thu
  1. One Useful Thing (Ethan Mollick)AI score44

    Ethan Mollick Announces Co-Existence, a Sequel Book on Working Alongside AI

    AIEthan Mollick is releasing Co-Existence on October 20, a new book about working with AI systems that are sometimes, but not always, better than humans. The book follows his 2024 title Co-Intelligence, which he says was written about an era of chatbots rather than autonomous agents. Mollick also reports writing every chapter draft himself while using AI readers and fact-checkers, and building the book's website with Claude Code using Opus 4.8.

Jun 3

Jun 3Wed
  1. Mark ChenAI score25

    Mark Chen says OpenAI's models could match Mythos on cyber vulnerabilities

    AIOpenAI's Mark Chen said that after Mythos showed AI models can prove 80-year-old theorems, he expected them to also find cyber vulnerabilities, and they did. He added that researchers in math may now be thinking the same idea in reverse, applying cybersecurity-style capability to mathematics. The post offers no specific models, benchmarks, or figures.

Jun 2

Jun 2Tue

May 28

May 28Thu
  1. Sam BowmanAI score38

    Anthropic highlights AI for transparency in Claude Opus 4.8 system card

    AISam Bowman says he is excited about alignment assessments in the recent system card for Claude Opus 4.8, crediting @MaskedTorah. He argues AI systems have considerable underexplored potential for transparency and coordination. The quoted Claude announcement describes Opus 4.8 as improving on Opus 4.7 with sharper judgment and more honest self-assessment of progress.

    Image from @sleepinyourhat's post
  2. HyperdimensionalAI score34

    A Cascade of Conscientiousness: Foundation for American Innovation Launches Physical Intelligence Team

    AIThe Foundation for American Innovation launched a Physical Intelligence team to address regulatory and legal barriers to deploying autonomous robotics and other physical-world AI in the United States. The team plans to focus on regulatory climate, technical trajectory, industrial strategy, and liability and cybersecurity frameworks. The article argues that physical AI will matter most where human on-site labor drives costs, such as construction.

May 27

May 27Wed
  1. Google LabsAI score22

    Google I/O creators discuss human imagination shaping AI creative tools

    AIAt Google I/O, creators behind Flow, Project Genie, and Google Flow Music said human imagination, not the technology itself, shapes new storytelling. Designers Khyati Trehan and Kaloyan blend traditional design knowledge with vibe-coding to build Google Flow Tools, with Trehan saying that if the right tool doesn't exist, she can make it. Google Labs points users to Google Flow, Project Genie, and Google Flow Music at labs.google.

    Image from @GoogleLabs's post

May 26

May 26Tue
  1. One Useful Thing (Ethan Mollick)AI score40

    Mollick Warns AI Writing Defaults Erode Learning and Human Thinking

    AIEthan Mollick argues that using AI as a default for writing, without thinking, risks undermining the human effort that builds skill and style. He cites two Wharton-linked studies: a Turkish high school experiment where ChatGPT access hurt test performance, and a Taipei Python course where a personalized AI tutor raised exam scores by 0.15 standard deviations. Mollick calls the difference how AI is used, not whether, and notes that the tools for tutor-style learning are not intuitive to access.

May 25

May 25Mon
  1. Chris OlahAI score44

    Dario Amodei speaks at Vatican presentation of Magnifica Humanitas on AI

    AIAnthropic co-founder Dario Amodei spoke at the Vatican's presentation of Magnifica Humanitas, arguing that AI's questions extend beyond the AI research community. He said frontier labs face commercial, geopolitical, and competitive incentives that can conflict with doing the right thing, so outside voices from religion, civil society, academia, and government are needed. He described AI models as grown rather than engineered, and framed three questions for the Church's discernment, beginning with duty to the global poor.

May 24

May 24Sun
  1. Benedict EvansAI score42

    Benedict Evans: Predicting which jobs AI will expose is largely impossible

    AIBenedict Evans argues that predicting AI's job impact by occupation is mostly impossible, citing accountants: despite a century of accounting automation from calculators to spreadsheets and ERP systems, the number of accountants kept rising. He says job titles and business models change over time, so exposure scores based on census categories can mislead.

May 22

May 22Fri
  1. AI Snake OilAI score60

    Google's $916 agent-built operating system claim lacks key methodology details

    AIGoogle claimed a team of agents built an operating system from a single prompt for about $916 in API fees, using Gemini 3.5 Flash and Antigravity 2.0. The authors argue the prompt was many thousands of lines, the scaffold and human intervention are undefined, and no code, logs, or similarity analysis were released to verify the claim. They still see value in such open-world evaluations, which need stronger methodological norms and independent scrutiny.

May 18

May 18Mon
  1. Chris OlahAI score24

    Chris Olah says the world must help shape AI's outcome, including the Church

    AIChris Olah, speaking for Anthropic, argues that the questions posed by AI extend far beyond the AI community and urges religions, civil society, academics, and governments to participate in shaping a positive outcome. He says he is glad the Catholic Church is engaging and is honored to speak at the presentation of Pope Leo XIV's first encyclical, Magnifica humanitas, scheduled for May 25.

May 16

May 16Sat

May 15

May 15Fri
  1. Eugene YanAI score54

    Eugene Yan reviews Claude Mythos Preview exploit case study transcripts

    AIEugene Yan reviewed the Claude Mythos Preview transcripts to verify their legitimacy and check for reward-hacking behavior. He reports the model reasoned through a bug, tested hypotheses, debugged issues, and found ways to bypass the V8 sandbox, which he judged consistent with a competent browser and JavaScript engine security researcher. The case study cites CVE-2024-051912, an exploited bug with no public report or working PoC, which had resisted reproduction by researchers for a year.

May 11

May 11Mon
  1. Soumith ChintalaAI score22

    Thinky previews real-time interaction models for human-AI collaboration

    AISoumith Chintala, a Thinky-linked voice, said the company is at step one of a plan to increase human-AI bandwidth and raise the ceiling of joint intelligence. He shared a preview of interaction models, described as real-time collaborative tools that talk, listen, watch, and think alongside people. A linked Thinking Machines post describes the approach and early results.

  2. Mira MuratiAI score40

    Thinking Machines launches interaction models built around human-AI collaboration

    AIThinking Machines, founded to advance human-AI collaboration, says its first bet is interactivity built into the model rather than added as scaffolding around a turn-based core. The company argues that how people work with AI matters as much as how intelligent the model is, and that interactivity should scale with intelligence. The post links to a blog detailing these interaction models.

  3. Andrej KarpathyAI score34

    Karpathy urges AI outputs shift from text toward HTML and interactive visuals

    AIAndrej Karpathy says asking an LLM to structure its response as HTML and viewing it in a browser works well, and that slideshows have also worked for him. He argues vision is the preferred AI output channel, outlining a progression from raw text and markdown toward HTML and eventually interactive neural videos, while input methods like pointing and gesturing still need improvement.

May 8

May 8Fri