Skip to contentSkip to stories

Updated

#Expert opinion

Showing low-relevance items too. Hide low-relevance items

Jul 23

Jul 23Thu

Jul 21

Jul 21Tue
  1. Eugene YanAI score36

    Eugene Yan argues evals should weigh tail tasks, not median performance

    AIEugene Yan argues that model evals anchor on median tasks, but tail tasks determine project completion, making reliable models like Fable and Opus the difference between success and failure. He recommends treating models as collaborators who handle multi-hour or multi-day work with intent and success criteria, not as narrow-spec tools. Steve Yegge adds that Fable's carefulness is the dimension that matters most for production work.

    Image from @eugeneyan's post
  2. Soumith ChintalaAI score45

    Soumith Chintala says Poolside's Laguna S 2.1 suits agentic work on DGX Spark

    AISoumith Chintala praised Poolside's Laguna S 2.1 as looking strong for agentic use and said it fits on a single NVIDIA DGX Spark. The quoted Poolside release describes it as a 118B total-parameter Mixture-of-Experts model with 8B active per token, up to 1M-token context, and thinking and no-thinking modes, with weights openly available under OpenMDW-1.1.

  3. Rowan CheungAI score34

    Frontier AI models raise growing cybersecurity challenges, Demis warns

    AIRowan Cheung says AI models pushing the frontier are creating a growing challenge for cybersecurity. Quoting Demis, he reports that security must be addressed alongside the agentic era, with cyber worries about some models being just the beginning. Demis suggests this may be the time to push for standards and international cooperation.

    Video from @rowancheung's post
  4. Air Street PressAI score67

    DeepMind's Raia Hadsell argues AI should move beyond language to world models and robotics

    AIAt RAAIS, DeepMind VP of Research Raia Hadsell argued that the field focuses too much on language and should apply large-model training to worlds, robots, biology, and weather. The article cites DeepMind's DiffusionGemma, a 26-billion-parameter open text model that generates blocks by denoising rather than one token at a time, and the Genie-3 world model, which runs in real time for several minutes. It also describes world models as a source of synthetic training data for robots.

Jul 20

Jul 20Mon
  1. Bryan CatanzaroAI score28

    Open models enable forensic analysis that commercial guardrails blocked

    AIA security team found commercial frontier model APIs blocked their incident-response log analysis, which required submitting real attack commands and exploit payloads. They ran the forensic analysis on GLM 5.2, an open-weight model, on their own infrastructure, which also kept attacker data and referenced credentials inside their environment.

Jul 18

Jul 18Sat

Jul 17

Jul 17Fri

Jul 16

Jul 16Thu

Jul 15

Jul 15Wed
  1. Leandro von WerraAI score42

    Thinking Machines releases Inkling, a multimodal model with open weights

    AIThinking Machines has introduced Inkling, a model that reasons across text, image, and audio, with full weights made available. It is available for fine-tuning on Tinker and can be tried in the Inkling Playground. Hugging Face's Leandro von Werra praised the release for its grounded writing, interesting details, and strong ecosystem integration.

Jul 14

Jul 14Tue

Jul 13

Jul 13Mon
  1. AI Snake OilAI score57

    Narayanan argues AI job change will unfold over decades, not with one model release

    AIArvind Narayanan's ICML keynote argues that AI's labor impact will depend on slow organizational adaptation rather than a single lab milestone. He cites reliability measurements showing agent accuracy rose much faster than reliability over the last 24 months, and points to software engineering and past technologies like electricity and ATMs. He concludes that evaluation work and human judgment will become more central as building tasks are increasingly automated.

  2. Liquid AI NewsletterAI score6

    Liquid AI invites developers to introduce themselves and share their AI projects

    AILiquid AI is asking developers and researchers who build efficient, general-purpose AI for on-device hardware such as phones, laptops, cars, robots, and enterprise systems to introduce themselves in the comments. It invites readers to describe what they are building or studying, whether shipping products, publishing research, or working on side projects.

Jul 12

Jul 12Sun
  1. Jazzyear · ArticlesAI score67

    Peking University mathematician Dong Bin on AI solving the Anderson conjecture

    AIIn a long interview, Peking University professor Dong Bin describes his team's AI framework autonomously solving the Anderson conjecture, reportedly the first such domestic result with large-scale formal verification. He argues AI can accelerate mathematical theory but worries about verification bottlenecks, the pace of change, and how education and research evaluation must adapt.

Jul 10

Jul 10Fri
  1. AI Futures ProjectAI score38

    AI Futures Project Proposes Further Research Into Plan A and Alternative Scenarios

    AIAI Futures Project released AI 2040: Plan A and outlined further research areas, including building competing prescriptive scenarios such as Plan S, a domestic-first Plan A, GPU arms control, and CERN for AI. The group also flagged covert-project modeling and US domestic governance as areas of substantial uncertainty needing further work.

  2. Soumith ChintalaAI score29

    Thinking Machines outlines personalization, human participation, and decentralization goals

    AISoumith Chintala, a Thinking Machines figure, says the lab focuses on personalization and sovereignty, human participation, and decentralization to democratize AI. He argues these reduce society's dependence on centralized AGI companies, including his own. He points to Tinker, interaction models, and openly published research as previews, with more coming soon.

  3. Sebastien BubeckAI score73

    Bubeck says GPT-5.6 matches humans on a self-contracted curve bound

    AISebastien Bubeck reports that GPT-5.6-pro reproduced the 2^n lower bound and reached a 2.31^n upper bound on self-contracted gradient flow curve length. He compares these results with prior human work, where the best known upper bound is 2.29^n, and suggests the question may stop being useful for tracking AI progress within about six months.

Jul 9

Jul 9Thu
  1. Thinking Machines LabAI score44

    Thinking Machines Argues the Future Worth Building Keeps Humans Central to AI Decisions

    AIThinking Machines Lab says AI should extend human will and judgment, with people shaping its goals through continuous feedback rather than relying on models trained once and frozen. The company outlines three technical directions: training strong models, building tools for customization including training model weights, and developing interfaces that let personal judgment influence AI work. It also says it will publish research for the scientific community.

  2. AI Snake OilAI score62

    AI labs may escape the commodity trap by moving up the stack

    AIThe essay argues that AI labs selling model inference face commodity pricing pressure, but may achieve durable profits by moving into products, enterprise deployments, and switching-cost moats. It cites historical infrastructure industries and the Bertrand paradox to support the view that value capture depends on climbing the stack. The authors also warn that successful lock-in could raise enterprise costs and concentrate power, making early interoperability and portability standards important.

  3. Andrew NgAI score49

    Andrew Ng warns government pre-approval threatens open source AI innovation

    AIAndrew Ng argues that innovation thrives when inventors need not seek government permission in advance, citing Adam Thierer's "Permissionless Innovation." He says protecting open source AI is now a critical part of preserving that principle. Thierer's post, cited as background, describes an informal, opaque model-review regime in the US that could threaten open source models.

  4. Benedict EvansAI score60

    Benedict Evans argues AI token prices face unstable, commodity-leaning equilibrium

    AIBenedict Evans argues that token prices are unstable amid a supply crunch, and that foundation models may end up as low-margin commodity infrastructure rather than holding lasting pricing power. He cites inference gross margins of 40-50% that exclude training costs, which currently exceed revenue, and compares the outlook with mobile data and semiconductor manufacturing. He concludes that the outcome remains uncertain and that value capture above the model layer would require changes not yet visible.

Jul 7

Jul 7Tue
  1. Berkeley AI ResearchAI score62

    Berkeley researchers outline how data systems must change as agents take over knowledge work

    AIBerkeley AI Research authors argue that near-free inference will make agents the dominant workload for data systems, requiring redesign for agentic speculation, agent-run state and coordination, and agent-synthesized systems. The post cites inference prices falling 9x to 900x per year with a median near 50x, and reports that about 80-90% of sub-queries in a text-to-SQL benchmark were duplicates. It frames the three directions as data systems for, of, and by agents.

    Why it matters: The piece maps three concrete data-system challenges posed by near-free inference, useful for anyone designing infrastructure for agent workloads and memory.

Jul 6

Jul 6Mon