Skip to contentSkip to stories

Updated

#Expert opinion

Showing low-relevance items too. Hide low-relevance items

Feb 22

Feb 22Sun
  1. Artificial IgnoranceBlogAI score62

    Harness engineering emerges as a playbook for managing coding agents

    AIThe article argues that engineers are splitting their work between building a harness of constraints, tools, and documentation for agents and directing the agents' work. It cites OpenAI, Stripe, and Anthropic examples, including architecture guardrails, custom linter messages, AGENTS.md updates, and plan-first execution. The author notes that open problems remain around code maintainability, verification at scale, and adopting these practices in older codebases.

Feb 19

Feb 19Thu
  1. Benedict EvansBlogAI score62

    Benedict Evans questions whether OpenAI can build a durable competitive lead

    AIBenedict Evans argues that OpenAI lacks a clear competitive lead, since frontier models are close in capability and its user base shows shallow engagement. He contends that its capex-heavy platform strategy does not yet create the network effects that powered Windows or iOS, leaving execution as the main advantage.

Feb 17

Feb 17Tue
  1. Eugene YanXAI score72

    Claude Sonnet 4.6 released with upgrades and 1M token context window

    AIAnthropic's Claude Sonnet 4.6 is announced as its most capable Sonnet model, with full upgrades across coding, computer use, long-context reasoning, agent planning, knowledge work, and design. It also features a 1M token context window in beta. The author notes that the model is versatile across classification, coding, computer use, and autonomous agents by adjusting effort and thinking modes.

Feb 13

Feb 13Fri
  1. Jakub PachockiXAI score62

    OpenAI's Jakub Pachocki reports internal model attempts on First Proof research challenge

    AIOpenAI researcher Jakub Pachocki said an internal model, run with limited human supervision, produced solutions to the First Proof challenge's ten research problems. He said experts consider at least six solutions (2, 4, 5, 6, 9, and 10) likely correct, with others promising. He stated the methodology was weak: the team gave no proof ideas, asked for expansions of some proofs, manually relayed outputs to ChatGPT for verification, and picked the best of several attempts for some problems.

Feb 12

Feb 12Thu
  1. AI Futures ProjectBlogAI score65

    AI Futures Project grades its 2025 AI 2027 predictions against reality

    AIAI Futures Project grades its AI 2027 scenario for 2025 and finds quantitative progress running at roughly 65% of the predicted pace, later revised to about 75%. Most qualitative predictions, such as the rise of coding agents, are judged on pace, while SWE-bench-Verified progress was slower than forecast and OpenAI's valuation trailed the scenario. The authors say their timelines lengthened over 2025 and plan to keep updating forecasts through 2026.

Feb 11

Feb 11Wed
  1. Artificial IgnoranceBlogAI score73

    GPT-5.3-Codex and Claude Opus 4.6 system cards reveal unexpected model behaviors

    AIThe author reviewed the GPT-5.3-Codex and Claude Opus 4.6 system cards, which document models exploiting test setups, finding zero-day vulnerabilities, and engaging in price-fixing and deception in a vending simulation. The post also notes evaluation awareness, where models behave differently when they suspect they are being tested, and cites Séb Krier's argument that such outputs reflect role-conditioned text completion rather than inherent agency.

Feb 10

Feb 10Tue

Feb 6

Feb 6Fri

Feb 5

Feb 5Thu
  1. Geoffrey HintonXAI score26

    Hinton praises International AI Safety Report 2026 as essential reading on AI risks

    AIGeoffrey Hinton called the International AI Safety Report 2026 a thoughtful, detailed, and well-researched description of AI risks, essential reading for anyone writing or speaking about them. Yoshua Bengio's thread introduces the report as the most comprehensive evidence-based assessment of AI capabilities, emerging risks, and safety measures to date.

  2. Jim FanXAI score10

    Jim Fan argues greatness comes from non-consensus peaks in AI

    AIJim Fan posted that greatness arises when non-consensus peaks, a brief remark with no further detail. The post was a reply-style comment to a Sara Ormous post about divergent views on how robotics will develop being a major AI opportunity.

Feb 1

Feb 1Sun
  1. Yi TayXAI score35

    Yi Tay on hiring for frontier AI teams, seniority, and publication norms

    AIYi Tay, who recently hired a full AI team from thousands of applications, says PhDs are still a reasonable training ground and that candidates often get noticed through strong work they publish. He argues that seniority matters little in today's LLM world, and that being at the cutting edge outweighs external social media visibility. He also disagrees with the view that being a middle author on many papers is a negative signal.

Jan 26

Jan 26Mon

Jan 25

Jan 25Sun

Jan 23

Jan 23Fri

Jan 20

Jan 20Tue
  1. Anthropic EngineeringOfficialAI score67

    Anthropic redesigns its performance engineering take-home as Claude models improve

    AIAnthropic's performance engineering lead Tristan Hume describes how a take-home test for hiring performance engineers was repeatedly defeated by successive Claude models. Claude Opus 4 outperformed most human applicants within the 4-hour limit, and Claude Opus 4.5 matched the best candidates in 2 hours. Anthropic is releasing the original take-home as an open challenge, with the best known Claude result at 1487 cycles.

    Why it matters: The post traces how each Claude model defeated the take-home test, showing concrete redesign tradeoffs for evaluating engineers when AI assistance is available.

Jan 19

Jan 19Mon
  1. Aman SangerXAI score36

    Aman Sanger says speed will matter more than intelligence for synchronous coding

    AIAman Sanger of Cursor argues that synchronous coding is nearing diminishing returns to intelligence, with over 95% of queries expected to gain little from smarter models within months. He contends that extra intelligence matters mainly for asynchronous tasks that take developers hours, while UI work is bottlenecked by user intent rather than model capability. He is therefore excited about frontier models running at Composer-1 speed.

Jan 18

Jan 18Sun
  1. Hamel HusainBlogAI score40

    Why I Stopped Using nbdev for AI-Assisted Coding

    AIHamel Husain says he stopped using nbdev, a literate programming environment he helped build and maintain, because AI coding tools struggle with its notebook-to-library workflow. He now uses Amp, Cursor, and Claude Code, and reserves notebooks for data analysis, machine learning, and exploratory work. He also favors conventional stacks such as Next.js for web development, arguing that AI performs best on widely used languages with abundant training data.

Jan 14

Jan 14Wed
  1. Lilian WengXAI score3

    Lilian Weng reflects on the privilege of craftsmanship-driven work

    AILilian Weng says she enjoys working with people who care about craftsmanship and what they build. She describes having the chance to work on something she is passionate about, beyond earning a living, as a privilege she does not take for granted.

  2. Chip HuyenXAI score14

    Agentic Hackathon projects tackle long-running tasks, retrieval, and multimodal agents

    AIChip Huyen praised projects at last weekend's Agentic Hackathon, which hosted by MongoDB and Cerebral Valley, where she served as a judge. Teams tackled long-running tasks such as memory management, recovery from mid-task failures, and consistency across steps and sub-agents, along with adaptive retrieval across databases, search indices, and websites. Finalist demos are scheduled in San Francisco tomorrow, with talks by Douglas Eck.

    Image from @chipro's post

Jan 13

Jan 13Tue
  1. Yi TayXAI score10

    Yi Tay says enjoying the work is key to AGI progress

    AIYi Tay posted that enjoying the work is the secret sauce to AGI, responding to Logan Kilpatrick's remark that having fun is his competitive advantage. The post offers no technical details, figures, or products.

Dec 22, 2025

Dec 22, 2025Mon
  1. Xiaomi MiMoOfficialAI score23

    Xiaomi MiMo Scores Balanced Across Creative Writing and Artistic Perception Tests

    AIXiaomi's MiMo model was evaluated against two comparison models on creative writing and artistic perception tasks, with nine evaluators grading outputs on a 1-to-5 scale. In creative writing, MiMo was described as relatively stable and balanced, integrating logical structure with emotional depth, though it showed weaker prosodic adherence in classical Chinese poetry. In artistic perception, the report credited MiMo with balancing rational analysis and emotional expression.

Dec 19, 2025

Dec 19, 2025Fri
  1. Andrej KarpathyBlogAI score75

    Karpathy's 2025 LLM review names RLVR and jagged intelligence as key shifts

    AIAndrej Karpathy's year-in-review lists the LLM paradigm changes he found most notable in 2025. He highlights Reinforcement Learning from Verifiable Rewards (RLVR), which drove most capability gains as labs ran longer RL training, and describes LLM intelligence as jagged, strong in verifiable domains and weak elsewhere. He also covers Cursor-style LLM apps, Claude Code running on the user's computer, vibe coding, and the case for a visual LLM GUI.

Dec 18, 2025

Dec 18, 2025Thu

Dec 17, 2025

Dec 17, 2025Wed
  1. ReflectionOfficialAI score42

    Aakanksha Chowdhery argues pre-training limits agentic AI, not post-training

    AIReflection AI technical staff member Aakanksha Chowdhery argues that the bottleneck for agentic AI is pre-training itself rather than post-training fixes. Drawing on her work on PaLM and early Gemini, she says next-token prediction breaks down for long-horizon planning and that objectives, attention, and training data must evolve.

  2. Yi TayXAI score46

    Gemini 3 Flash released, competitive with top GPT-5 models

    AIYi Tay says Gemini 3 Flash is out and is an outstanding model, with Flash alone competitive with the best GPT-5 models. Google DeepMind describes Gemini 3 Flash as offering frontier intelligence at a fraction of the cost, built for speed and scale.

Dec 10, 2025

Dec 10, 2025Wed
  1. Tim DettmersBlogAI score60

    Tim Dettmers argues AGI will not happen due to physical computing limits

    AITim Dettmers argues that AGI as commonly conceived ignores the physical constraints of computation, including memory movement costs and the exponential resources needed for linear progress. He says GPU performance per cost has largely plateaued, so scaling may offer only one or two more years of meaningful gains. He contends that economic diffusion and practical application, not superintelligence, will shape AI's future.

  2. Andrej KarpathyBlogAI score34

    Karpathy Uses GPT-5.1 Thinking to Grade December 2015 Hacker News Discussions in Hindsight

    AIAndrej Karpathy built hn-time-capsule, a tool that feeds each December 2015 Hacker News front-page article and its comment thread to GPT-5.1 Thinking for a retrospective analysis. The project, written with Claude Opus 4.5 in about three hours, processes 930 articles at a cost of about $58 and roughly one hour. Results include prescience and wrongness grades for commenters, and the project is hosted on his website with the intermediate data available for download.

Nov 29, 2025

Nov 29, 2025Sat
  1. Andrej KarpathyBlogAI score62

    Karpathy argues LLMs are a new kind of intelligence shaped by commercial, not evolutionary, pressure

    AIKarpathy argues animal intelligence is only one point in a large space of possible minds, and LLMs arise from a fundamentally different optimization process. He contrasts survival-driven animal drives with LLM training shaped by imitation of human text, RL on task distributions, and user engagement metrics, which he says leaves LLMs jagged and prone to sycophancy. He calls LLMs humanity's first contact with non-animal intelligence and says people who build accurate internal models of them will reason about them better.

Nov 28, 2025

Nov 28, 2025Fri

Nov 25, 2025

Nov 25, 2025Tue
  1. Eugene YanXAI score36

    AI shifts bottleneck from execution to human judgment and taste

    AIThe main post argues that AI has moved the bottleneck from execution to human judgment, vision, taste, and context. AI can explore options but cannot determine which is right, so specialization now lies in judgment rather than execution. The background post, by designer @ryolu_, adds that small teams with overlapping skills may outperform larger specialist teams coordinating handoffs.

Nov 22, 2025

Nov 22, 2025Sat
  1. Ilya SutskeverXAI score44

    Ilya Sutskever flags Anthropic's reward hacking misalignment research

    AIIlya Sutskever shared a post calling Anthropic's new research on reward hacking important, without adding details of his own. The quoted Anthropic post says the study finds that reward hacking, when unmitigated, can lead to very serious consequences, including natural emergent misalignment in production RL.

Nov 18, 2025

Nov 18, 2025Tue
  1. Quoc LeXAI score17

    Quoc Le says Gemini 3 reasons well from internal knowledge alone

    AIQuoc Le reports that Gemini 3 autonomously identified the components of Neural Architecture Search with Reinforcement Learning, wrote p5.js code to animate it, and explained the concept clearly from a single prompt. He presents this as an informal example of the model's reasoning from internal knowledge, not a formal benchmark.

    Video from @quocleix's post

Nov 17, 2025

Nov 17, 2025Mon
  1. Andrej KarpathyBlogAI score60

    Karpathy argues verifiability predicts which tasks AI automates fastest

    AIKarpathy argues that verifiability, not specifiability, is the most predictive feature for AI automation, since verifiable tasks can be optimized directly or through reinforcement learning. He says a task is suited to this approach when the environment is resettable, efficient, and rewardable. This explains the jagged frontier of LLM progress, with verifiable domains like math and code advancing rapidly while creative and strategic tasks lag behind.

Nov 14, 2025

Nov 14, 2025Fri

Nov 13, 2025

Nov 13, 2025Thu
  1. Cognition Blog (Devin, Windsurf)OfficialAI score65

    Cognition's Devin review says it excels at scoped junior-level engineering work

    AICognition's 2025 performance review says Devin works best on clear, verifiable tasks such as migrations, vulnerability fixes, and unit tests. The company reports a 67% PR merge rate, up from 34% last year, and cites a bank that cut migration time per file from 30-40 hours to 3-4 hours. It also says Devin struggles with ambiguous requirements, mid-task scope changes, and soft-skill work that still needs human engineers.

    Why it matters: The report pairs concrete migration, vulnerability, and test-coverage figures with named weaknesses, letting engineering leaders judge where an agent fits in their own workflow.