Geoffrey Hinton recounts early AI days in Jeff Dean chat
AIGeoffrey Hinton says he recently talked about the old days with Jeff Dean, in a fireside chat moderated by Jordan Jacobs. The full recording of that discussion is now available on Spotify.
Updated
Updated
AIGeoffrey Hinton says he recently talked about the old days with Jeff Dean, in a fireside chat moderated by Jordan Jacobs. The full recording of that discussion is now available on Spotify.
AIYi Tay says Gemini 3 Flash is out and is an outstanding model, with Flash alone competitive with the best GPT-5 models. Google DeepMind describes Gemini 3 Flash as offering frontier intelligence at a fraction of the cost, built for speed and scale.
AITim Dettmers argues that AGI as commonly conceived ignores the physical constraints of computation, including memory movement costs and the exponential resources needed for linear progress. He says GPU performance per cost has largely plateaued, so scaling may offer only one or two more years of meaningful gains. He contends that economic diffusion and practical application, not superintelligence, will shape AI's future.
AIAndrej Karpathy built hn-time-capsule, a tool that feeds each December 2015 Hacker News front-page article and its comment thread to GPT-5.1 Thinking for a retrospective analysis. The project, written with Claude Opus 4.5 in about three hours, processes 930 articles at a cost of about $58 and roughly one hour. Results include prescience and wrongness grades for commenters, and the project is hosted on his website with the intermediate data available for download.
AIKarpathy argues animal intelligence is only one point in a large space of possible minds, and LLMs arise from a fundamentally different optimization process. He contrasts survival-driven animal drives with LLM training shaped by imitation of human text, RL on task distributions, and user engagement metrics, which he says leaves LLMs jagged and prone to sycophancy. He calls LLMs humanity's first contact with non-animal intelligence and says people who build accurate internal models of them will reason about them better.
AIIlya Sutskever says scaling current AI methods will keep producing improvements and will not stall. He adds that something important will still be missing.
AIThe main post argues that AI has moved the bottleneck from execution to human judgment, vision, taste, and context. AI can explore options but cannot determine which is right, so specialization now lies in judgment rather than execution. The background post, by designer @ryolu_, adds that small teams with overlapping skills may outperform larger specialist teams coordinating handoffs.
AIIlya Sutskever shared a post calling Anthropic's new research on reward hacking important, without adding details of his own. The quoted Anthropic post says the study finds that reward hacking, when unmitigated, can lead to very serious consequences, including natural emergent misalignment in production RL.
AIQuoc Le reports that Gemini 3 autonomously identified the components of Neural Architecture Search with Reinforcement Learning, wrote p5.js code to animate it, and explained the concept clearly from a single prompt. He presents this as an informal example of the model's reasoning from internal knowledge, not a formal benchmark.
AIKarpathy argues that verifiability, not specifiability, is the most predictive feature for AI automation, since verifiable tasks can be optimized directly or through reinforcement learning. He says a task is suited to this approach when the environment is resettable, efficient, and rewardable. This explains the jagged frontier of LLM progress, with verifiable domains like math and code advancing rapidly while creative and strategic tasks lag behind.
AIChip Huyen responded to Sam Altman's post saying ChatGPT now follows custom instructions to avoid em-dashes. The main post itself only reads "Sam!!!", so its reaction is conveyed mainly through the quoted context.
AICognition's 2025 performance review says Devin works best on clear, verifiable tasks such as migrations, vulnerability fixes, and unit tests. The company reports a 67% PR merge rate, up from 34% last year, and cites a bank that cut migration time per file from 30-40 hours to 3-4 hours. It also says Devin struggles with ambiguous requirements, mid-task scope changes, and soft-skill work that still needs human engineers.
Why it matters: The report pairs concrete migration, vulnerability, and test-coverage figures with named weaknesses, letting engineering leaders judge where an agent fits in their own workflow.
AIAman Sanger of Cursor argues that heavy compute spent at indexing time can be reused to improve performance without raising inference-time compute, with embeddings as the simplest mechanism. Cursor's background post says semantic search improves its agent's accuracy across frontier models, especially in large codebases where grep alone falls short.
AIChip Huyen argues that many AI product failures stem from user experience, data quality, and organizational structure rather than the AI itself. She cites a chatbot whose traction improved after adding pre-populated questions and a voice option for users whose hands were busy, and a lead scoring model that was broken because marketing wasn't asking the right questions. She also notes that senior engineers gain the most from AI coding while resisting it more, and recommends building small tools for daily frustrations to solve the "idea crisis."
AILilian Weng says on-policy distillation lets a teacher model act as a process reward model, providing dense rewards during training. The approach also prevents the out-of-distribution shock that SFT-style training can cause during rollouts. Thinking Machines' related post reports it outperforms other approaches for math reasoning and an internal chat assistant at a fraction of the cost.
AIA hiring manager told Chip Huyen that a software engineering candidate who has not experimented with vibe coding is a red flag. Huyen posted the remark on X and asked for readers' thoughts.
AIGeoffrey Hinton recorded a podcast with Jon Stewart, whom he describes as a longtime hero, focused on explaining how AI works. Hinton says Stewart was especially eager to understand the underlying mechanics.
AIEight Sleep integrated Cognition's Devin into its data workflows, letting staff tag Devin in Slack to query Snowflake, dbt, and Looker and check Amplitude. The company says it is now shipping 3x as many data features and investigations each week, with its ad-hoc data request queue near zero. Devin was used to trace a suspicious revenue spike to a better-than-expected email campaign.
AICognition argues that parallel multi-agent architectures are fragile because subagents act on conflicting, unshared assumptions. It proposes two principles for reliable agents: share context and full agent traces, and treat actions as carrying implicit decisions. The post recommends simpler single-threaded designs for most cases and notes that context compression and fine-tuned models can extend long-running tasks.
Why it matters: The post explains concrete failure modes of parallel multi-agent setups and offers two context-sharing principles, useful for anyone designing long-running agent systems.
AICognition tested OpenAI's o1-mini and o1-preview in a simplified Devin-Base agent, comparing them with GPT-4o on its internal cognition-golden benchmark. The chart reports Devin-Base scores of 25.9% with GPT-4o, 34.6% with o1-mini, and 51.8% with o1-preview, versus 74.2% for the production Devin. The post also describes the benchmark's realistic environments, simulated users, and agent-based evaluation.
Why it matters: The post explains how Cognition evaluates coding agents with autonomous, environment-based tests, which shows how base-model swaps are measured in practice.