Skip to contentSkip to stories

Updated

#Trend

Showing low-relevance items too. Hide low-relevance items

May 8

May 8Fri

May 6

May 6Wed

May 3

May 3Sun

Apr 30

Apr 30Thu
  1. Andrej KarpathyAI score66

    Karpathy on agentic engineering, Software 3.0, and jagged AI capability

    AIAndrej Karpathy describes a December 2025 shift in which coding agents began producing larger, more reliable chunks of work, changing programming toward orchestrating agents. He argues that models automate what can be verified and that their capability is jagged, depending on verifiability and what labs emphasize in training, so users need to stay in the loop. He also says hiring, founder opportunities, and agent-native infrastructure should adapt to this shift.

Apr 29

Apr 29Wed
  1. Cognition Blog (Devin, Windsurf)AI score34

    Cognition opens Singapore headquarters for Asia-Pacific push with Devin

    AICognition has opened its Asia-Pacific headquarters in Singapore to expand its autonomous software engineering platform, Devin, across the region. The company says OCBC saw up to 30% improvement in code and test case generation, and its system integration test first-pass rate rose from below 50% to over 80% after deployment. Cognition is building its Singapore team across engineering, go-to-market, and partnerships, with Richard Spence leading APAC.

Apr 28

Apr 28Tue

Apr 24

Apr 24Fri

Apr 22

Apr 22Wed
  1. Cognition Blog (Devin, Windsurf)AI score54

    Cognition says building cloud agents requires VM isolation, state snapshots, and org change

    AICognition argues that enterprises building cloud agents face three problems: shared container kernels, the inability to persist agent state across async gaps, and the scale of orchestration, governance, and integrations. The post says VM-level isolation with hypervisor-level snapshots was needed for Devin, and that organizations must also rebuild engineering processes around agent execution.

Apr 20

Apr 20Mon
  1. Soumith ChintalaAI score36

    Soumith Chintala Critiques Dwarkesh's AGI Framing After Jensen Huang Podcast

    AISoumith Chintala says Jensen Huang understood AI ecosystems, trade, and policy far better than host Dwarkesh Patel in their podcast. He argues that no single model such as Mythos marks a critical phase change, since a state-of-the-art Chinese open-source model with three orders of magnitude more test-time compute and unpublished post-training advances would be a more realistic baseline. He also says American policy should use measured, continuous levers across a Western-controlled ecosystem rather than abrupt interventions.

Apr 18

Apr 18Sat

Apr 16

Apr 16Thu

Apr 14

Apr 14Tue
  1. Jan LeikeAI score14

    Jan Leike outlines top-down approach to automating alignment research

    AIJan Leike distinguishes two ways to automate alignment research: bottom-up, where researchers automate more of their existing work, and top-down, where specific subproblems are carved out for AI to solve. He says Anthropic's work mostly follows the bottom-up path, such as using Claude for coding, while this post focuses on the top-down approach.

Apr 13

Apr 13Mon
  1. Mira MuratiAI score26

    Thinking Machines welcomes Workshop Labs founders Luke Drago and @LRudL_

    AIThinking Machines has welcomed Luke Drago and @LRudL_, who co-founded Workshop Labs to build AI that keeps humans relevant. Mira Murati says they will continue that mission at Thinking Machines, which builds powerful AI systems that think alongside humans and extend human agency. She ties the hire to the company's broader work, including Tinker, research grants, and advancing the frontier.

Apr 7

Apr 7Tue

Apr 4

Apr 4Sat
  1. Andrej KarpathyAI score62

    Andrej Karpathy outlines an LLM-maintained markdown wiki workflow for personal research

    AIKarpathy describes using LLMs to compile raw source documents into a markdown wiki that he views in Obsidian, with the LLM writing and maintaining most of the wiki. He reports that at about 100 articles and 400K words, the LLM agent can answer complex questions directly from the wiki, and he also runs LLM health checks to find inconsistencies and gaps. He shares the underlying idea as an "idea file" that users can give to their own agents to build a customized version.

Apr 2

Apr 2Thu
  1. AI Futures ProjectAI score62

    AI Futures Project shortens Automated Coder timelines to mid 2028

    AIAI Futures Project moved Daniel Kokotajlo's Automated Coder median from late 2029 to mid 2028 and Eli's from early 2032 to mid 2030. The main reasons cited are a faster METR time horizon doubling time and the impressive results of Claude Opus 4.6. The authors also say progress in agentic coding has been faster than expected over the past 3 to 5 months.

Mar 30

Mar 30Mon
  1. Mckay WrigleyAI score22

    AI tools may soon use, clone, and extend any software autonomously

    AIMckay Wrigley predicts AI tools will within 6-12 months autonomously use any software, clone it in a weekend, monitor it for updates, and add custom features. He frames this as a future where users never need to operate their computer themselves. The prediction follows a referenced Claude Code update adding computer use in research preview for Pro and Max plans.

Mar 27

Mar 27Fri

Mar 26

Mar 26Thu
  1. Andrej KarpathyAI score47

    Karpathy wants agents to handle full app DevOps from one command

    AIAndrej Karpathy argues that the hardest part of building a deployed app is not the code but the DevOps work of assembling services, API keys, payments, auth, and deployment. He says the goal is for agents to handle this entire lifecycle as code, with agent-native CLI and API access instead of manual web clicking. He calls it a from-scratch redesign that is only now barely technically possible.

Mar 23

Mar 23Mon

Mar 17

Mar 17Tue
  1. Tri DaoAI score49

    Mamba-3 linear model released, outperforming Mamba-2 and Gated DeltaNet

    AITri Dao announced Mamba-3, which he described as the most powerful linear sequence model to date, as hybrid architectures increasingly rely on strong linear models. The post cites Qwen, Kimi-Linear, and NVIDIA's Nemotron-3 Super as examples of this trend. According to co-author Albert Gu, Mamba-3 shows noticeable performance gains over Mamba-2 and Gated DeltaNet at all sizes while maintaining speed.

Mar 1

Mar 1Sun
  1. Artificial IgnoranceAI score46

    Build Your Own Benchmark: Why Public AI Evals Are Saturating and What Replaces Them

    AIPublic AI benchmarks such as MMLU, SWE-bench Verified, and GPQA Diamond are saturating or showing contamination, prompting OpenAI to call SWE-bench Verified "no longer suitable" in late February and recommend SWE-bench Pro. OpenAI's audit found 59.4% of the problems its best model failed had flawed test cases, and GPT-5.2, Claude Opus 4.5, and Gemini 3 Flash could reproduce original fixes from memory. The article argues that behavioral tests, such as Vending-Bench's simulated vending machine business, may be more useful for everyday model choice.

Feb 27

Feb 27Fri

Feb 25

Feb 25Wed
  1. Jim FanAI score75

    EgoScale trains a 22-DoF humanoid mostly on 20,000 hours of human video

    AIResearchers trained a humanoid with 22-DoF dexterous hands mainly on over 20,000 hours of egocentric human video, with no robot in the loop, to perform tasks such as assembling model cars and folding shirts. They report a log-linear scaling law (R² = 0.998) between human video volume and action prediction loss, and state that this loss predicts real-robot success rate. The recipe, called EgoScale, pre-trains GR00T N1.5 on the video, adds only 4 hours of robot play data, and reports a 54% gain over training from scratch across five dexterous tasks.

    Video from @DrJimFan's post

Feb 19

Feb 19Thu

Feb 12

Feb 12Thu
  1. AI Futures ProjectAI score65

    AI Futures Project grades its 2025 AI 2027 predictions against reality

    AIAI Futures Project grades its AI 2027 scenario for 2025 and finds quantitative progress running at roughly 65% of the predicted pace, later revised to about 75%. Most qualitative predictions, such as the rise of coding agents, are judged on pace, while SWE-bench-Verified progress was slower than forecast and OpenAI's valuation trailed the scenario. The authors say their timelines lengthened over 2025 and plan to keep updating forecasts through 2026.

Feb 10

Feb 10Tue

Feb 9

Feb 9Mon

Feb 5

Feb 5Thu
  1. Geoffrey HintonAI score26

    Hinton praises International AI Safety Report 2026 as essential reading on AI risks

    AIGeoffrey Hinton called the International AI Safety Report 2026 a thoughtful, detailed, and well-researched description of AI risks, essential reading for anyone writing or speaking about them. Yoshua Bengio's thread introduces the report as the most comprehensive evidence-based assessment of AI capabilities, emerging risks, and safety measures to date.

Feb 1

Feb 1Sun
  1. Yi TayAI score35

    Yi Tay on hiring for frontier AI teams, seniority, and publication norms

    AIYi Tay, who recently hired a full AI team from thousands of applications, says PhDs are still a reasonable training ground and that candidates often get noticed through strong work they publish. He argues that seniority matters little in today's LLM world, and that being at the cutting edge outweighs external social media visibility. He also disagrees with the view that being a middle author on many papers is a negative signal.

Jan 29

Jan 29Thu
  1. Chip HuyenAI score18

    Chip Huyen launches GoodAIList.com to track trending open-source AI repos

    AIChip Huyen built GoodAIList.com, which tracks 14K open-source AI repositories with contributions from over 145K developers. Each day it searches for new repos using 123 keywords and topics, surfaces those gaining traction, and categorizes them with AI-generated annotations that she notes are not highly accurate. The site also maps contributor locations, which she uses to find people doing interesting AI work when she travels.

    Image from @chipro's post

Jan 26

Jan 26Mon

Jan 25

Jan 25Sun