Cursor agents build a 3M+ line browser in a week
AI🤯 Quoted post: Watch Cursor build a 3M+ line browser in a week.
Updated
Updated
AI🤯 Quoted post: Watch Cursor build a 3M+ line browser in a week.
AIHamel Husain says he stopped using nbdev, a literate programming environment he helped build and maintain, because AI coding tools struggle with its notebook-to-library workflow. He now uses Amp, Cursor, and Claude Code, and reserves notebooks for data analysis, machine learning, and exploratory work. He also favors conventional stacks such as Next.js for web development, arguing that AI performs best on widely used languages with abundant training data.
AITim Dettmers, a professor who has used Claude Code for eight months to automate his own work, argues that more than 90% of code and text should be written by agents. He says the coding-focused hype on Twitter overstates parallel sessions and autonomy, which translate poorly to most non-software tasks. The post offers a balanced guide to what actually works in agent-based automation.
AIMiniMax has released MiniMax-M2.1 model weights on Hugging Face, with API access on the MiniMax Open Platform and the MiniMax Agent product. The company reports gains over M2 on coding and agent benchmarks such as SWE-bench Verified (74.0) and VIBE average (88.6), and says it outperforms Claude Sonnet 4.5 on multilingual scenarios.
Why it matters: The release pairs open weights with a broad benchmark table against Claude and GPT models, letting readers compare coding and agent claims directly.
AIAndrej Karpathy's year-in-review lists the LLM paradigm changes he found most notable in 2025. He highlights Reinforcement Learning from Verifiable Rewards (RLVR), which drove most capability gains as labs ran longer RL training, and describes LLM intelligence as jagged, strong in verifiable domains and weak elsewhere. He also covers Cursor-style LLM apps, Claude Code running on the user's computer, vibe coding, and the case for a visual LLM GUI.
AIXiaomi released and open-sourced MiMo-V2-Flash, a Mixture-of-Experts model with 309B total and 15B active parameters, under the MIT license. The company reports 73.4% on SWE-Bench Verified, the top score among open-source models, and inference at 150 tokens per second for $0.1 per million input tokens and $0.3 per million output tokens. It supports a hybrid thinking mode and a 256k context window.
Why it matters: The post gives architecture, speculative decoding speedup, and pricing figures, which help readers judge how the efficiency claims are achieved and what they cost.
AIOpenAI is introducing GPT-5.2 in ChatGPT, describing it as its most advanced model series for professional work. GPT-5.2 Thinking is positioned for tasks such as building spreadsheets and presentations, writing and reviewing production code, and analyzing long documents. The post says it beats or ties industry professionals on well-specified knowledge work tasks spanning 44 occupations 70.9% of the time on GDPval, and GPT-5.2 Instant, Thinking, and Pro begin rolling out to all tiers, starting with paid plans.
Why it matters: The post links the model's professional-work focus to GDPval results across 44 occupations, showing how the claimed capability was measured.
AIAndrej Karpathy built hn-time-capsule, a tool that feeds each December 2015 Hacker News front-page article and its comment thread to GPT-5.1 Thinking for a retrospective analysis. The project, written with Claude Opus 4.5 in about three hours, processes 930 articles at a cost of about $58 and roughly one hour. Results include prescience and wrongness grades for commenters, and the project is hosted on his website with the intermediate data available for download.
AIWe asked it to design a fractal explosion in HTML. By using extended thinking to strategize on recursion and aesthetics before coding, it turns a simple prompt into sophisticated digital art.
AIGoogle's Gemini 3 Deep Think mode is now available in the Gemini app for Ultra users. The post says it uses parallel thinking for difficult coding and scientific tasks and builds on technology that reached gold-medal level at the ICPC World Finals and IMO.
AIAI explores options, but can't tell you which is right. that's where specialization matters now – in judgment, not execution.
AICognition's 2025 performance review says Devin works best on clear, verifiable tasks such as migrations, vulnerability fixes, and unit tests. The company reports a 67% PR merge rate, up from 34% last year, and cites a bank that cut migration time per file from 30-40 hours to 3-4 hours. It also says Devin struggles with ambiguous requirements, mid-task scope changes, and soft-skill work that still needs human engineers.
Why it matters: The report pairs concrete migration, vulnerability, and test-coverage figures with named weaknesses, letting engineering leaders judge where an agent fits in their own workflow.
AIWindsurf has launched Codemaps, AI-annotated structured maps of a codebase powered by SWE-1.5 and Claude Sonnet 4.5, which users can generate from a task prompt using a Fast (SWE-1.5) or Smart (Sonnet 4.5) model. Codemaps links grouped code sections to exact lines and can be referenced in Cascade with @{codemap} to give agents more specific context.
AIChip Huyen argues that many AI product failures stem from user experience, data quality, and organizational structure rather than the AI itself. She cites a chatbot whose traction improved after adding pre-populated questions and a voice option for users whose hands were busy, and a lead scoring model that was broken because marketing wasn't asking the right questions. She also notes that senior engineers gain the most from AI coding while resisting it more, and recommends building small tools for daily frustrations to solve the "idea crisis."
AICognition has released SWE-1.5, a model optimized for software engineering that it says reaches near-frontier coding performance while running at up to 950 tok/s with Cerebras inference. The company reports it is 6x faster than Haiku 4.5 and 13x faster than Sonnet 4.5, and it is available now in Windsurf. The post's SWE-Bench Pro chart places SWE-1.5 at 40.08%, behind Sonnet 4.5 at 43.60%, and it notes that the model was trained with reinforcement learning on the Cascade agent harness.
Why it matters: The post pairs a benchmark chart with a 950 tok/s speed claim and describes how harness, RL environments, and inference were co-designed, useful context for judging the speed-versus-quality tradeoff.
AICognition says its autonomous coding agent Devin can complete a .NET Framework to .NET Core migration in as little as two weeks, using a Strangler Fig approach adapted from Jimmy Bogard's guide. The post says Devin handles planning via Ask Devin and DeepWiki, dependency sharing, controller and view conversion, and session state adaptation through a remote app.
AICognition introduces SWE-grep and SWE-grep-mini, fast agentic models trained with reinforcement learning for multi-turn context retrieval in coding tasks. The company says they match frontier coding models at retrieval while taking an order of magnitude less time, and they power the Fast Context subagent in Windsurf. The models issue up to 8 parallel tool calls per turn within 4 turns, and Cerebras serves SWE-grep-mini at over 2,800 tokens per second and SWE-grep at over 650 tokens per second.
Why it matters: The post explains the speed-intelligence tradeoff in agentic code search, showing how parallel tool calls and RL training change the cost of retrieving context for coding agents.
AICognition rebuilt its Devin coding agent for Claude Sonnet 4.5, reporting 2x faster performance and 12% better results on its Junior Developer Evals, now available in Agent Preview. The team found the model is aware of its context window, which led to premature wrap-up behavior that they countered with repeated prompts and a 200k usage cap within a 1M token beta.
Why it matters: The post explains which agent behaviors changed under Sonnet 4.5, such as context-window awareness and note-taking, that forced a rebuild rather than a simple model swap.
AICognition says Claude Sonnet 4.5 is available in Devin starting today, improving its planning performance by 18% and end-to-end eval scores by 12%. Cognition says the model's testing of its own code lets Devin run longer and handle harder tasks.
AICognition, maker of the AI software engineer Devin, raised over $400M at a $10.2B post-money valuation led by Founders Fund. The company says its acquisition of Windsurf more than doubled its ARR, with combined enterprise ARR up over 30% in the seven weeks after the deal. It also reports Devin ARR grew from $1M in September 2024 to $73M in June 2025, with total net burn under $20M.
AICognition Blog published "Coding Agents 101: The Art of Actually Getting Things Done" on June 26, 2025, as a guide to using coding agents effectively. The available source text contains only navigation and a list of other Cognition posts, so no specific techniques, features, or results can be verified from it.
AICognition is bringing back its Devin Open Source Initiative, offering $500 in Devin ACU credits to open-source GitHub projects with over 100 forks. Projects below that threshold will still be considered. Eligible projects must have an OSI-approved license and be actively maintained, and maintainers can apply through a linked form.
AICognition has released Devin 2.1, which reports its confidence in completing tasks using green, yellow, and red ratings. The company says green scores led to twice the likelihood of a merged PR compared with red, and Devin now also answers codebase questions and scores Linear and Jira issues.
Why it matters: The post explains how Devin now shows confidence scores and asks clarifying questions, which changes how teams can decide which tasks to hand over.
AIStanford and Cognition AI researchers introduced Kevin-32B, a 32B-parameter model trained with multi-turn reinforcement learning to write CUDA kernels. On KernelBench, it solves 89% of tasks at best@16 and achieves 65% average correctness over eight refinement steps, versus 53% for o4-mini and 51% for o3. Its best@16 speedup is 1.41x, and multi-turn training outperforms single-turn training as refinement steps increase.
AICognition has launched DeepWiki, a free public version of its Devin Wiki and Devin Search tools that helps developers understand codebases. Users can view docs for any repo by replacing github.com with deepwiki.com in the URL, and more than 50,000 top public GitHub repositories are already indexed. Private repositories require a Devin account.
AICognition has released Devin 2.0, a new agent-native IDE experience with a flexible plan starting at $20. The update lets users run multiple parallel Devins, each with its own cloud-based IDE, and adds Interactive Planning, Devin Search, and Devin Wiki.
Why it matters: The release adds planning, codebase search, and auto-generated wikis to Devin, showing how an agent can prepare work before executing it.
AICognition's February 2025 Devin update adds parallel batch edits, beta GitLab support, and Sonnet 3.7, which Cognition says is the best model it has tested for debugging, codebase search, and agentic planning. Devin is about 2x faster than in October 2024, taking about 7.8 minutes on average to complete junior developer tasks in Cognition's internal evaluations. Other changes include copy-paste in Devin's browser and proactive feedback on suboptimal prompts.
AILinktree has used Devin, Cognition's AI software engineer, to merge roughly 100 pull requests in a month, mostly fixing customer-reported bugs and implementing small features. The engineering team also used Devin to add support for new social media platforms, launching five Devins, one per repo and PR, and later used the Devin API with a Playbook script to spawn multiple Devins for multi-repo features. The team says Devin works best on tasks an engineer could finish in a couple of hours.
AICognition's Devin can be triggered through its External API by GitHub Actions to automatically review pull requests, typically within five to ten minutes. The setup involves adding a workflow file, storing a DEVIN_API_KEY secret, and customizing the review prompt to match team conventions. Cognition recommends treating Devin as an extra reviewer rather than a replacement for human oversight, since it does not catch every bug.
AICognition's January 2025 Devin update improves its ability to find relevant files and reuse existing code in repositories, with changes rolling out to all users. It also adds enterprise accounts for centralized management of multiple organizations, audio message support in Slack, and pay-as-you-go billing after monthly ACU capacity is used, starting January 9.
AICrossmint said Devin became its top contributor to the open-source GOAT SDK during an initial trial, merging 8 pull requests versus 4 for the next contributor. Examples included a DEXScreener plugin built from a documentation URL and a harder Sui blockchain integration that needed three rounds of feedback and about an hour of human involvement. The company said its results depended on proper training, clear task context, and planned validation, not on treating Devin as superhuman.
AICognition has made Devin generally available to all engineering teams, with subscriptions starting at $500 per month. Over the past two weeks, Devin was made about 10% faster and about 10% more cost-efficient, especially for tasks requiring many code edits. The update also adds fixes for stuck or hanging sessions, more options to customize filters and Slack notifications, and larger machine settings for disk, RAM, and CPU.
AICognition is launching the Devin Open Source Initiative, giving selected open source maintainers 500 free ACUs on a Devin Teams plan as part of Devin's general availability launch. The post shows Devin contributing pull requests to projects including Anthropic's MCP Inspector, Dagger, and nanoGPT, with maintainers still reviewing the results. Devin's GitHub integration forwards PR comments and CI checks to help refine changes, though the company warns a human should still verify final quality.
AICognition is making Devin generally available to engineering teams starting at $500 a month, with no seat limits and access to its Slack integration, IDE extension, and API. The post recommends starting with small frontend bugs, first-draft PRs for backlog tasks, and targeted refactors, and shares open-source PR sessions where Devin resolved issues for projects including Anthropic MCP, Zod, and nanoGPT.
Why it matters: The post shows concrete open-source PR examples and the tasks where Devin works best, helping teams judge where an autonomous coding agent fits their workflow.
AICognition's December 2024 Devin update lets users tag Devin in Slack threads and ask it to create PRs, with Devin automatically responding to PR comments and lint failures. The update adds Repo Knowledge that Devin generates by scanning repositories, an Agency setting that makes Devin propose plans before executing complex tasks, and a REST API for structured input and output. Devin also gained faster session startup and enterprise options such as Okta single sign-on.
AICognition tested OpenAI's o1-mini and o1-preview in a simplified Devin-Base agent, comparing them with GPT-4o on its internal cognition-golden benchmark. The chart reports Devin-Base scores of 25.9% with GPT-4o, 34.6% with o1-mini, and 51.8% with o1-preview, versus 74.2% for the production Devin. The post also describes the benchmark's realistic environments, simulated users, and agent-based evaluation.
Why it matters: The post explains how Cognition evaluates coding agents with autonomous, environment-based tests, which shows how base-model swaps are measured in practice.
AICognition's Devin gained a September 2024 update with MultiDevin, which lets a manager Devin delegate work to up to 10 worker Devins, currently available on the Enterprise plan. Devin now automatically responds to comments on its pull requests, suggests Knowledge additions, and can restore earlier checkpoints, and the company reports up to an 80% reduction in time for common tasks.