Skip to contentSkip to stories

Updated

#Deployment/Engineering

Items with an AI score under 20 are hidden. Show low-relevance items

Jan 23

Jan 23Fri

Jan 21

Jan 21Wed
  1. Mistral AI · new models on Hugging FaceAI score65

    Mistral releases open-weight Voxtral Mini 4B Realtime 2602 speech model

    AIMistral AI released Voxtral Mini 4B Realtime 2602, a multilingual realtime speech-transcription model with 13 supported languages under the Apache 2.0 license. The model has a configurable transcription delay from 240ms to 2.4s, and it matches leading offline open-source models at a 480ms delay. The source says it is optimized for on-device deployment and is currently supported only in vLLM.

    Why it matters: The source specifies the 480ms delay operating point, 4B size, Apache 2.0 license, and vLLM serving path, which matter for teams weighing realtime transcription deployment.

Jan 20

Jan 20Tue
  1. Anthropic EngineeringAI score67

    Anthropic redesigns its performance engineering take-home as Claude models improve

    AIAnthropic's performance engineering lead Tristan Hume describes how a take-home test for hiring performance engineers was repeatedly defeated by successive Claude models. Claude Opus 4 outperformed most human applicants within the 4-hour limit, and Claude Opus 4.5 matched the best candidates in 2 hours. Anthropic is releasing the original take-home as an open challenge, with the best known Claude result at 1487 cycles.

    Why it matters: The post traces how each Claude model defeated the take-home test, showing concrete redesign tradeoffs for evaluating engineers when AI assistance is available.

  2. Cognition Blog (Devin, Windsurf)AI score54

    Cognition launches Devin Review to help humans review AI-generated code

    AICognition introduced Devin Review, a free early-release code review tool that works on any public or private GitHub PR, with features for organizing diffs, chatting about changes, and flagging AI-detected bugs. The company says code review, not code generation, is now the bottleneck as coding agents increase the volume and size of pull requests.

Jan 19

Jan 19Mon
  1. Factory NewsAI score47

    Factory Introduces Agent Readiness to Score Codebases for Autonomous Coding Agents

    AIFactory's new Agent Readiness tool evaluates repositories across eight technical pillars and five maturity levels, using 60+ binary criteria run via the /readiness-report command. The company says it can also open pull requests to fix foundational gaps such as missing AGENTS.md files, linter configuration, and pre-commit hooks. Factory says scores are now more consistent, with variance dropping from an average of 7% to 0.6%.

  2. Aman SangerAI score36

    Aman Sanger says speed will matter more than intelligence for synchronous coding

    AIAman Sanger of Cursor argues that synchronous coding is nearing diminishing returns to intelligence, with over 95% of queries expected to gain little from smarter models within months. He contends that extra intelligence matters mainly for asynchronous tasks that take developers hours, while UI work is bottlenecked by user intent rather than model capability. He is therefore excited about frontier models running at Composer-1 speed.

  3. Z.ai (GLM) · new models on Hugging FaceAI score62

    Z.ai releases GLM-4.7-Flash, a 30B-A3B MoE model for lightweight deployment

    AIZ.ai has released GLM-4.7-Flash, a 30B-A3B MoE model that it positions as the strongest model in the 30B class. The model reports SWE-bench Verified 59.2 and τ²-Bench 79.5, and supports local deployment through vLLM and SGLang.

    Why it matters: The source lists benchmark scores against Qwen3-30B-A3B-Thinking-2507 and GPT-OSS-20B, letting readers compare the 30B-class MoE model directly with its named rivals.

Jan 18

Jan 18Sun
  1. Hamel HusainAI score40

    Why I Stopped Using nbdev for AI-Assisted Coding

    AIHamel Husain says he stopped using nbdev, a literate programming environment he helped build and maintain, because AI coding tools struggle with its notebook-to-library workflow. He now uses Amp, Cursor, and Claude Code, and reserves notebooks for data analysis, machine learning, and exploratory work. He also favors conventional stacks such as Next.js for web development, arguing that AI performs best on widely used languages with abundant training data.

Jan 14

Jan 14Wed
  1. Black Forest Labs · new models on Hugging FaceAI score62

    Black Forest Labs releases FLUX.2 [klein] 4B image model under Apache 2.0

    AIBlack Forest Labs released FLUX.2 [klein] 4B, a 4 billion parameter model that unifies text-to-image generation and image editing with multi-reference support. The source says it runs on consumer GPUs such as the RTX 3090 or 4070 with about 13GB VRAM, and its open weights are available under the Apache 2.0 license.

    Why it matters: The source specifies a 4 billion parameter model running on about 13GB VRAM under Apache 2.0, which helps readers judge whether local image generation fits their hardware.

Jan 13

Jan 13Tue
  1. Tim DettmersAI score36

    Tim Dettmers Argues Agents Should Automate Most Personal Work, Not Just Code

    AITim Dettmers, a professor who has used Claude Code for eight months to automate his own work, argues that more than 90% of code and text should be written by agents. He says the coding-focused hype on Twitter overstates parallel sessions and autonomy, which translate poorly to most non-software tasks. The post offers a balanced guide to what actually works in agent-based automation.

Jan 7

Jan 7Wed
  1. Nick TurleyAI score72

    OpenAI launches ChatGPT Health for connecting medical records

    AIOpenAI is launching ChatGPT Health, a dedicated and private space where users can securely connect apps and medical records. The launch starts with a small group of users from the waitlist, with access expanding over the coming weeks.

    Why it matters: The post names the access path and a dedicated space for health records, which matters for judging how sensitive data would be handled.

Jan 6

Jan 6Tue
  1. Cognition Blog (Devin, Windsurf)AI score42

    Infosys partners with Cognition to deploy Devin AI software engineer across its enterprise

    AIInfosys will deploy Cognition's Devin, an autonomous AI software engineer, across its own teams and global client base to expand delivery capacity. The rollout begins in its Financial Services practice, covering banking, payments, capital markets, insurance, and wealth management, and is planned to extend to retail, energy, and healthcare. Over the past six months, Infosys reports material productivity gains, including COBOL and JCP servlet migrations completed in record time.

Dec 19, 2025

Dec 19, 2025Fri

Dec 12, 2025

Dec 12, 2025Fri

Dec 11, 2025

Dec 11, 2025Thu
  1. Runway ResearchAI score62

    Runway Introduces GWM-1, a Real-Time General World Model Family

    AIRunway announced GWM-1, its first general world model family, built on Gen-4.5 and generating frames autoregressively in real time under interactive control. It comes in three variants: GWM Worlds for explorable environments, GWM Avatars for conversational characters, and GWM Robotics for robotic manipulation. Runway also says it is working toward unifying these domains under a single base world model, and GWM Robotics includes a Python SDK.

    Why it matters: The post separates three GWM-1 variants and ties each to a concrete use, which clarifies where a general world model would fit compared with a single model.

Dec 10, 2025

Dec 10, 2025Wed
  1. Tim DettmersAI score60

    Tim Dettmers argues AGI will not happen due to physical computing limits

    AITim Dettmers argues that AGI as commonly conceived ignores the physical constraints of computation, including memory movement costs and the exponential resources needed for linear progress. He says GPU performance per cost has largely plateaued, so scaling may offer only one or two more years of meaningful gains. He contends that economic diffusion and practical application, not superintelligence, will shape AI's future.

Dec 4, 2025

Dec 4, 2025Thu

Dec 2, 2025

Dec 2, 2025Tue
  1. Apple · new models on Hugging FaceAI score36

    Apple releases CLaRa-7B-E2E, an end-to-end RAG model with 16x and 128x compression

    AIApple's CLaRa-7B-E2E is a fully end-to-end unified RAG model that jointly optimizes retrieval and generation, with 16x and 128x document compression. It is trained with end-to-end finetuning using differentiable top-k retrieval and a unified language-modeling objective. The model is available on Hugging Face with example end-to-end inference code.

Dec 1, 2025

Dec 1, 2025Mon

Nov 13, 2025

Nov 13, 2025Thu
  1. Cognition Blog (Devin, Windsurf)AI score65

    Cognition's Devin review says it excels at scoped junior-level engineering work

    AICognition's 2025 performance review says Devin works best on clear, verifiable tasks such as migrations, vulnerability fixes, and unit tests. The company reports a 67% PR merge rate, up from 34% last year, and cites a bank that cut migration time per file from 30-40 hours to 3-4 hours. It also says Devin struggles with ambiguous requirements, mid-task scope changes, and soft-skill work that still needs human engineers.

    Why it matters: The report pairs concrete migration, vulnerability, and test-coverage figures with named weaknesses, letting engineering leaders judge where an agent fits in their own workflow.

Nov 3, 2025

Nov 3, 2025Mon
  1. Cognition Blog (Devin, Windsurf)AI score47

    Windsurf Codemaps Adds AI-Annotated Code Maps to Help Engineers Understand Code

    AIWindsurf has launched Codemaps, AI-annotated structured maps of a codebase powered by SWE-1.5 and Claude Sonnet 4.5, which users can generate from a task prompt using a Fast (SWE-1.5) or Smart (Sonnet 4.5) model. Codemaps links grouped code sections to exact lines and can be referenced in Cascade with @{codemap} to give agents more specific context.

Oct 30, 2025

Oct 30, 2025Thu
  1. Chip HuyenAI score27

    Chip Huyen's AI product lessons: UX, data, and team structure matter most

    AIChip Huyen argues that many AI product failures stem from user experience, data quality, and organizational structure rather than the AI itself. She cites a chatbot whose traction improved after adding pre-populated questions and a voice option for users whose hands were busy, and a lead scoring model that was broken because marketing wasn't asking the right questions. She also notes that senior engineers gain the most from AI coding while resisting it more, and recommends building small tools for daily frustrations to solve the "idea crisis."

  2. Moonshot AI (Kimi) · new models on Hugging FaceAI score60

    Moonshot AI releases Kimi Linear 48B hybrid linear attention models on Hugging Face

    AIMoonshot AI released Kimi Linear, a hybrid linear attention architecture with 48B total and 3B activated parameters and a 1M-token context length, on Hugging Face. The model card reports up to 6.3x faster TPOT than MLA at 1M tokens and up to 75% lower KV cache needs, and says it outperforms full attention on long-context and RL-style benchmarks.

    Why it matters: The model card gives concrete long-context speed and memory figures for a hybrid attention design, useful for judging whether linear attention can replace full attention in practice.

  3. Moonshot AI (Kimi) · new models on Hugging FaceAI score72

    Moonshot AI releases Kimi Linear 48B-A3B hybrid attention models on Hugging Face

    AIMoonshot AI has released Kimi-Linear-Base and Kimi-Linear-Instruct, both 48B total and 3B activated parameters with a 1M context length, on Hugging Face. The models use Kimi Delta Attention in a 3:1 hybrid ratio with global MLA, cutting KV cache by up to 75% and boosting decoding throughput by up to 6x at 1M tokens. The KDA kernel is open-sourced in FLA, and the checkpoints were trained on 5.7T tokens.

    Why it matters: The model card gives concrete throughput and KV cache figures for a hybrid attention design, which helps readers weigh its long-context tradeoffs against full attention.

Oct 28, 2025

Oct 28, 2025Tue
  1. Cognition Blog (Devin, Windsurf)AI score72

    Cognition releases SWE-1.5, a coding agent model served at up to 950 tok/s

    AICognition has released SWE-1.5, a model optimized for software engineering that it says reaches near-frontier coding performance while running at up to 950 tok/s with Cerebras inference. The company reports it is 6x faster than Haiku 4.5 and 13x faster than Sonnet 4.5, and it is available now in Windsurf. The post's SWE-Bench Pro chart places SWE-1.5 at 40.08%, behind Sonnet 4.5 at 43.60%, and it notes that the model was trained with reinforcement learning on the Cascade agent harness.

    Why it matters: The post pairs a benchmark chart with a 950 tok/s speed claim and describes how harness, RL environments, and inference were co-designed, useful context for judging the speed-versus-quality tradeoff.

Oct 27, 2025

Oct 27, 2025Mon
  1. Cognition Blog (Devin, Windsurf)AI score36

    Devin Automates .NET Framework to .NET Core Migration in Weeks, Not Months

    AICognition says its autonomous coding agent Devin can complete a .NET Framework to .NET Core migration in as little as two weeks, using a Strangler Fig approach adapted from Jimmy Bogard's guide. The post says Devin handles planning via Ask Devin and DeepWiki, dependency sharing, controller and view conversion, and session state adaptation through a remote app.

Oct 26, 2025

Oct 26, 2025Sun
  1. Factory NewsAI score36

    AWS and Factory Announce Partnership, Factory Available on AWS Marketplace

    AIFactory has announced a partnership with Amazon Web Services and made its Droids agent platform available on the AWS Marketplace. Enterprise teams can use existing AWS Enterprise Discount Program commitments to buy Factory, with Droids accessible from CLI, Terminal UI, Web, Slack, Linear, and an IDE overlay. The source cites 31× faster feature development, 96.1%+ reduction in migration times, and 95.8% reduction in incident resolution times.

Oct 15, 2025

Oct 15, 2025Wed
  1. Cognition Blog (Devin, Windsurf)AI score73

    Cognition releases SWE-grep models for fast parallel code context retrieval

    AICognition introduces SWE-grep and SWE-grep-mini, fast agentic models trained with reinforcement learning for multi-turn context retrieval in coding tasks. The company says they match frontier coding models at retrieval while taking an order of magnitude less time, and they power the Fast Context subagent in Windsurf. The models issue up to 8 parallel tool calls per turn within 4 turns, and Cerebras serves SWE-grep-mini at over 2,800 tokens per second and SWE-grep at over 650 tokens per second.

    Why it matters: The post explains the speed-intelligence tradeoff in agentic code search, showing how parallel tool calls and RL training change the cost of retrieving context for coding agents.

Sep 28, 2025

Sep 28, 2025Sun
  1. Cognition Blog (Devin, Windsurf)AI score72

    Cognition rebuilds Devin around Claude Sonnet 4.5 for 2x speed

    AICognition rebuilt its Devin coding agent for Claude Sonnet 4.5, reporting 2x faster performance and 12% better results on its Junior Developer Evals, now available in Agent Preview. The team found the model is aware of its context window, which led to premature wrap-up behavior that they countered with repeated prompts and a 200k usage cap within a 1M token beta.

    Why it matters: The post explains which agent behaviors changed under Sonnet 4.5, such as context-window awareness and note-taking, that forced a rebuild rather than a simple model swap.

Sep 7, 2025

Sep 7, 2025Sun
  1. Cognition Blog (Devin, Windsurf)AI score53

    Cognition raises over $400M at $10.2B valuation after Windsurf acquisition

    AICognition, maker of the AI software engineer Devin, raised over $400M at a $10.2B post-money valuation led by Founders Fund. The company says its acquisition of Windsurf more than doubled its ARR, with combined enterprise ARR up over 30% in the seven weeks after the deal. It also reports Devin ARR grew from $1M in September 2024 to $73M in June 2025, with total net burn under $20M.

Sep 3, 2025

Sep 3, 2025Wed
  1. Cognition Blog (Devin, Windsurf)AI score38

    Eight Sleep Uses Devin AI as Data Analyst to Clear Ad-Hoc Requests

    AIEight Sleep integrated Cognition's Devin into its data workflows, letting staff tag Devin in Slack to query Snowflake, dbt, and Looker and check Amplitude. The company says it is now shipping 3x as many data features and investigations each week, with its ad-hoc data request queue near zero. Devin was used to trace a suspicious revenue spike to a better-than-expected email campaign.

Aug 4, 2025

Aug 4, 2025Mon
  1. Cognition Blog (Devin, Windsurf)AI score22

    Devin Can Automate Migrating Jenkins Pipelines to GitHub Actions at Scale

    AICognition says generative AI agents such as Devin can read internal docs and convert Jenkins pipelines to GitHub Actions syntax, replacing custom plugins with Actions or APIs and validating the results. The company claims enterprises can cut multi-year migration efforts to a few months, with Devin running inside the customer's secure environment so code and secrets stay internal.

Jul 21, 2025

Jul 21, 2025Mon
  1. Cognition Blog (Devin, Windsurf)AI score50

    Devin Adds MCP Support and a Marketplace for Connecting External Servers

    AICognition's Devin is now compatible with the Model Context Protocol (MCP), letting users connect favorite MCP servers through a new MCP Marketplace found in Settings. The source cites example uses including querying Datadog and Sentry logs, creating Notion docs, Google Docs, and Linear tickets, and interacting with Figma, Airtable, Stripe, and Hubspot.

Jul 13, 2025

Jul 13, 2025Sun
  1. Cognition Blog (Devin, Windsurf)AI score62

    Cognition signs definitive agreement to acquire Windsurf agentic IDE

    AICognition has signed a definitive agreement to acquire Windsurf, including its IP, product, trademark, brand, and team. The deal includes Windsurf's IDE, now with full access to the latest Claude models, and $82M of ARR with enterprise ARR doubling quarter-over-quarter. Cognition says it will invest heavily in integrating Windsurf's capabilities into its products while Windsurf continues operating as before.

    Why it matters: The post shows how an acquisition of an AI coding IDE was structured around staff treatment and product integration, with specific ARR and customer figures.

Jun 26, 2025

Jun 26, 2025Thu

Jun 22, 2025

Jun 22, 2025Sun
  1. Cognition Blog (Devin, Windsurf)AI score57

    Cognition details blockdiff, an open-source file format for instant VM disk snapshots

    AICognition built and open-sourced blockdiff, a file format that creates block-level diffs of VM disks using only filesystem metadata. The company reports its otterlink hypervisor cut snapshot times from 30 to 60 minutes on EC2 to about 5 to 10 seconds for a 128 GB disk with a 5 GB diff, roughly a 200x speedup. The post also covers the sparse file and copy-on-write concepts behind the approach and why OverlayFS, ZFS, and qcow2 were not chosen.

Jun 11, 2025

Jun 11, 2025Wed
  1. Cognition Blog (Devin, Windsurf)AI score62

    Cognition argues multi-agent architectures are fragile and proposes context-sharing principles

    AICognition argues that parallel multi-agent architectures are fragile because subagents act on conflicting, unshared assumptions. It proposes two principles for reliable agents: share context and full agent traces, and treat actions as carrying implicit decisions. The post recommends simpler single-threaded designs for most cases and notes that context compression and fine-tuned models can extend long-running tasks.

    Why it matters: The post explains concrete failure modes of parallel multi-agent setups and offers two context-sharing principles, useful for anyone designing long-running agent systems.

May 21, 2025

May 21, 2025Wed
  1. Cognition Blog (Devin, Windsurf)AI score62

    Cognition launches official DeepWiki MCP server for indexed GitHub repos

    AICognition launched the official DeepWiki Model Context Protocol server, which is free and requires no login or authentication. It gives programmatic access to ask_question, read_wiki_contents, and read_wiki_structure for GitHub repositories indexed on DeepWiki.com. Private repositories require a Devin account with GitHub connected, and open-source maintainers can apply for $500 in Devin credits.

    Why it matters: The source names the three tools and the access path, showing how indexed GitHub repositories can be queried programmatically without login.