Skip to contentSkip to stories

Updated

#Agent

May 20

May 20Wed
  1. Cognition Blog (Devin, Windsurf)AI score46

    Devin Gains Native Windows Environment for Building, Testing, and Migrating Apps

    AICognition's Devin AI software engineer can now build, run, and test code natively in its own Windows virtual machine, including migrating .NET Framework apps to .NET Core. The Windows capability is in beta for Enterprise Cloud and Dedicated Deployment customers, with the same SOC 2 Type II and ISO 27001 controls as the Linux version. Citi and Mercedes-Benz are named as existing Devin users.

May 19

May 19Tue
  1. koray kavukcuogluAI score72

    Google's Gemini 3.5 Flash beats Gemini 3.1 Pro on coding and agentic benchmarks

    AIGoogle's Gemini 3.5 Flash outperforms Gemini 3.1 Pro on Terminal-Bench 2.1 (76.2%), GDPval-AA (1656 Elo), and MCP Atlas (83.6%). The post also claims it is 4x faster than other frontier models, or 12x in Antigravity, and reports 83.6% on MMMU-Pro for multimodal performance.

    Why it matters: The post gives specific benchmark scores against Gemini 3.1 Pro, letting readers compare coding, agentic, and multimodal results directly.

  2. koray kavukcuogluAI score62

    Google introduces Gemini 3.5 Flash, used with agents to rebuild AlphaZero

    AIAt Google I/O, Google introduced Gemini 3.5 Flash, which the author says has become part of the daily research cycle. The author says a team of agents in Antigravity 2.0 recreated the original AlphaZero paper and built a playable web version from two prompts, coding the reinforcement learning pipeline in JAX/Flax and training a ResNet model via self-play on multi-TPU pods.

May 18

May 18Mon
  1. Michael TruellAI score40

    Cursor's Composer 2.5 is a significant upgrade over Composer 2

    AIMichael Truell of Cursor says Composer 2.5 is a significant step up from Composer 2. He adds that this is only the start of work with SpaceXAI, with more improvements expected soon. Cursor's announcement describes the model as more intelligent, better at long-running tasks, and more reliable at complex instructions, with doubled included usage for the next week.

  2. Eugene YanAI score62

    Cloudflare outlines an eight-stage agent harness for vulnerability discovery

    AIEugene Yan shares Cloudflare's description of a vulnerability discovery harness that runs eight stages, from reconnaissance to report writing. The pipeline uses about 50 concurrent agents to hunt for bugs, independent agents to try to disprove findings, and a trace step to confirm whether attacker input reaches each bug. Reachable findings feed back into new hunt tasks before a report is written against a predefined schema.

May 17

May 17Sun
  1. Cognition Blog (Devin, Windsurf)AI score60

    Cognition launches Auto-Triage, letting Devin investigate alerts and open fixes

    AICognition has released Auto-Triage in Devin Automations, which lets Devin respond to Slack messages, Linear events, GitHub activity, schedules, and webhooks. Devin can investigate with connected observability tools and the codebase, then post a summary, tag an owner, or open a PR. Devin runs in network-sandboxed environments with added protections against prompt injection and data exfiltration, and a limited-time offer gives $200 in credits for a first automation.

    Why it matters: The post shows how an agent handles alerts and bug reports from existing team channels, a practical pattern for teams weighing automated incident response.

May 14

May 14Thu

May 13

May 13Wed
  1. Eugene YanAI score72

    Mythos completes 32-step network attack in six of ten UK AISI trials

    AIEugene Yan relays two evaluations of Mythos: UK AISI reports it completed a 32-step network attack, estimated at about 20 expert hours, in 6 of 10 tries and was the first model to solve its end-to-end cyber ranges. XBOW's evaluation describes its performance as token-for-token and unprecedented in precision. The post links both AISI and XBOW blog posts for details.

May 11

May 11Mon
  1. Mira MuratiAI score40

    Thinking Machines launches interaction models built around human-AI collaboration

    AIThinking Machines, founded to advance human-AI collaboration, says its first bet is interactivity built into the model rather than added as scaffolding around a turn-based core. The company argues that how people work with AI matters as much as how intelligent the model is, and that interactivity should scale with intelligence. The post links to a blog detailing these interaction models.

May 10

May 10Sun
  1. Cognition Blog (Devin, Windsurf)AI score39

    Devin Automates HIL/SIL Failure Triage and Scales Test Generation at Automotive Firms

    AICognition reports that deploying its Devin agent on hardware-in-the-loop and software-in-the-loop workflows cut failure triage time and multiplied test generation at automotive customers. One team reclaimed 2K–4K engineering hours monthly across about 4,000 tickets, while RV Tech rose from 1–2 to 10–15 generated tests per day. Devin also helps convert bottlenecked HIL tests into SIL equivalents to catch failures earlier.

May 9

May 9Sat
  1. PaddlePaddleAI score60

    Baidu releases ERNIE 5.1 with reduced pretraining cost and parameter scale

    AIBaidu's PaddlePaddle account announced ERNIE 5.1, which it says cuts total parameters to about one-third and activated parameters to about one-half, using roughly 6% of the pretraining cost of models at similar scale. The post reports benchmark results including 99.6 on AIME26 with tools, surpassing DeepSeek-V4-Pro on τ3-bench and SpreadsheetBench-Verified, and ranking #4 globally on Arena Search. ERNIE 5.1 is available through the ERNIE website and Baidu AI Studio Model Playground.

May 8

May 8Fri

May 5

May 5Tue
  1. Eugene YanAI score14

    Eugene Yan shares five principles for working with AI models

    AIEugene Yan outlines five principles for working effectively with AI models: treating context as infrastructure, taste as configuration, verification as the basis for autonomy, scaling through delegation, and closing the loop. The post is a short list of themes linked to a longer essay, and no further detail is given in the post itself.

May 1

May 1Fri

Apr 30

Apr 30Thu
  1. OpenAI Alignment Research BlogAI score79

    OpenAI's Auto-review lets Codex agents act without constant human approval

    AIOpenAI released Auto-review in Codex, which replaces user approval at the sandbox boundary with a separate agent that approves or denies boundary-crossing actions. In internal deployment, Codex sessions stopped for human approval about 200x less often than in manual mode, and Auto-review approved around 99% of escalated actions. The post also states that Auto-review is not a guarantee of security and cannot protect against model scheming.

    Why it matters: The post explains how Auto-review replaces human approval at the sandbox boundary, with internal deployment figures and stated limits that help readers judge the tradeoff for coding agents.

  2. Andrej KarpathyAI score66

    Karpathy on agentic engineering, Software 3.0, and jagged AI capability

    AIAndrej Karpathy describes a December 2025 shift in which coding agents began producing larger, more reliable chunks of work, changing programming toward orchestrating agents. He argues that models automate what can be verified and that their capability is jagged, depending on verifiability and what labs emphasize in training, so users need to stay in the loop. He also says hiring, founder opportunities, and agent-native infrastructure should adapt to this shift.

Apr 29

Apr 29Wed
  1. Cognition Blog (Devin, Windsurf)AI score34

    Cognition opens Singapore headquarters for Asia-Pacific push with Devin

    AICognition has opened its Asia-Pacific headquarters in Singapore to expand its autonomous software engineering platform, Devin, across the region. The company says OCBC saw up to 30% improvement in code and test case generation, and its system integration test first-pass rate rose from below 50% to over 80% after deployment. Cognition is building its Singapore team across engineering, go-to-market, and partnerships, with Richard Spence leading APAC.

Apr 28

Apr 28Tue

Apr 27

Apr 27Mon
  1. Cognition Blog (Devin, Windsurf)AI score38

    Mercedes-Benz Deploys Devin and Windsurf Across Global Engineering Teams

    AIMercedes-Benz is deploying Cognition's Devin and Windsurf across its global engineering teams, from the United States to Europe and Asia. In a four-week pilot, Devin analyzed over 200,000 lines of COBOL code and cut modernization time from an estimated eight months to eight days. The company is now rolling out the full suite, with Windsurf for development, Devin as an autonomous cloud agent, and Devin for Terminal for the most complex tasks.

  2. Xiaomi MiMo · new models on Hugging FaceAI score72

    Xiaomi releases MiMo-V2.5, an open omnimodal model with 1M context

    AIXiaomi's MiMo-V2.5 is a native omnimodal model that understands text, image, video, and audio within one architecture. It is a sparse MoE with 310B total and 15B activated parameters, and supports up to 1M tokens of context. The repository also notes a config.json and tokenizer_config.json update that users who downloaded before commit 4da2748 should re-pull.

    Why it matters: The repository documents a 310B-parameter omnimodal MoE with a hybrid attention design, useful for comparing long-context efficiency against other open multimodal models.

  3. Mistral AI · new models on Hugging FaceAI score36

    Mistral Medium 3.5 EAGLE draft model released for speculative decoding on Hugging Face

    AIMistral AI has released mistralai/Mistral-Medium-3.5-128B-EAGLE, an EAGLE draft model for speculative decoding with the 128B dense Mistral Medium 3.5. The companion model, which the source says replaces Mistral Medium 3.1 and Magistral in Le Chat and Devstral 2 in Vibe, has a 256k context window, handles text and image input with text output, and is served with vLLM or SGLang using three speculative tokens. The model is released under a Modified MIT License that allows commercial use with exceptions for companies with large revenue.

  4. Soumith ChintalaAI score15

    Chintala Suggests Anthropic Account Support May Need Scaling Up

    AISoumith Chintala comments on a Reddit report that Anthropic banned organizations without warning, suggesting Anthropic may need to scale Account Support using Claude or human account managers. He also argues that enterprises may increasingly adopt multiple AI providers with open harnesses, facing cloud-era vendor problems that would likely affect all AI providers.

Apr 26

Apr 26Sun
  1. Xiaomi MiMoAI score87

    Xiaomi releases open-source MiMo-V2.5-Pro for long-horizon agentic coding

    AIXiaomi released and open-sourced MiMo-V2.5-Pro, a 1.02T-parameter Mixture-of-Experts model with 42B active parameters and a 1M-token context window. The company reports gains in agentic tasks, software engineering, and long-horizon work, including a Rust SysY compiler task finished in 4.3 hours across 672 tool calls. Weights and tokenizer are on Hugging Face, and API pricing is unchanged.

    Why it matters: The release pairs a 1.02T-parameter open-weight model with long-horizon agent results and token-efficiency claims, useful for judging its fit in coding and agent workflows.

Apr 24

Apr 24Fri
  1. Ahmad Al-DahleAI score82

    Ahmad Al-Dahle says DeepSeek-V4's efficient 1M context is its key bet

    AIAhmad Al-Dahle argues that the most interesting part of DeepSeek-V4 is its bet on efficient ultra-long context rather than its benchmarks. He says this is the precondition for test-time scaling and long-horizon agents, and cites 27% of V3's FLOPs at 1M tokens. The quoted DeepSeek post announces DeepSeek-V4-Pro (1.6T total, 49B active) and DeepSeek-V4-Flash (284B total, 13B active), both open-sourced with 1M context and API access.

    Why it matters: The post argues that efficient 1M-token context, not benchmark scores, is the key bet behind DeepSeek-V4's design for test-time scaling and long-horizon agents.

Apr 22

Apr 22Wed
  1. Factory NewsAI score38

    Factory's Automated QA Skill Tests Apps Like Real Users and Posts Reports to PRs

    AIFactory has released an Automated QA skill that drives an app as a real user would, filling forms, typing into terminals, and calling endpoints, then posts a structured report with screenshots, terminal snapshots, and API traces as a single updating comment on each pull request. Teams can run it on every push or make it an optional CI check triggered by a PR label, comment command, or manual dispatch, and developers can run /qa locally in any Droid session. Automated QA is available today in all Factory plans.

  2. Cognition Blog (Devin, Windsurf)AI score54

    Cognition says building cloud agents requires VM isolation, state snapshots, and org change

    AICognition argues that enterprises building cloud agents face three problems: shared container kernels, the inability to persist agent state across async gaps, and the scale of orchestration, governance, and integrations. The post says VM-level isolation with hypervisor-level snapshots was needed for Devin, and that organizations must also rebuild engineering processes around agent execution.

Apr 21

Apr 21Tue
  1. Cognition Blog (Devin, Windsurf)AI score72

    Cognition says multi-agent systems work when only one agent writes

    AICognition reports that multi-agent setups work best when writes stay single-threaded and extra agents contribute intelligence instead of actions. It describes a code-review loop where a clean-context review agent catches bugs in Devin-written PRs, averaging 2 bugs per PR with roughly 58% severe. The post also says the smart-friend pattern, pairing a smaller primary model with a stronger one, has not yet worked well with asymmetrically weaker primaries and is an open training problem.

    Why it matters: The post gives concrete findings on which multi-agent setups work, including clean-context code review and smart-friend escalation, and where they still fail.

  2. Xiaomi MiMoAI score67

    Xiaomi releases MiMo-V2.5, an open multimodal agent model with 1M context

    AIXiaomi released MiMo-V2.5, a 310B-parameter sparse MoE model with 15B active parameters that adds native visual and audio understanding. The model supports up to 1 million tokens of context, and its weights, tokenizer, and model card are available on Hugging Face. Xiaomi says it surpasses MiMo-V2-Pro on agentic performance and reports a Claw-Eval score of 62.3 on the general subset.

    Why it matters: The release pairs native visual and audio understanding with a 1M-token context window and open weights, a combination worth checking against your own multimodal workflows.

  3. Michael TruellAI score62

    Cursor partners with SpaceX to scale up Composer, with an option to acquire

    AICursor's Michael Truell says the company is partnering with the SpaceX team to scale up Composer, calling it a meaningful step toward building the best place to code with AI. The quoted SpaceX post says Cursor gives SpaceX the right to acquire Cursor later this year for $60 billion, or pay $10 billion for the work together. It also cites SpaceX's Colossus training supercomputer, described as a million H100-equivalent system, as a source of training capacity.