Skip to contentSkip to stories

Updated

All AI news

Items with an AI score under 20 are hidden. Show low-relevance items

Apr 30

Apr 30Thu
  1. OpenAI Alignment Research BlogAI score79

    OpenAI's Auto-review lets Codex agents act without constant human approval

    AIOpenAI released Auto-review in Codex, which replaces user approval at the sandbox boundary with a separate agent that approves or denies boundary-crossing actions. In internal deployment, Codex sessions stopped for human approval about 200x less often than in manual mode, and Auto-review approved around 99% of escalated actions. The post also states that Auto-review is not a guarantee of security and cannot protect against model scheming.

    Why it matters: The post explains how Auto-review replaces human approval at the sandbox boundary, with internal deployment figures and stated limits that help readers judge the tradeoff for coding agents.

Apr 29

Apr 29Wed
  1. Cognition Blog (Devin, Windsurf)AI score34

    Cognition opens Singapore headquarters for Asia-Pacific push with Devin

    AICognition has opened its Asia-Pacific headquarters in Singapore to expand its autonomous software engineering platform, Devin, across the region. The company says OCBC saw up to 30% improvement in code and test case generation, and its system integration test first-pass rate rose from below 50% to over 80% after deployment. Cognition is building its Singapore team across engineering, go-to-market, and partnerships, with Richard Spence leading APAC.

Apr 27

Apr 27Mon
  1. Cognition Blog (Devin, Windsurf)AI score38

    Mercedes-Benz Deploys Devin and Windsurf Across Global Engineering Teams

    AIMercedes-Benz is deploying Cognition's Devin and Windsurf across its global engineering teams, from the United States to Europe and Asia. In a four-week pilot, Devin analyzed over 200,000 lines of COBOL code and cut modernization time from an estimated eight months to eight days. The company is now rolling out the full suite, with Windsurf for development, Devin as an autonomous cloud agent, and Devin for Terminal for the most complex tasks.

  2. Xiaomi MiMo · new models on Hugging FaceAI score72

    Xiaomi releases MiMo-V2.5, an open omnimodal model with 1M context

    AIXiaomi's MiMo-V2.5 is a native omnimodal model that understands text, image, video, and audio within one architecture. It is a sparse MoE with 310B total and 15B activated parameters, and supports up to 1M tokens of context. The repository also notes a config.json and tokenizer_config.json update that users who downloaded before commit 4da2748 should re-pull.

    Why it matters: The repository documents a 310B-parameter omnimodal MoE with a hybrid attention design, useful for comparing long-context efficiency against other open multimodal models.

  3. Mistral AI · new models on Hugging FaceAI score36

    Mistral Medium 3.5 EAGLE draft model released for speculative decoding on Hugging Face

    AIMistral AI has released mistralai/Mistral-Medium-3.5-128B-EAGLE, an EAGLE draft model for speculative decoding with the 128B dense Mistral Medium 3.5. The companion model, which the source says replaces Mistral Medium 3.1 and Magistral in Le Chat and Devstral 2 in Vibe, has a 256k context window, handles text and image input with text output, and is served with vLLM or SGLang using three speculative tokens. The model is released under a Modified MIT License that allows commercial use with exceptions for companies with large revenue.

Apr 26

Apr 26Sun
  1. Xiaomi MiMoAI score87

    Xiaomi releases open-source MiMo-V2.5-Pro for long-horizon agentic coding

    AIXiaomi released and open-sourced MiMo-V2.5-Pro, a 1.02T-parameter Mixture-of-Experts model with 42B active parameters and a 1M-token context window. The company reports gains in agentic tasks, software engineering, and long-horizon work, including a Rust SysY compiler task finished in 4.3 hours across 672 tool calls. Weights and tokenizer are on Hugging Face, and API pricing is unchanged.

    Why it matters: The release pairs a 1.02T-parameter open-weight model with long-horizon agent results and token-efficiency claims, useful for judging its fit in coding and agent workflows.

Apr 24

Apr 24Fri
  1. DeepSeek API NewsAI score67

    DeepSeek API adds V4-Pro and V4-Flash, retiring legacy model names in July 2026

    AIThe DeepSeek API now supports V4-Pro and V4-Flash through both the OpenAI ChatCompletions and Anthropic interfaces. Developers keep the same base_url and set the model parameter to deepseek-v4-pro or deepseek-v4-flash. The legacy names deepseek-chat and deepseek-reasoner will be discontinued on 2026-07-24, and until then they map to the non-thinking and thinking modes of deepseek-v4-flash, respectively.

    Why it matters: The source gives exact model names, an unchanged base URL, and a July 2026 discontinuation date, so developers can plan their migration from legacy names.

Apr 23

Apr 23Thu
  1. Apple · new models on Hugging FaceAI score40

    Apple releases CADD-Base-7B, a masked diffusion model for code generation

    AIApple has released CADD-Base-7B on Hugging Face, a 7B masked diffusion language model for code generation that uses Continuously Augmented Discrete Diffusion (CADD) to guide discrete denoising with a continuous flow-matching signal. The model loads through Transformers with trust_remote_code, and its diffusion_generate method supports CADD sampling modes "weighted" and "argmax" with alg options such as "entropy" and "maskgit_plus". The release builds on DiffuCoder and reuses Dream's modeling architecture and generation utilities.

  2. OpenAI Alignment Research BlogAI score44

    OpenAI Open-Sources Chain-of-Thought Monitorability Evaluation Datasets and Code

    AIOpenAI is releasing a subset of datasets, reference code, and the g-mean 2 metric for evaluating chain-of-thought monitorability. The release includes most datasets from its monitorability suite, while some evaluations relying on private or restricted data were omitted. The company says it will keep reporting monitorability results in future frontier reasoning model system cards.

Apr 22

Apr 22Wed
  1. Factory NewsAI score38

    Factory's Automated QA Skill Tests Apps Like Real Users and Posts Reports to PRs

    AIFactory has released an Automated QA skill that drives an app as a real user would, filling forms, typing into terminals, and calling endpoints, then posts a structured report with screenshots, terminal snapshots, and API traces as a single updating comment on each pull request. Teams can run it on every push or make it an optional CI check triggered by a PR label, comment command, or manual dispatch, and developers can run /qa locally in any Droid session. Automated QA is available today in all Factory plans.

  2. Cognition Blog (Devin, Windsurf)AI score54

    Cognition says building cloud agents requires VM isolation, state snapshots, and org change

    AICognition argues that enterprises building cloud agents face three problems: shared container kernels, the inability to persist agent state across async gaps, and the scale of orchestration, governance, and integrations. The post says VM-level isolation with hypervisor-level snapshots was needed for Devin, and that organizations must also rebuild engineering processes around agent execution.

  3. Anthropic EngineeringAI score78

    Anthropic traces Claude Code quality complaints to three product changes

    AIAnthropic says three changes to Claude Code, the Claude Agent SDK, and Claude Cowork caused recent quality complaints, and the API was not affected. The fixes were resolved by April 20 (v2.1.116), and the company is resetting usage limits for all subscribers as of April 23.

    Why it matters: The postmortem traces three separate changes to specific dates and versions, showing how a bug in context management can look like broad degradation to users.

Apr 21

Apr 21Tue
  1. Cognition Blog (Devin, Windsurf)AI score72

    Cognition says multi-agent systems work when only one agent writes

    AICognition reports that multi-agent setups work best when writes stay single-threaded and extra agents contribute intelligence instead of actions. It describes a code-review loop where a clean-context review agent catches bugs in Devin-written PRs, averaging 2 bugs per PR with roughly 58% severe. The post also says the smart-friend pattern, pairing a smaller primary model with a stronger one, has not yet worked well with asymmetrically weaker primaries and is an open training problem.

    Why it matters: The post gives concrete findings on which multi-agent setups work, including clean-context code review and smart-friend escalation, and where they still fail.

  2. Xiaomi MiMoAI score67

    Xiaomi releases MiMo-V2.5, an open multimodal agent model with 1M context

    AIXiaomi released MiMo-V2.5, a 310B-parameter sparse MoE model with 15B active parameters that adds native visual and audio understanding. The model supports up to 1 million tokens of context, and its weights, tokenizer, and model card are available on Hugging Face. Xiaomi says it surpasses MiMo-V2-Pro on agentic performance and reports a Claw-Eval score of 62.3 on the general subset.

    Why it matters: The release pairs native visual and audio understanding with a 1M-token context window and open weights, a combination worth checking against your own multimodal workflows.

Apr 20

Apr 20Mon
  1. Berkeley AI ResearchAI score44

    GRASP: A Gradient-Based Planner for Long-Horizon World Model Planning

    AIBerkeley AI Research introduces GRASP, a gradient-based planner for learned world models that aims to make long-horizon planning more robust. GRASP lifts trajectories into virtual states for parallel optimization across time, adds stochasticity to state iterates for exploration, and reshapes gradients to avoid brittle state-input gradients through high-dimensional vision models. The post identifies ill-conditioned gradients and non-greedy loss landscapes as core failure modes of standard rollout-based planning.

Apr 17

Apr 17Fri
  1. OpenAI · new models on Hugging FaceAI score41

    OpenAI Releases Privacy Filter, an Open-Weight PII Detection Model on Hugging Face

    AIOpenAI released Privacy Filter, a bidirectional token-classification model that detects and masks personally identifiable information in text under the Apache 2.0 license. The model has 1.5B total parameters with 50M active, supports a 128,000-token context window, and can run in a web browser or on a laptop. Users can fine-tune it and adjust precision/recall tradeoffs through preset operating points.

Apr 14

Apr 14Tue
  1. Moonshot AI (Kimi) · new models on Hugging FaceAI score78

    Moonshot AI releases open-source Kimi K2.6 multimodal agentic model

    AIMoonshot AI released Kimi K2.6, an open-source native multimodal agentic model with 1T total and 32B activated parameters and a 256K context length. The model card reports benchmark results against GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro across agentic, coding, reasoning, and vision tasks, and supports swarms of up to 300 sub-agents.

    Why it matters: The model card gives specific agent swarm scale, context length, and benchmark comparisons against several frontier models, useful for judging its coding and agent capabilities.

Apr 13

Apr 13Mon
  1. ARC PrizeAI score58

    ARC Prize Releases Human Performance Dataset for ARC-AGI-3 Benchmark

    AIARC Prize Foundation released an open-source human dataset for ARC-AGI-3, covering 342 step-by-step replays across 25 public environments from a study of 458 participants. The source reports that every environment was solved by at least two humans, and it updates scoring by moving the per-level baseline to the median human player and raising the per-level cap from 100% to 115%.

  2. Cognition Blog (Devin, Windsurf)AI score62

    Cognition introduces SWE-check, a fast RL-trained bug detection model for Windsurf

    AICognition and Applied Compute RL-trained SWE-check, a specialized bug detection model for the Windsurf IDE. It matches frontier performance on in-distribution evals and is an order of magnitude faster with cheaper inference, though it trails frontier models on out-of-distribution evals (delta F1 0.29 versus 0.49 before training). A preview is available in Windsurf Next, with a mainstream release planned.

    Why it matters: The post explains how production environment replication, reward linearization, and two-phase post-training trade bug-detection quality against latency for an IDE specialist model.

  3. Cognition Blog (Devin, Windsurf)AI score49

    Devin Introduces New Self-Serve Plans and Charges for Ask Devin and Devin Review

    AIDevin is retiring its Core and Team plans for a new lineup of Free, Pro at $20/month, Max at $200/month, Teams with usage-based billing and an $80/month minimum, and custom-priced Enterprise. Ask Devin's Deep Mode, Devin Review after a 2-week free trial, and higher-quality DeepWiki generation will move to usage-based billing, with DeepWiki's existing generation and open-source Devin Review remaining free. Self-serve usage beyond included quota will be billed in dollars rather than ACUs.

Apr 8

Apr 8Wed
  1. MiniMax · new models on Hugging FaceAI score78

    MiniMax releases open-weight MiniMax-M2.7 with agent and coding gains

    AIMiniMax has released MiniMax-M2.7 on Hugging Face, describing it as its first model to participate in its own evolution. The source reports 56.22% on SWE-Pro, 46.3% on Toolathon, and 62.7% on MM ClawBench, and says an internal version autonomously optimized a programming scaffold over 100+ rounds for a 30% performance improvement.

    Why it matters: The source ties its benchmark claims to a self-evolution process and a named comparison set, which helps readers weigh how the reported gains were achieved.

  2. Cognition Blog (Devin, Windsurf)AI score31

    Cognition Expands to Japan, Appoints Takumi Masai to Lead Devin Launch

    AICognition is expanding into Japan, its first step into Asia, and has appointed Takumi Masai as Japan President and General Manager to lead a local team working with Japanese enterprises. DeNA has used Devin to more than double operational efficiency across multiple engineering functions, and Mizuho Securities has deployed it as one of the first large-scale financial institutions in Japan.

  3. Stability AIAI score44

    Stability AI launches Brand Studio, a creative production platform built around brand identity

    AIStability AI has introduced Brand Studio, an end-to-end creative production platform for enterprise teams that builds around each brand's identity. Its Brand Central hub supports custom Brand ID models and Campaigns, while Producer Mode turns prompts into step-by-step production plans. Curated Model Routing selects models including Stable Diffusion, Nano Banana, and Seedream, and new Precision Inpainting and Product Insertion tools enable targeted edits.

Apr 7

Apr 7Tue
  1. Cognition Blog (Devin, Windsurf)AI score70

    How Devin Is Modernizing COBOL at Fortune 500 Companies

    AICognition describes how Devin handles COBOL modernization at several Fortune 500 companies, citing a shortage of COBOL developers and 68% failure rates for such efforts. The post identifies three obstacles for agents: untraceable data across copybooks, little COBOL in model training, and no way to run code on Linux-based VMs. It says Devin succeeds on documentation, batch migrations, and large-scale refactoring, while transactional workloads remain out of reach.

    Why it matters: The post explains why agents struggle with COBOL and which workloads they can migrate, giving a framework for judging where automation fits legacy systems.

  2. Anthropic EngineeringAI score67

    Anthropic decouples agent brain, hands, and session in Managed Agents

    AIAnthropic's Managed Agents separates the harness, sandbox, and session into independently replaceable interfaces. The source says this design let failed containers be replaced, kept tokens out of the sandbox, and reduced p50 time-to-first-token by roughly 60% and p95 by over 90%.

    Why it matters: The post explains how decoupling the harness, sandbox, and session changed failure recovery, credential security, and latency, offering a reusable architecture pattern for long-running agents.

Apr 6

Apr 6Mon
  1. Z.ai Release NotesAI score34

    Z.ai's GLM-5.3 and GLM-5.2 Lead Open-Source Coding and Long-Context Models

    AIZ.ai's GLM-5.3 delivers a 50% coding gain over GLM-5.2 on Z.ai Code Bench, reaching open-source state-of-the-art on public benchmarks including Terminal Bench 3.0. GLM-5.3-Flash uses 320B total parameters with 18B activated, combining linear and sparse attention to reduce compute and KV-cache needs. GLM-5.2 supports a 1M lossless context window for long-horizon tasks.

  2. Cognition Blog (Devin, Windsurf)AI score44

    Windsurf releases SWE-1.6, a software engineering model optimized for speed and user experience

    AIWindsurf has made SWE-1.6, its model for software engineering agents, generally available, with the company saying it improves on the SWE-1.6 Preview by reducing overthinking, looping, and sequential tool calls. The model is free for three months, with a free version offered at 200 tok/s through Fireworks and a faster paid version at 950 tok/s through Cerebras.

  3. Black Forest Labs · new models on Hugging FaceAI score41

    FLUX.2 Small Decoder offers faster, lower-VRAM drop-in replacement for FLUX.2 decoder

    AIBlack Forest Labs released FLUX.2 Small Decoder, a distilled VAE decoder that works as a drop-in replacement for the standard FLUX.2 decoder on Hugging Face. It decodes about 1.4x faster and uses about 1.4x less VRAM at decode time, with ~28M decoder parameters versus ~50M in the full decoder and minimal quality loss. It is available under the Apache 2.0 license and is compatible with FLUX.2-klein-4B, FLUX.2-klein-9B, FLUX.2-klein-9b-kv, and FLUX.2-dev.

  4. OpenAI Alignment Research BlogAI score31

    OpenAI opens applications for Safety Fellowship on AI safety and alignment research

    AIOpenAI announced applications for its Safety Fellowship, a pilot program supporting external researchers, engineers, and practitioners in safety and alignment research on advanced AI systems. The program runs from September 14, 2026 through February 5, 2027, with a monthly stipend, compute support, API credits, and mentorship, and fellows are expected to produce a substantial output such as a paper, benchmark, or dataset. Applications close May 3, and successful applicants will be notified by July 25.

Apr 3

Apr 3Fri
  1. Z.ai (GLM) · new models on Hugging FaceAI score73

    Z.ai releases GLM-5.1, a flagship model for agentic engineering

    AIZ.ai has released GLM-5.1, its next-generation flagship model for agentic engineering, with stronger coding than GLM-5. The model is described as staying effective over longer agentic tasks, sustaining optimization over hundreds of rounds and thousands of tool calls. The release lists benchmark results including SWE-Bench Pro at 58.4 and Terminal-Bench 2.0 at 63.5, and local deployment is supported through SGLang, vLLM, xLLM, Transformers, and KTransformers.

    Why it matters: The release gives benchmark tables against several rival models, letting readers compare GLM-5.1's coding and agentic results with GLM-5 and frontier systems.

Mar 31

Mar 31Tue
  1. Mistral AI · new models on Hugging FaceAI score76

    Mistral Medium 3.5 releases as a 128B dense merged model with vision

    AIMistral AI released Mistral Medium 3.5, a dense 128B model with a 256k context window that handles instruction-following, reasoning, and coding in a single set of weights. It replaces Mistral Medium 3.1, Magistral, and Devstral 2, and reasoning effort is configurable per request. The model accepts text and image input and is released under a Modified MIT License that excludes companies with large revenue.

    Why it matters: The release merges instruction, reasoning, and coding into one 128B model with per-request reasoning control, giving developers one set of weights to compare against separate specialized models.

  2. Alibaba NLP (Tongyi) · new models on Hugging FaceAI score26

    LaSER-Qwen3-8B: Alibaba NLP's 8B dense retriever with latent reasoning released on Hugging Face

    AIAlibaba NLP released LaSER-Qwen3-8B, an 8B-parameter dense retriever built on Qwen/Qwen3-8B that internalizes explicit reasoning into latent space through continuous latent thinking tokens. The model scores 29.3 nDCG@10 on the BRIGHT benchmark, ahead of the rewrite-then-retrieve pipeline's 28.1, and carries a 4096-dimension embedding with an 8192-token maximum sequence length. It is licensed under MIT and adds about 1.7× latency over standard single-pass dense retrievers.

Mar 24

Mar 24Tue
  1. ARC PrizeAI score70

    ARC Prize announces ARC-AGI-3, an interactive benchmark for frontier agents

    AIARC Prize has released ARC-AGI-3, a set of hundreds of interactive, turn-based environments with thousands of game-style levels, with no instructions or stated goals. Humans score 100% while frontier AI scores 0.51%. ARC Prize 2026 offers over $2 million in prizes for open-source solutions to ARC-AGI-2 and ARC-AGI-3.

    Why it matters: The benchmark's human versus frontier AI gap and its interactive design show how agent evaluation is shifting from instruction-following toward exploration and adaptation.

  2. Anthropic EngineeringAI score78

    How Anthropic built Claude Code auto mode to replace skipped permissions

    AIAnthropic describes Claude Code auto mode, which delegates approval of agent actions to model-based classifiers instead of manual prompts or skipped permissions. The classifier reviews tool calls before execution and a separate probe screens tool outputs for prompt injection. Anthropic reports a 0.4% false positive rate on real internal traffic and a 17% false negative rate on real overeager actions.

    Why it matters: The post explains the layered classifier design and its measured tradeoffs, showing how autonomous coding agents can cut approval fatigue without fully removing risk.

Mar 23

Mar 23Mon
  1. Anthropic EngineeringAI score78

    Anthropic shows a three-agent harness for long-running app development

    AIAnthropic's Labs team describes a three-agent harness with planner, generator, and evaluator agents for building full-stack applications over multi-hour autonomous coding sessions. The evaluator uses Playwright to test the running app against sprint contracts, and a retro game maker built with the harness worked end to end where a single-agent run's core feature did not. The author later removed the sprint construct and kept only the components still needed on Opus 4.6.

    Why it matters: The post shows how a generator-evaluator loop, with explicit grading criteria and a tuned QA agent, turned a solo run's broken output into a working app, and how the harness was pruned as models improved.

Mar 22

Mar 22Sun
  1. FunAudioLLM (Alibaba Tongyi) · new models on Hugging FaceAI score32

    PrismAudio Adds Reinforcement Learning to Video-to-Audio Generation with Chain-of-Thought Planning

    AIPrismAudio is a framework that integrates reinforcement learning into video-to-audio generation, using a Chain-of-Thought planning mechanism. It builds on ThinkSound by splitting single-step reasoning into four CoT modules for semantic, temporal, aesthetic, and spatial dimensions, each with targeted reward functions. Code, model weights, and datasets are released for research and educational use under the MIT License, and commercial use requires explicit author authorization.

Mar 19

Mar 19Thu
  1. Cognition Blog (Devin, Windsurf)AI score50

    Devin can now schedule recurring sessions that carry state between runs

    AIDevin can now schedule its own recurring sessions from a plain-language description, such as running a weekly feature-flag cleanup every Monday at 9am. Devin keeps its own notes across runs, so each scheduled session builds on earlier results rather than starting over. The feature can also be combined with Managed Devins to run parallel recurring tasks, such as a weekly QA pass reported to Slack.

Mar 18

Mar 18Wed
  1. Cognition Blog (Devin, Windsurf)AI score72

    Devin can now break tasks down and run a team of managed Devins

    AIDevin can now break large tasks into scoped pieces and delegate them to a team of managed Devins that run in parallel. Each managed Devin runs in its own isolated virtual machine with its own terminal, browser, and development environment, and has its own session link. The main coordinator session monitors progress, resolves conflicts, and compiles results, and managed Devins are available now for all users.

    Why it matters: The post explains how a coordinator session splits work across isolated managed sessions, giving readers a concrete pattern for running agent tasks in parallel.