Skip to contentSkip to stories

Updated

All AI news

May 6

May 6Wed

May 2

May 2Sat

Apr 30

Apr 30Thu
  1. Mark ChenAI score62

    OpenAI's Mark Chen says GPT-5.5 performs like Mythos in UK AISI cyber range

    AIMark Chen says GPT-5.5 performs similarly to Mythos on UK AISI's cyber range, which tests long-horizon, agentic capability, and calls it one eval rather than a full picture. He adds that frontier model risks are real and that OpenAI aims to deploy AI people can actually use through mitigations. The attached chart shows completed steps per cumulative token spent for GPT-5.5, Mythos Preview, and several Claude and GPT models, from M1 reconnaissance up to M9 full network takeover.

  2. koray kavukcuogluAI score12

    Alphabet named to TIME's 2026 Most Influential Companies list

    AIAlphabet has been recognized in TIME's 2026 Most Influential Companies list, with Google's Koray Kavukcuoglu crediting teams for turning AI breakthroughs into everyday user features. TIME's cover story says Sundar Pichai's 2016 "AI-first" push, including custom chips, Cloud, YouTube, and deep AI research, has paid off.

Apr 29

Apr 29Wed
  1. Cognition Blog (Devin, Windsurf)AI score34

    Cognition opens Singapore headquarters for Asia-Pacific push with Devin

    AICognition has opened its Asia-Pacific headquarters in Singapore to expand its autonomous software engineering platform, Devin, across the region. The company says OCBC saw up to 30% improvement in code and test case generation, and its system integration test first-pass rate rose from below 50% to over 80% after deployment. Cognition is building its Singapore team across engineering, go-to-market, and partnerships, with Richard Spence leading APAC.

Apr 27

Apr 27Mon
  1. Cognition Blog (Devin, Windsurf)AI score38

    Mercedes-Benz Deploys Devin and Windsurf Across Global Engineering Teams

    AIMercedes-Benz is deploying Cognition's Devin and Windsurf across its global engineering teams, from the United States to Europe and Asia. In a four-week pilot, Devin analyzed over 200,000 lines of COBOL code and cut modernization time from an estimated eight months to eight days. The company is now rolling out the full suite, with Windsurf for development, Devin as an autonomous cloud agent, and Devin for Terminal for the most complex tasks.

Apr 24

Apr 24Fri

Apr 23

Apr 23Thu

Apr 22

Apr 22Wed
  1. Anthropic EngineeringAI score78

    Anthropic traces Claude Code quality complaints to three product changes

    AIAnthropic says three changes to Claude Code, the Claude Agent SDK, and Claude Cowork caused recent quality complaints, and the API was not affected. The fixes were resolved by April 20 (v2.1.116), and the company is resetting usage limits for all subscribers as of April 23.

    Why it matters: The postmortem traces three separate changes to specific dates and versions, showing how a bug in context management can look like broad degradation to users.

Apr 21

Apr 21Tue
  1. Michael TruellAI score62

    Cursor partners with SpaceX to scale up Composer, with an option to acquire

    AICursor's Michael Truell says the company is partnering with the SpaceX team to scale up Composer, calling it a meaningful step toward building the best place to code with AI. The quoted SpaceX post says Cursor gives SpaceX the right to acquire Cursor later this year for $60 billion, or pay $10 billion for the work together. It also cites SpaceX's Colossus training supercomputer, described as a million H100-equivalent system, as a source of training capacity.

Apr 18

Apr 18Sat

Apr 17

Apr 17Fri

Apr 13

Apr 13Mon
  1. Mira MuratiAI score26

    Thinking Machines welcomes Workshop Labs founders Luke Drago and @LRudL_

    AIThinking Machines has welcomed Luke Drago and @LRudL_, who co-founded Workshop Labs to build AI that keeps humans relevant. Mira Murati says they will continue that mission at Thinking Machines, which builds powerful AI systems that think alongside humans and extend human agency. She ties the hire to the company's broader work, including Tinker, research grants, and advancing the frontier.

Apr 8

Apr 8Wed
  1. Cognition Blog (Devin, Windsurf)AI score31

    Cognition Expands to Japan, Appoints Takumi Masai to Lead Devin Launch

    AICognition is expanding into Japan, its first step into Asia, and has appointed Takumi Masai as Japan President and General Manager to lead a local team working with Japanese enterprises. DeNA has used Devin to more than double operational efficiency across multiple engineering functions, and Mizuho Securities has deployed it as one of the first large-scale financial institutions in Japan.

Apr 7

Apr 7Tue
  1. Dario AmodeiAI score72

    Dario Amodei backs Project Glasswing to counter AI-driven cyber threats

    AIDario Amodei said many of the world's leading companies have joined Project Glasswing, an effort to address cyber threats posed by increasingly capable AI systems. The initiative was introduced by Anthropic and is powered by its newest frontier model, Claude Mythos Preview, which the quoted post says can find software vulnerabilities better than all but the most skilled humans.

    Why it matters: The post gives a concrete example of how a frontier AI lab is organizing industry partners around AI-driven software vulnerability discovery.

Apr 6

Apr 6Mon
  1. OpenAI Alignment Research BlogAI score31

    OpenAI opens applications for Safety Fellowship on AI safety and alignment research

    AIOpenAI announced applications for its Safety Fellowship, a pilot program supporting external researchers, engineers, and practitioners in safety and alignment research on advanced AI systems. The program runs from September 14, 2026 through February 5, 2027, with a monthly stipend, compute support, API credits, and mentorship, and fellows are expected to produce a substantial output such as a paper, benchmark, or dataset. Applications close May 3, and successful applicants will be notified by July 25.

Mar 27

Mar 27Fri

Mar 26

Mar 26Thu

Mar 20

Mar 20Fri
  1. Aman SangerAI score55

    Cursor's Composer 2 is built on Kimi k2.5 base model with added training

    AIAman Sanger says Cursor's team evaluated many base models on perplexity-based evals and found Kimi k2.5 the strongest. Composer 2 was then built with continued pretraining and a 4x scale-up of high-compute RL, with Fireworks providing inference and RL samplers. The author admits Cursor should have named the Kimi base in its launch blog and says it will do so for the next model.

  2. Aman SangerAI score22

    Cursor's Composer 2 model praised, built on an open-source base

    AIAman Sanger of Cursor says Composer 2 is a really good model and he is excited for more people to try it. The quoted reply from Lee Robinson says Composer 2 started from an open-source base, with only about one-quarter of the final model's compute coming from that base. Cursor plans full pretraining in the future and says it is following the license through its inference partner terms.

Mar 10

Mar 10Tue

Mar 2

Mar 2Mon

Mar 1

Mar 1Sun

Feb 27

Feb 27Fri
  1. Mckay WrigleyAI score80

    Pentagon Secretary moves to label Anthropic a supply-chain risk

    AIMckay Wrigley reposted a statement from @SecWar accusing Anthropic of refusing the Department of War unrestricted access to its models for lawful purposes. The quoted statement directs the Department of War to designate Anthropic a Supply-Chain Risk to National Security, bars contractors from commercial activity with Anthropic, and allows Anthropic services for no more than six months. The author's own added text says only that he finds the situation horrifying and supports Anthropic.

    Why it matters: The quoted statement is a direct government action against a named AI lab, giving readers a primary-source view of a dispute over military access to AI models.

Feb 25

Feb 25Wed
  1. Quoc LeAI score53

    Google's Aletheia math agent solves 6 of 10 FirstProof problems

    AIQuoc Le announced that Aletheia, a math research agent, autonomously solved 6 of 10 FirstProof problems, the best result in the inaugural challenge. The post says this exceeds last year's IMO-gold achievement and points to a paper and thread for full details. The accompanying figure shows 10 unmodified problems, 6 candidate solutions per agent, and expert evaluation yielding 6 solved problems on a best-of-2 basis.

Feb 24

Feb 24Tue

Feb 23

Feb 23Mon