Skip to contentSkip to stories

Updated

All AI news

Items with an AI score under 20 are hidden. Show low-relevance items

Jun 16

Jun 16Tue
  1. BAAIOfficialAI score38

    BAAI unveils WuJie physical-world AI architecture in 2026 report

    AIBAAI President Wang Zhongyuan announced a shift in AI from token prediction to physical state prediction in the institute's 2026 annual research report. The report unveiled the full-stack WuJie architecture spanning foundation models, autonomous agents, and hardware-software infrastructure, and noted that BAAI has open-sourced over 200 models with global downloads exceeding 1 billion.

    Image from @BAAIBeijing's post

Jun 15

Jun 15Mon
  1. Zed BlogOfficialAI score38

    Zed Guild Cohort 1 Ends with 148 Merged Pull Requests from 33 Contributors

    AIZed's 12-week Guild program, its first cohort run this spring, had 33 active contributors merge 148 pull requests into the open-source editor. The top contributor, feitreim, merged 23 PRs, including fixes for Vim mode screen flickering and terminal ANSI rendering, and won a trip to Rust Week in Utrecht. Zed plans to organize Cohort 2 work into tighter groups around specific parts of the codebase.

  2. BAAIOfficialAI score22

    Turing Award winners Diffie and Barto keynote BAAI Conference on AI security and RL

    AITuring Award winners Whitfield Diffie and Andrew Barto delivered keynotes at the BAAI Conference on AI security and reinforcement learning. Diffie argued that today's feedback-based approach only patches programs after they fail, and that formal methods offer a path to substantially more reliable intended behavior. Barto framed reinforcement learning around control, search, and associative memory, describing its core insight as caching search results rather than searching continuously.

    Image from @BAAIBeijing's post

Jun 12

Jun 12Fri
  1. Jeremy HowardXAI score72

    US export directive forces Anthropic to disable Fable 5 and Mythos 5 for customers

    AIThe US government issued an export control directive suspending access to Fable 5 and Mythos 5 for all foreign nationals, inside or outside the United States. Anthropic says the order forces it to disable both models for all customers, while other Claude models are unaffected. Anthropic calls the directive a misunderstanding and says it is working to restore access as soon as possible. The author disagrees with the decision and questions why Anthropic did not anticipate it, given its claim that only it can safely handle these models.

    Why it matters: The quoted Anthropic statement gives the directive's scope and the disruption to customers, which helps readers judge its practical effect on access to these models.

Jun 10

Jun 10Wed

Jun 5

Jun 5Fri
  1. BAAIOfficialAI score20

    BAAI's 8th Conference set for June 12–13 in Beijing

    AIThe 8th BAAI Conference will be held June 12–13 at the Zhongguancun Innovation Center in Beijing, with Turing Award laureates and China's large-model leaders attending. Core focuses include world models and agents, plus two new flagship sessions on AI-native education and the token economy. The event will feature 25 forums, over 200 speeches, and a first-ever on-site AI agent conference companion for real-time listening and summarization.

    Image from @BAAIBeijing's post

Jun 3

Jun 3Wed
  1. Cognition Blog (Devin, Windsurf)OfficialAI score60

    Cognition launches $10M AI Productivity Guarantee for enterprise Devin customers

    AICognition introduced the AI Productivity Guarantee, under which it will issue credits up to $10M if Devin delivers less engineering value than enterprise customers pay for. The company uses an AI estimator to measure hours of productive output, validated against engineers' own estimates of how long the same work would have taken by hand. Value is converted to dollars at a standard global rate and compared against each customer's consumption near the end of the annual contract.

    Why it matters: The post explains how Cognition estimates Devin's output in hours and backs the estimate with a $10M credit commitment, a concrete model for measuring AI vendor value.

Jun 2

Jun 2Tue

May 26

May 26Tue
  1. Cognition Blog (Devin, Windsurf)OfficialAI score46

    Cognition raises over $1B at $26B valuation as Devin usage surges

    AICognition raised over $1B at a $26B valuation, led by Lux Capital, General Catalyst, and 8VC. Enterprise Devin usage grew more than 10x since the start of this year, and run-rate revenue reached $492M. The company also said it launched SWE-1.6, which reaches up to 950 tok/s and has become the most used model in Devin Desktop.

May 22

May 22Fri
  1. ReflectionOfficialAI score38

    Reflection signs MOU with US Department of Energy for national labs

    AIReflection, an open-source AI lab, has signed a memorandum of understanding with the US Department of Energy to explore strategic collaborations under the Genesis Mission. Through the partnership, Reflection will provide open-weight models to the DOE's 17 National Laboratories, customizable on lab-specific scientific data and deployed into researcher workflows. The Genesis Mission focuses on advanced nuclear, fusion and grid modernization, quantum ecosystem growth, and national security AI.

May 21

May 21Thu
  1. Mark ChenXAI score92

    OpenAI model disproves Erdős's unit distance conjecture in planar geometry

    AIAn OpenAI model disproved Erdős's longstanding planar unit distance conjecture, which Paul Erdős posed in 1946, by discovering a new family of constructions that performs better than the square grids mathematicians had long assumed. Mark Chen says the proof draws on algebraic number theory and describes it as the first time AI has autonomously solved a prominent open problem central to a field of mathematics.

    Why it matters: The post names the specific open problem and the approach used, giving readers a concrete case of AI producing a research proof in mathematics.

May 19

May 19Tue

May 17

May 17Sun

May 11

May 11Mon

May 8

May 8Fri

May 6

May 6Wed
  1. OpenAI Alignment Research BlogOfficialAI score62

    OpenAI finds accidental chain-of-thought grading in several RL runs but no clear monitorability loss

    AIOpenAI reports that its automated system found accidental chain-of-thought grading in RL runs for several released models, including GPT-5.4 Thinking and GPT-5.4 mini. Its analysis found no clear reduction in CoT monitorability, though the company says subtler effects cannot be ruled out. OpenAI says it still avoids grading CoTs during RL and has fixed the affected reward pathways.

    Why it matters: The post shows how accidental chain-of-thought grading was detected and tested, giving a concrete method for checking monitorability risks in RL training.

  2. Eugene YanXAI score44

    Anthropic gains more compute via SpaceX, boosting Claude rate limits

    AIAnthropic announced a partnership with SpaceX that will give it more compute, according to the post. The related Claude post says Claude Code's 5-hour rate limits are doubled for Pro, Max, and Team plans, peak-hour reductions are removed for Pro and Max, and Opus API rate limits are substantially raised.

May 2

May 2Sat
  1. Nick TurleyXAI score27

    ChatGPT's new image feature sees usage up over 50% in weeks

    AIOpenAI's Nick Turley reports that usage of the new ChatGPT images feature rose more than 50% within a few weeks. Nearly 60% of daily users are newly logged-in users, and the feature is being used across home design, learning, work graphics, and creative projects.

May 1

May 1Fri
  1. ReflectionOfficialAI score22

    Reflection says shared Pentagon understanding sets AI deployment precedent

    AIReflection says its mission is to build open, safe, and accessible frontier intelligence, and that it has a responsibility to shape how AI is deployed. The company says a shared understanding with the Pentagon sets a precedent for how AI labs could work across the U.S. government, from supporting service members to scientists.

  2. ReflectionOfficialAI score38

    Reflection joins AI coalition on responsible U.S. government deployment

    AIReflection has joined a coalition including AWS, Microsoft, OpenAI, Google, and Nvidia on a framework governing how the U.S. government licenses and deploys AI. The agreement, which includes a non-binding memorandum of understanding with the DoW, commits to safety, red-teaming, and ongoing evaluation and explicitly prohibits unlawful mass surveillance and autonomous weapon use. Reflection says it will keep its commitment to open source while customizing its models for scientists in national labs.

Apr 30

Apr 30Thu
  1. Mark ChenXAI score62

    OpenAI's Mark Chen says GPT-5.5 performs like Mythos in UK AISI cyber range

    AIMark Chen says GPT-5.5 performs similarly to Mythos on UK AISI's cyber range, which tests long-horizon, agentic capability, and calls it one eval rather than a full picture. He adds that frontier model risks are real and that OpenAI aims to deploy AI people can actually use through mitigations. The attached chart shows completed steps per cumulative token spent for GPT-5.5, Mythos Preview, and several Claude and GPT models, from M1 reconnaissance up to M9 full network takeover.

    Why it matters: The post links a single cyber-range eval to OpenAI's own safety framing, so readers can weigh the result against the company's stated risk and deployment position.

Apr 29

Apr 29Wed
  1. Cognition Blog (Devin, Windsurf)OfficialAI score34

    Cognition opens Singapore headquarters for Asia-Pacific push with Devin

    AICognition has opened its Asia-Pacific headquarters in Singapore to expand its autonomous software engineering platform, Devin, across the region. The company says OCBC saw up to 30% improvement in code and test case generation, and its system integration test first-pass rate rose from below 50% to over 80% after deployment. Cognition is building its Singapore team across engineering, go-to-market, and partnerships, with Richard Spence leading APAC.

Apr 27

Apr 27Mon
  1. Cognition Blog (Devin, Windsurf)OfficialAI score38

    Mercedes-Benz Deploys Devin and Windsurf Across Global Engineering Teams

    AIMercedes-Benz is deploying Cognition's Devin and Windsurf across its global engineering teams, from the United States to Europe and Asia. In a four-week pilot, Devin analyzed over 200,000 lines of COBOL code and cut modernization time from an estimated eight months to eight days. The company is now rolling out the full suite, with Windsurf for development, Devin as an autonomous cloud agent, and Devin for Terminal for the most complex tasks.

  2. Andy JassyXAI score47

    OpenAI models coming to Amazon Bedrock in coming weeks

    AIAndy Jassy announced that OpenAI's models will become available directly to customers on Amazon Bedrock in the coming weeks. The availability will be alongside the upcoming Stateful Runtime Environment, giving builders more model choice. More details are expected at the AWS event in San Francisco tomorrow.

Apr 24

Apr 24Fri
  1. Andy JassyXAI score42

    Meta commits to tens of millions of AWS Graviton cores

    AIMeta has committed to tens of millions of AWS Graviton cores, according to Andy Jassy, Amazon's CEO. Jassy argues agentic AI is becoming as much a CPU story as a GPU story, since multi-step orchestration and real-time reasoning are CPU-intensive. He says the Graviton5 instances deliver up to 33% lower latency between cores.

    Video from @ajassy's post

Apr 23

Apr 23Thu

Apr 22

Apr 22Wed
  1. Anthropic EngineeringOfficialAI score78

    Anthropic traces Claude Code quality complaints to three product changes

    AIAnthropic says three changes to Claude Code, the Claude Agent SDK, and Claude Cowork caused recent quality complaints, and the API was not affected. The fixes were resolved by April 20 (v2.1.116), and the company is resetting usage limits for all subscribers as of April 23.

    Why it matters: The postmortem traces three separate changes to specific dates and versions, showing how a bug in context management can look like broad degradation to users.

Apr 21

Apr 21Tue
  1. Michael TruellXAI score62

    Cursor partners with SpaceX to scale up Composer, with an option to acquire

    AICursor's Michael Truell says the company is partnering with the SpaceX team to scale up Composer, calling it a meaningful step toward building the best place to code with AI. The quoted SpaceX post says Cursor gives SpaceX the right to acquire Cursor later this year for $60 billion, or pay $10 billion for the work together. It also cites SpaceX's Colossus training supercomputer, described as a million H100-equivalent system, as a source of training capacity.

    Why it matters: The post shows a partnership with conditional acquisition terms, which matters for judging how Cursor's coding products and model training may be developed.

Apr 18

Apr 18Sat
  1. Mark ChenXAI score42

    Mark Chen says science remains vital, names Bubeck and Helkky to lead

    AIOpenAI's Mark Chen rejected the claim that science is no longer a priority, saying AI should accelerate discovery in partnership with scientists. He said Sébastien Bubeck and Helkky are taking on this mandate, which the quoted post frames as a response to reports that OpenAI's VP of Science is leaving.

Apr 17

Apr 17Fri
  1. NVIDIA AI DeveloperOfficialAI score23

    NVIDIA's Cosmos Cookoff winners showcase Cosmos Reason 2 projects

    AINVIDIA highlighted the developers who won the Cosmos Cookoff and how they used Cosmos Reason 2 to build projects spanning disaster-response drones, explainable visual AI, and intelligent security systems. The post links to a YouTube showcase and a LinkedIn recap of the event.

    Image from @NVIDIAAIDev's post

Apr 13

Apr 13Mon
  1. Mira MuratiXAI score26

    Thinking Machines welcomes Workshop Labs founders Luke Drago and @LRudL_

    AIThinking Machines has welcomed Luke Drago and @LRudL_, who co-founded Workshop Labs to build AI that keeps humans relevant. Mira Murati says they will continue that mission at Thinking Machines, which builds powerful AI systems that think alongside humans and extend human agency. She ties the hire to the company's broader work, including Tinker, research grants, and advancing the frontier.

Apr 8

Apr 8Wed
  1. Cognition Blog (Devin, Windsurf)OfficialAI score31

    Cognition Expands to Japan, Appoints Takumi Masai to Lead Devin Launch

    AICognition is expanding into Japan, its first step into Asia, and has appointed Takumi Masai as Japan President and General Manager to lead a local team working with Japanese enterprises. DeNA has used Devin to more than double operational efficiency across multiple engineering functions, and Mizuho Securities has deployed it as one of the first large-scale financial institutions in Japan.

Apr 7

Apr 7Tue
  1. Dario AmodeiXAI score62

    Anthropic's Dario Amodei says new Mythos Preview model shows a large jump in cyber capabilities

    AIDario Amodei says the company has tracked growing cyber capabilities in AI models for years, which arise from their general coding proficiency. He states that the new model, Mythos Preview, represents a particularly large step up in those capabilities.

    Why it matters: The post links rising cyber capability to general coding skill, and names a notable jump in a new model, Mythos Preview.

  2. Dario AmodeiXAI score72

    Dario Amodei backs Project Glasswing to counter AI-driven cyber threats

    AIDario Amodei said many of the world's leading companies have joined Project Glasswing, an effort to address cyber threats posed by increasingly capable AI systems. The initiative was introduced by Anthropic and is powered by its newest frontier model, Claude Mythos Preview, which the quoted post says can find software vulnerabilities better than all but the most skilled humans.

    Why it matters: The post gives a concrete example of how a frontier AI lab is organizing industry partners around AI-driven software vulnerability discovery.