Skip to contentSkip to stories

Updated

#OpenAI

Showing low-relevance items too. Hide low-relevance items

Oct 8

Oct 8Thu
  1. MIT Technology Review · AIAI score44

    AI advances won't quickly make robots useful in everyday life, researchers say

    AIResearchers at robotics labs say that AI advances behind chatbots like ChatGPT and Claude will not quickly produce robots that are useful in everyday life. Many skeptics argue that using language- and image-based intelligence to master the physical world is far harder than it sounds, despite bold predictions from Elon Musk about Tesla's Optimus. Progress is real but incremental, as shown by Google DeepMind's Gemini Robotics controlling ALOHA 2 arms to pack a lunchbox.

  2. Hacker News · AI (150+ points)AI score40

    OpenAI Withdraws Three Math Papers Over Sign Error, Revises 14 Others

    AIOpenAI withdrew three math manuscripts, including "Algebraicity of Weil classes on split abelian eightfolds," after a sign error invalidated a stabilization-trace cancellation argument used by two dependent papers. The withdrawn papers now carry notices linking to archived manuscripts, and 14 other manuscripts were revised with proof repairs and corrected statements.

  3. Air Street PressAI score60

    Nathan Benaich's 2026 State of AI Report covers agents, robotics, and AI control

    AINathan Benaich's 9th annual State of AI Report covers agents, robotics, AI for science, inference economics, and government control over frontier AI access. The report also records a 2025 prediction scorecard and lists nine predictions for the next 12 months. It cites an OpenAI cyber evaluation in which agents compromised Hugging Face's production infrastructure, and it says Anthropic and OpenAI's combined annualized revenue run rate reached $105B by late summer.

  4. meng shaoAI score39

    Claude Haiku 5.5 tops GPT-6 Luna on benchmarks, with 2x faster token output

    AIAnthropic's Claude Haiku 5.5, released alongside Claude Opus 5.5 and Claude Sonnet 5.5, is reported to lead GPT-6 Luna across benchmarks, with OpenRouter measuring roughly twice the token output speed. Anthropic says Haiku 5.5 is its cheapest, fastest, and most capable small model, costing about 75% less to run than Claude Haiku 4.5 on average. The post also notes some CodeX users are reportedly migrating to Claude Code.

  5. Artificial Analysis ArticlesAI score62

    GPT-6 Sol Daybreak Blue leads the Artificial Analysis Cyber Index

    AIArtificial Analysis is adding trusted-access models to its Cyber Index, starting with GPT-6 Sol (Daybreak Blue, max), which is available only through OpenAI's Daybreak program. The model hits no safety blocks across the Index and scores 32 points higher overall than the publicly available GPT-6 Sol (max), with its largest gains on CyberGym-E2E.

    Why it matters: The source shows how safety refusals shape cyber benchmark scores, with the trusted-access model's gains concentrated on CyberGym-E2E, useful for comparing guarded and unguarded models.

  6. Artificial Analysis ArticlesAI score50

    Harvey LAB-AA v1.1 adds hallucination checks to legal AI benchmark

    AIHarvey LAB-AA v1.1 adds hallucination checks that audit every model deliverable against task source documents, with material hallucinations zeroing a task's score. GPT-6 Astra averaged 0.03 material hallucinations per task across 120 tasks, while Gemini 3.8 Flash averaged 13.96. Harvey uses GPT-6 Sol (high) as the hallucination checker, separate from its three-judge rubric panel.

Oct 7

Oct 7Wed
  1. KhazixAI score88

    OpenAI Releases 722 Unpublished AI-Generated Math Manuscripts on GitHub

    AIOpenAI published 722 math manuscripts covering 372 result groups in a new GitHub repository, openai/math, all produced by an unreleased internal model. The author describes the results as including a near-Riemann hypothesis claim pushed to 0.875, and notes that 25 Fields Medal winners criticized the company's approach to AI math research.

    Why it matters: The piece traces how AI math results moved from benchmarks to open problems, offering context on verification and the mathematicians' pushback.

  2. Jensen HuangAI score40

    Awesome day, @satyanadella!

    AIWindows sparked a platform shift that created a new industry for NVIDIA. Then we invented programmable shading GPUs for DirectX, which led to CUDA. Then we partnered to bring GPU supercomputers to Azure, which helped OpenAI train GPT. That collaboration inspired us to reinvent Windows for the age of personal agents. 4 years. Thousands of engineering years between us. So proud of what we built together.

  3. Andrew CurranAI score52

    AI Labs Reportedly Test Internal Models Against Cryptographic Protocols

    AIScott Aaronson reports, based on his sources, that some AI companies have begun discreetly investigating whether their latest internal models can break important cryptographic protocols and primitives. He notes that cryptography is conspicuously absent from OpenAI's list of 376 papers, and the quoted post adds that the US government has censored academic quantum cryptanalysis results.

    Image from @AndrewCurran_'s post
  4. meng shaoAI score75

    Microsoft positions Windows as the home for hybrid AI agents across four layers

    AIMicrosoft has repositioned Windows as the home for hybrid intelligence, where AI agents can run locally or in the cloud. The announcement covers four layers: MXC reaching general availability for agent isolation, local models such as MAI Code 1.1 Flash, Copilot on Copilot+ PCs gaining local context and actions in coming months, and new hardware including RTX Spark PCs and DGX Station for Windows.

    Image from @shao__meng's post
  5. meng shaoAI score88

    OpenAI rolls out GPT-6 with Intelligent UI to over 1.2 billion weekly ChatGPT users

    AIOpenAI is rolling out GPT-6 to ChatGPT's over 1.2 billion weekly users, adding Intelligent UI, which lets replies include charts, buttons, forms, and interactive tools. The post's image cites tiered access, with Free/Go and Plus/Pro/Business/Enterprise sharing the Sol and Luna model splits, and says the feature is progressively rendered as the model generates it.

    Image from @shao__meng's post
  6. Epoch AIAI score67

    Epoch tests six AI models on real Epoch work and finds they cannot yet fully automate it

    AIEpoch gave six models 11 real work tasks from its own operations, including graphic design, data insights, and research design, and graded outputs against employee standards. Fable 5.1 and GPT-6 Astra led on average task performance, reliably handling well-defined work such as coding and computational analysis. The report finds that all models still fail on open-ended judgment, including matching Epoch's standards, designing informative experiments, and generating diverse ideas, so the authors conclude AI cannot yet replace workers at Epoch.

    Why it matters: The report separates well-defined task reliability from open-ended judgment failures, which benchmark scores on easily verifiable tasks would miss.

  7. Ethan MollickAI score60

    Mathematicians react to hundreds of AI-generated proofs released by OpenAI

    AIEthan Mollick shares early first-hand accounts from mathematicians grappling with hundreds of AI proofs released by OpenAI. He highlights problems solved in ways no human has yet understood, raising questions about what it means to know something. The linked Scott Aaronson post quotes a researcher, Dana, describing the proofs as unclear and hard to read without AI help, with some possibly verified by a Lean certificate.

    Image from @emollick's post
  8. TypeSafe AIAI score25

    Jev-killer OpenAI Decisions API benchmarked against Jev for HiringCafe

    AIThe main post is a short reply saying reports of a company's death have been greatly exaggerated, with no details about products or figures. The background post from @h_nilforoshan reports that OpenAI's Decisions API, billed as a "Jev-killer," was benchmarked against Jev for HiringCafe, which serves 2.5 million users. On the task of scoring job-description relevance from 1 to 10, the author reports OpenAI costing 2x more and performing 5-10% worse.

  9. TechRadar · AIAI score42

    Trump creates Super Intelligence Force and renames AI to "SI" in federal communications

    AIPresident Trump announced a White House-led "Super Intelligence Force" that will spend 120 days examining AI risks and federal responses, and signed an executive order directing agencies to use "Super Intelligence" and "SI" instead of "Artificial Intelligence" and "AI." The order asks officials to develop a possible new federal definition within 60 days, but the source says the change is linguistic rather than architectural. Critics quoted in the article argue that renaming does not change the technology itself.

  10. Ethan MollickAI score58

    Ethan Mollick Tries Intelligent UI in ChatGPT, Finds It Beats Text Walls

    AIEthan Mollick had early access to Intelligent UI and found it a welcome change from long blocks of text. He suggests interfaces will increasingly be built on demand for each user's problem. The quoted OpenAI post says GPT-6 and Intelligent UI are rolling out in ChatGPT for everyone, delivering fast, interactive answers with visual explanations and task tools.