Skip to content

#OpenAI

Oct 8

TodayOct 8Thu117 items
  1. Lewis TunstallAI score62

    Lewis Tunstall Shares a Physics Paper Proof Developed with OpenAI's Astra Model

    Lewis Tunstall quotes Kyle Cranmer's post about a paper by Nate Gunnarsson on a non-perturbative approach to chiral fermions in the Standard Model, extending Lüscher's abelian result. The paper's acknowledgments state that OpenAI's GPT-6 Astra model was essential, proposing refinement strategies, writing rewrites of the proof, and carrying out Lean verification.

  2. OpenAI · YouTubeAI score67

    OpenAI launches GPT-6 Intelligent UI for interactive ChatGPT answers

    OpenAI's GPT-6 in ChatGPT adds Intelligent UI, which lets ChatGPT answer with interactive interfaces and build quick tools for a task. The feature is rolling out to Plus, Pro, Business, and Enterprise first, with Free and Go tiers following, and Enterprise access depends on workplace admin settings. The update covers only the Chat experience, and the models powering Work and Codex are not changing.

    This story has a top pick“OpenAI rolls out GPT-6 and Intelligent UI to all ChatGPT users”

  3. The Guardian · AIAI score42

    AI Boom Drives San Francisco Rents Up 25% as Evictions Rise 44%

    The AI industry's boom is pushing San Francisco rents sharply higher, with the average one-bedroom now costing $4,400, up more than 25% from last year and nearly three times the national average. Eviction notices citywide are up 44% compared with last year, and advocates say landlords are using Ellis Act evictions, renovictions and "self-eviction" tactics to exploit the market. Mayor Daniel Lurie declared a "rent emergency" in September, proposing eviction legal aid and caps on rent increases for newly vacant rent-controlled units.

  4. Philipp SchmidAI score46

    SynthID Detector (http://synthid.com) is now publicly available, supporting the Nano Banana 2.1, @OpenAI, @nvidia, Kakao, and soon @Apple. 🌐 Verify if an image, video, or audio file was generated: 1️⃣ Upload or paste your file at https://synthid.com 2️⃣ Scans for watermarks from Google or our partners (files are deleted right after)

    SynthID Detector (http://synthid.com) is now publicly available, supporting the Nano Banana 2.1, @OpenAI, @nvidia, Kakao, and soon @Apple. 🌐 Verify if an image, video, or audio file was generated: 1️⃣ Upload or paste your file at https://synthid.com 2️⃣ Scans for watermarks from Google or our partners (files are deleted right after)

  5. IEEE Spectrum · AIAI score46

    Nuclear Plants Adopt AI Tools, Led by Atomic Canyon's NIVA Assistant

    Atomic Canyon's Nuclear Industry Virtual Assistant (NIVA), developed with nuclear-industry groups, is now available to the entire U.S. fleet of 94 reactors after pilot testing at Constellation Energy plants. Nuclearn says its products have reached more than 65 U.S. partners, and the article says the industry is turning to AI to help manage regulatory paperwork and a shrinking, aging workforce.

  6. a16z NewsAI score45

    CFOs Are Becoming Builders as AI Reshapes Finance Operations

    AI-native tools are removing the data bottleneck that long constrained CFOs, shifting the role toward designing the operating systems that turn data into decisions. Finance teams are adopting AI-native software for ERP, forecasting, procurement, and audit, and "finance engineers" are building custom automations and agents. OpenAI's CFO Sarah Friar describes finance moving toward a zero-day close and continuously updated forecasts.

  7. TransformerAI score67

    Bengio urges safety-minded AI researchers to leave frontier labs

    Yoshua Bengio, co-president of LawZero, writes to researchers urging those who prioritize safety to leave frontier AI companies for safety institutes or mission-driven organizations. He argues that safety efforts at the labs are not sufficiently slowing a dangerous race toward recursive self-improvement, and cites LawZero's recent C$200 million-plus funding from Canada and Germany as an alternative path.

  8. The Verge · AIAI score41

    Meta's Muse and OpenAI's Dots: can consumers trust AI agents with their lives?

    Meta's Muse and OpenAI's Dots are always-on AI agents with animated mascots, pitched to consumers and businesses for tasks like restaurant reservations and inbox triage. Muse is free, while Dots is not, and OpenAI also offers "specialist" Dots for marketing, legal work, and accounting. The discussion centers on privacy and security concerns about giving agents access to credit card details and email.

  9. QbitAI (量子位)AI score44

    PaperBenchX Shows Top Model Reproduces Only 13.98% of 93 Scientific Papers End-to-End

    UniPat AI's PaperBenchX benchmark found the strongest model, GPT-6 Astra, fully reproduced only 13.98% of 93 real research-paper tasks across 12 scientific fields. Reproduction was judged by regenerating outputs in an isolated environment, with 3,168 expert-verified scoring items. UniPat has open-sourced 12 test tasks and kept 81 tasks closed to preserve long-term evaluation validity.

  10. The Guardian · AIAI score42

    Co-op places legal services staff under AI monitoring of customer phone calls

    The Co-op's legal services arm is using AI models to record and score staff phone calls with customers seeking probate and will advice, with a model from OpenAI assessing more than 50 aspects of each call. Managers use pass and fail scores to analyse employee performance, and the Co-op says the system is a support tool, not a decision-maker. Trade unions and a whistleblower have criticised the monitoring as oppressive.

  11. Gizmodo · AIAI score45

    OpenAI Reportedly Regains Nearly Half of AI Compute Market Share From Anthropic in 2026

    According to a Wall Street Journal report citing OpenRouter data, OpenAI's share of AI compute routed through the platform rose from under 25% at the start of 2026 to nearly 50% last month. The data comes largely from AI-native startups, with some legacy tech companies also included. The report comes as OpenAI reportedly shifted focus from Sora and an erotica generator toward productivity and business customers.

  12. MIT Technology Review · AIAI score44

    AI advances won't quickly make robots useful in everyday life, researchers say

    Researchers at robotics labs say that AI advances behind chatbots like ChatGPT and Claude will not quickly produce robots that are useful in everyday life. Many skeptics argue that using language- and image-based intelligence to master the physical world is far harder than it sounds, despite bold predictions from Elon Musk about Tesla's Optimus. Progress is real but incremental, as shown by Google DeepMind's Gemini Robotics controlling ALOHA 2 arms to pack a lunchbox.

  13. Gergely OroszAI score48

    You can either hold crypto and keep being stressed out if a math breakthrough would drain your wallet; or someone stealing your keys would drain your wallet; or someone kidnapping you and forcing you to hand over your keys would drain your wallet Or you can just not hold crypto

    You can either hold crypto and keep being stressed out if a math breakthrough would drain your wallet; or someone stealing your keys would drain your wallet; or someone kidnapping you and forcing you to hand over your keys would drain your wallet Or you can just not hold crypto

  14. Hacker News · AI (150+ points)AI score40

    OpenAI Withdraws Three Math Papers Over Sign Error in Weil Classes Proof

    OpenAI withdrew three math manuscripts, including "Algebraicity of Weil classes on split abelian eightfolds," after a sign error invalidated a stabilization-trace cancellation argument. The withdrawal also affects two papers that depended on that construction, and the withdrawn papers now carry notices linking to archived manuscripts. The update also revised 14 other manuscripts with proof repairs and added six formalizations.

  15. Air Street PressAI score60

    Nathan Benaich's 2026 State of AI Report covers agents, robotics, and AI control

    Nathan Benaich's 9th annual State of AI Report covers agents, robotics, AI for science, inference economics, and government control over frontier AI access. The report also records a 2025 prediction scorecard and lists nine predictions for the next 12 months. It cites an OpenAI cyber evaluation in which agents compromised Hugging Face's production infrastructure, and it says Anthropic and OpenAI's combined annualized revenue run rate reached $105B by late summer.

  16. Hacker News · AI (150+ points)AI score54

    OpenAI withdraws three mathematical results

    OpenAI has withdrawn three mathematical results, according to a Hacker News post linking to a history file in its openai/math GitHub repository. The feed supplied only the link and discussion counts, so the specific results, reasons for withdrawal, and any corrected conclusions are not described here.

  17. DeedyAI score24

    “Cryptography is a subfield that’s extremely conspicuous by its absence from OpenAI’s list of 376 papers! But my sources tell me that the AI cos have now started, gingerly and discreetly, investigating whether their latest internal models can break important cryptographic protocols and primitives.” - Scott Aaronson, CS chair at UT Austin

    “Cryptography is a subfield that’s extremely conspicuous by its absence from OpenAI’s list of 376 papers! But my sources tell me that the AI cos have now started, gingerly and discreetly, investigating whether their latest internal models can break important cryptographic protocols and primitives.” - Scott Aaronson, CS chair at UT Austin

  18. meng shaoAI score39

    Claude Haiku 5.5 tops GPT-6 Luna on benchmarks, with 2x faster token output

    Anthropic's Claude Haiku 5.5, released alongside Claude Opus 5.5 and Claude Sonnet 5.5, is reported to lead GPT-6 Luna across benchmarks, with OpenRouter measuring roughly twice the token output speed. Anthropic says Haiku 5.5 is its cheapest, fastest, and most capable small model, costing about 75% less to run than Claude Haiku 4.5 on average. The post also notes some CodeX users are reportedly migrating to Claude Code.

  19. Artificial Analysis ArticlesAI score62

    GPT-6 Sol Daybreak Blue leads the Artificial Analysis Cyber Index

    Artificial Analysis is adding trusted-access models to its Cyber Index, starting with GPT-6 Sol (Daybreak Blue, max), which is available only through OpenAI's Daybreak program. The model hits no safety blocks across the Index and scores 32 points higher overall than the publicly available GPT-6 Sol (max), with its largest gains on CyberGym-E2E.

    AIWhy it matters: The source shows how safety refusals shape cyber benchmark scores, with the trusted-access model's gains concentrated on CyberGym-E2E, useful for comparing guarded and unguarded models.

  20. Artificial Analysis ArticlesAI score50

    Harvey LAB-AA v1.1 adds hallucination checks to legal AI benchmark

    Harvey LAB-AA v1.1 adds hallucination checks that audit every model deliverable against task source documents, with material hallucinations zeroing a task's score. GPT-6 Astra averaged 0.03 material hallucinations per task across 120 tasks, while Gemini 3.8 Flash averaged 13.96. Harvey uses GPT-6 Sol (high) as the hallucination checker, separate from its three-judge rubric panel.

  21. LangChain BlogAI score67

    LangChain's Restock agent shows how to build a payment-capable AI agent

    LangChain built Restock, a sample office-supply agent that runs in Slack on Managed Deep Agents and pays through Stripe's Link wallet. The agent searches products, builds a cart, and pays over the Machine Payments Protocol, with the user approving the purchase in Slack and the payment in Link. The post uses a pens order at $22.18 to show the flow from request to confirmed order.

    AIWhy it matters: The post walks through how an agent handles search, budget limits, Slack review, and Link approval, showing where each control sits outside the model.

Oct 7

Oct 7Wed
  1. Khazix (数字生命卡兹克)AI score88

    OpenAI Releases 722 Unpublished AI-Generated Math Manuscripts on GitHub

    OpenAI published 722 math manuscripts covering 372 result groups in a new GitHub repository, openai/math, all produced by an unreleased internal model. The author describes the results as including a near-Riemann hypothesis claim pushed to 0.875, and notes that 25 Fields Medal winners criticized the company's approach to AI math research.

    AIWhy it matters: The piece traces how AI math results moved from benchmarks to open problems, offering context on verification and the mathematicians' pushback.

  2. Andrew CurranAI score52

    AI Labs Reportedly Test Internal Models Against Cryptographic Protocols

    Scott Aaronson reports, based on his sources, that some AI companies have begun discreetly investigating whether their latest internal models can break important cryptographic protocols and primitives. He notes that cryptography is conspicuously absent from OpenAI's list of 376 papers, and the quoted post adds that the US government has censored academic quantum cryptanalysis results.