Skip to contentSkip to stories

Updated

#OpenAI

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 9

TodayOct 9Fri3 items
  1. MIT Technology Review · AIAI score62

    AI refusal is probabilistic and unreliable, and it raises censorship risks

    AIThe article argues that AI refusal, the main safety mechanism in modern models, is unreliable and hard to draw lines for. It cites jailbreaks, classifier stacks, and studies showing refusal skewed toward repressive governments. It warns that governments and companies could use refusal to censor speech, and that refusal behavior remains poorly understood.

  2. Bloomberg · TechnologyAI score48

    How AI Is Upending the World of Mathematics

    AIOpenAI announced last month that it had produced an AI-generated proof for the Navier-Stokes problem, a result the source says is hard even for experts to parse. The source also says LLMs now tackle math problems that have stumped humans for decades, while teachers struggle to keep up with AI-completed homework.

Oct 8

Oct 8Thu
  1. The Verge · AIAI score46

    Artificial, Guadagnino's Sam Altman satire, sticks closely to OpenAI's real events

    AILuca Guadagnino's film Artificial, a satirical biopic of OpenAI CEO Sam Altman, follows the factual events of Altman's rise and brief ouster, with Andrew Garfield playing Altman. The film is structured around real milestones, including Ilya Sutskever's meeting with Geoffrey Hinton and the 2017 Dota 2 win, and relies little on invented scenes.

  2. CNBC · TechnologyAI score22

    Cramer Says AI Stock Sell-Off Shows Value of Diversification

    AICNBC's Jim Cramer said Thursday's AI stock sell-off, which followed a Financial Times report that OpenAI's annualized revenue was about $20 billion below previously signaled levels, shows why investors need diversification. Oracle fell 5.5% and Broadcom dropped 4.35%, while Home Depot rallied as Treasury yields eased. Cramer argued that concentrated portfolios risk panic-selling and missing a recovery.

  3. Artificial AnalysisAI score42

    More output tokens don't guarantee higher scores in AI benchmarks

    AIArtificial Analysis reports that generating more output tokens does not necessarily yield a higher score. GPT-6 Astra (max) scored 8.6% using about 81k output tokens per task, while Grok 4.7 (xhigh) used roughly 180k yet scored lower. Three Claude models produced the most output tokens, about 202k to 562k per task, but scored between 2.8% and 6.4%.

    Image from @ArtificialAnlys's post
  4. Ethan MollickAI score42

    Community Rapidly Advances OpenAI-Linked Proofs, Tightening Bound to 2⁻¹⁵

    AIEthan Mollick notes that some OpenAI proofs have sparked rapid iterative advances from a wide community of collaborators amid debate over their implications for mathematics. A related post reports that a collaborative effort tightened the bound κ from 2⁻¹⁸² to 2⁻¹⁵, a roughly 500-thousand-fold improvement on the previous result.

  5. a16z NewsAI score45

    CFOs Are Becoming Builders as AI Reshapes Finance Operations

    AIAI-native tools are removing the data bottleneck that long constrained CFOs, shifting the role toward designing the operating systems that turn data into decisions. Finance teams are adopting AI-native software for ERP, forecasting, procurement, and audit, and "finance engineers" are building custom automations and agents. OpenAI's CFO Sarah Friar describes finance moving toward a zero-day close and continuously updated forecasts.

  6. TransformerAI score53

    Yoshua Bengio urges AI researchers to leave frontier labs for safety work

    AIYoshua Bengio, co-president of LawZero, asks researchers at frontier AI companies to reconsider whether they should keep working there, arguing that safety efforts are not slowing a dangerous race. He cites the recent UN Security Council briefing on AI incidents and says he left his earlier research path after ChatGPT made the risks feel immediate. He urges researchers to join AI Safety Institutes or mission-driven organizations such as LawZero.

  7. The Verge · AIAI score41

    Meta's Muse and OpenAI's Dots: can consumers trust AI agents with their lives?

    AIMeta's Muse and OpenAI's Dots are always-on AI agents with animated mascots, pitched to consumers and businesses for tasks like restaurant reservations and inbox triage. Muse is free, while Dots is not, and OpenAI also offers "specialist" Dots for marketing, legal work, and accounting. The discussion centers on privacy and security concerns about giving agents access to credit card details and email.

  8. MIT Technology Review · AIAI score44

    AI advances won't quickly make robots useful in everyday life, researchers say

    AIResearchers at robotics labs say that AI advances behind chatbots like ChatGPT and Claude will not quickly produce robots that are useful in everyday life. Many skeptics argue that using language- and image-based intelligence to master the physical world is far harder than it sounds, despite bold predictions from Elon Musk about Tesla's Optimus. Progress is real but incremental, as shown by Google DeepMind's Gemini Robotics controlling ALOHA 2 arms to pack a lunchbox.

  9. Air Street PressAI score60

    Nathan Benaich's 2026 State of AI Report covers agents, robotics, and AI control

    AINathan Benaich's 9th annual State of AI Report covers agents, robotics, AI for science, inference economics, and government control over frontier AI access. The report also records a 2025 prediction scorecard and lists nine predictions for the next 12 months. It cites an OpenAI cyber evaluation in which agents compromised Hugging Face's production infrastructure, and it says Anthropic and OpenAI's combined annualized revenue run rate reached $105B by late summer.

  10. meng shaoAI score39

    Claude Haiku 5.5 tops GPT-6 Luna on benchmarks, with 2x faster token output

    AIAnthropic's Claude Haiku 5.5, released alongside Claude Opus 5.5 and Claude Sonnet 5.5, is reported to lead GPT-6 Luna across benchmarks, with OpenRouter measuring roughly twice the token output speed. Anthropic says Haiku 5.5 is its cheapest, fastest, and most capable small model, costing about 75% less to run than Claude Haiku 4.5 on average. The post also notes some CodeX users are reportedly migrating to Claude Code.

Oct 7

Oct 7Wed
  1. Andrew CurranAI score52

    AI Labs Reportedly Test Internal Models Against Cryptographic Protocols

    AIScott Aaronson reports, based on his sources, that some AI companies have begun discreetly investigating whether their latest internal models can break important cryptographic protocols and primitives. He notes that cryptography is conspicuously absent from OpenAI's list of 376 papers, and the quoted post adds that the US government has censored academic quantum cryptanalysis results.

    Image from @AndrewCurran_'s post
  2. Ethan MollickAI score60

    Mathematicians react to hundreds of AI-generated proofs released by OpenAI

    AIEthan Mollick shares early first-hand accounts from mathematicians grappling with hundreds of AI proofs released by OpenAI. He highlights problems solved in ways no human has yet understood, raising questions about what it means to know something. The linked Scott Aaronson post quotes a researcher, Dana, describing the proofs as unclear and hard to read without AI help, with some possibly verified by a Lean certificate.

    Image from @emollick's post
  3. Ethan MollickAI score58

    Ethan Mollick Tries Intelligent UI in ChatGPT, Finds It Beats Text Walls

    AIEthan Mollick had early access to Intelligent UI and found it a welcome change from long blocks of text. He suggests interfaces will increasingly be built on demand for each user's problem. The quoted OpenAI post says GPT-6 and Intelligent UI are rolling out in ChatGPT for everyone, delivering fast, interactive answers with visual explanations and task tools.

  4. Marcus on AIAI score62

    Marcus Says OpenAI's Math Result Lacks Details Needed to Judge Its Generality

    AIGary Marcus argues that OpenAI's math announcement omits the procedure, the model architecture, and the failure rate, so its generalizability cannot be assessed. He says it could be a step toward AGI or a Lean-based verification trick in a verifiable domain, and the initial report cannot distinguish the two. The post includes a quoted Terence Tao post that shares a satirical press release about a fictional film-endings repository.

  5. Andrew CurranAI score42

    Andrew Curran says OpenAI's 722 math results omit cryptography breakthroughs

    AIAndrew Curran notes OpenAI's 722 published mathematical results show striking under-representation of cryptographic breakthroughs, and he says he has personally witnessed US government censorship of academic quantum cryptanalysis results. He calls backroom government interventionism his base case and says rumors suggest yesterday's OpenAI math release was only the first of three batches.

    Image from @AndrewCurran_'s post
  6. Exponential ViewAI score72

    OpenAI's 722 machine-generated math results may split mathematics into two layers

    AIOpenAI released 722 mathematical manuscripts in 372 families, produced by an unreleased frontier model, with the average result taking the equivalent of three hours of ChatGPT Pro thinking. The author notes many results are verified in Lean but not all, and suggests mathematics could divide into vast machine-verified work and a compressed human 'effective theory' that people can actually understand.

  7. The SequenceAI score37

    The Sequence Learning Loop: OpenAI DevDay and Gemini 4 Argon Show Workflow Competition

    AIThe newsletter argues that AI competition is shifting toward completed workflows, citing OpenAI's September 29 DevDay announcements on cost and infrastructure and Google's September 30 introduction of Gemini 4 Argon for longer, more demanding reasoning tasks. It says coding agents must inspect repositories, edit code, run tests, and deliver reviewable work, so cost, context, and supervision matter alongside model intelligence.

Oct 6

Oct 6Tue