Skip to content

#OpenAI

Oct 8

TodayOct 8Thu119 items
  1. CNBC · TechnologyAI score52

    AI stocks fall after OpenAI revenue figure comes in below earlier reports

    OpenAI told investors it reached roughly $50 billion in annualized revenue at the end of September, below the $68 billion figure widely reported last month. Nvidia, Oracle, and CoreWeave shares fell on Thursday, with a person familiar saying the $68 billion figure included partner gross revenue. OpenAI is also weighing a possible funding round of around $30 billion and has not finalized a term sheet.

  2. Epoch AIAI score31

    This isn’t conclusive evidence of a new architecture, but it suggests something has changed in how GPT-6.1 Sol handles long contexts. This is a follow-up to our earlier report on latency scaling in frontier models: https://epoch.ai/publications/long-context-latency-scaling-gpt-vs-claude#appendix-h-gpt-6-sol-and-gpt-61-sol

    This isn’t conclusive evidence of a new architecture, but it suggests something has changed in how GPT-6.1 Sol handles long contexts. This is a follow-up to our earlier report on latency scaling in frontier models: https://epoch.ai/publications/long-context-latency-scaling-gpt-vs-claude#appendix-h-gpt-6-sol-and-gpt-61-sol

  3. Wired · AIAI score36

    Elon Musk's America PAC Spends Millions on 2026 Midterm Senate Races

    Elon Musk's America PAC is spending millions on the most consequential Senate races of the 2026 midterms, with most of its money opposing Democrats rather than supporting Republicans. WIRED reports that for every dollar the PAC spent backing a Republican candidate, more than two went to oppose a Democrat. In Texas, the group has spent about $9 million attacking Democrat James Talarico versus roughly $1 million promoting Ken Paxton.

  4. NVIDIA NewsroomAI score46

    Developers Use Frontier AI Agents to Build NVIDIA Omniverse Simulations

    NVIDIA developers are pairing frontier AI models, including GPT-6 Astra and Claude Fable 5, with Omniverse libraries to turn simulation ideas into working applications. Examples include a humanoid warehouse simulator, an autonomous-driving testing workflow, and sensor-matching digital twins. The projects are guided through natural-language instructions and reviewed by developers.

  5. NVIDIA BlogAI score49

    How Developers Use Frontier AI Agents to Build Omniverse Simulations

    Developers are pairing frontier AI models with NVIDIA Omniverse libraries to turn simulation ideas into working applications, from humanoid warehouse simulators to autonomous-driving test environments. In the examples, developers direct AI agents through natural-language instructions and review results, while Omniverse provides GPU-accelerated physics, rendering and sensor simulation. One experiment reported a simulated Unitree G1 humanoid clearing a hurdle in 64 of 100 trials.

  6. TechCrunch · AIAI score62

    Fired OpenAI safety researchers dispute misconduct claims and warn of chilling effect

    Three OpenAI safety researchers, Jasmine Wang, Tomek Korbak, and Mikita Balesni, were fired after OpenAI said they mishandled sensitive information by sharing it with an outside AI safety organization. In an open letter, they deny the claims, argue the dismissals will deter employees from raising safety concerns, and call on OpenAI to keep its public commitments on third-party safety auditing. OpenAI says the firings followed an investigation into a pattern of misconduct and denies they were retaliation for safety concerns.

  7. Artificial AnalysisAI score38

    GPT-6 Sol (Daybreak Blue, max) has been added to the Artificial Analysis Cyber Index as a trusted-access model, achieving the #1 spot on the Index. Compared to the publicly available GPT-6 Sol, the Daybreak Blue model has the largest gains on CyberGym-E2E, which is the benchmark where we observe the most safety refusals

    GPT-6 Sol (Daybreak Blue, max) has been added to the Artificial Analysis Cyber Index as a trusted-access model, achieving the #1 spot on the Index. Compared to the publicly available GPT-6 Sol, the Daybreak Blue model has the largest gains on CyberGym-E2E, which is the benchmark where we observe the most safety refusals

  8. Artificial AnalysisAI score46

    GPT-6 Sol (Daybreak Blue, max) takes the top position on the Artificial Analysis Cyber Index at much lower cost than other leading models. At a Cost per Task of $1.77, it is significantly more cost-effective than other leading models, including Grok 4.7 (xhigh) which costs $11.67 per task

    GPT-6 Sol (Daybreak Blue, max) takes the top position on the Artificial Analysis Cyber Index at much lower cost than other leading models. At a Cost per Task of $1.77, it is significantly more cost-effective than other leading models, including Grok 4.7 (xhigh) which costs $11.67 per task

  9. Artificial AnalysisAI score62

    GPT-6 Sol Daybreak Blue leads the Artificial Analysis Cyber Index

    Artificial Analysis added trusted-access models to its Cyber Index, and GPT-6 Sol (Daybreak Blue, max) now ranks first. The model is available only through OpenAI's Daybreak program and records no safety blocks across the Index. Its overall score is 32 points higher than the publicly available GPT-6 Sol (max), at a cost of $1.77 per task versus $11.67 for Grok 4.7 (xhigh).

  10. Sherwin WuAI score62

    Harvey LAB-AA v1.1 adds hallucination gate, reshaping legal benchmark rankings

    Artificial Analysis and Harvey released LAB-AA v1.1, which credits a legal task only when deliverables pass every rubric criterion with no material hallucinations. Grok 4.7 (xhigh) leads at 9.4%, ahead of Muse Spark 1.3 (max) at 8.9% and GPT-6 Astra (max) at 8.6%, while over 60% of otherwise passing results contained a material hallucination. The sharper reordering appears in the hallucination counts, where GPT-6 Astra averages 0.03 material hallucinations per task against 13.96 for Gemini 3.8 Flash (high).

  11. Andrew CurranAI score62

    Three fired OpenAI safety researchers publish open letter to leadership

    Three OpenAI safety and alignment employees, Tomek Korbak, Jasmine Wang, and Mikita Balesni, were fired last week and have published an open letter to OpenAI's safety and governance committees. The letter argues that OpenAI cannot make AI safe on its own, calls for open debate, third-party collaboration, and clear internal procedures, and says the firing and its handling bear directly on safety oversight.

  12. Codex · GitHub ReleasesAI score36

    Codex 0.162.0 adds managed worktree tools and clickable URLs in the TUI

    OpenAI's Codex 0.162.0 release adds tools for creating and listing managed Git worktrees from trusted local projects when the worktrees feature is enabled. The update also lets users pin tasks in the agent Command Center, copy transcript blocks with /copy, and make URLs clickable in approval headers, questions, and warnings, along with several Linux and Windows sandbox fixes.

  13. Testing CatalogAI score50

    OPENAI 🔥: GPT-6.1 Sol Ultrafast is rolling out on ChatGPT Work, Codex, and the API. > GPT-6.1 Sol in Ultrafast mode is $12 per million input tokens and $60 per million output tokens. > Ultrafast mode delivers 8x faster speeds than Sol Standard. Ultrafast testing time 👀 https://x.com/OpenAIDevs/status/2108262812489531498/video/1

    OPENAI 🔥: GPT-6.1 Sol Ultrafast is rolling out on ChatGPT Work, Codex, and the API. > GPT-6.1 Sol in Ultrafast mode is $12 per million input tokens and $60 per million output tokens. > Ultrafast mode delivers 8x faster speeds than Sol Standard. Ultrafast testing time 👀 https://x.com/OpenAIDevs/status/2108262812489531498/video/1

  14. Artificial AnalysisAI score42

    Generating more output tokens doesn’t necessarily translate to a higher score. GPT-6 Astra (max) scores 8.6% on ~81k output tokens per task, under half the ~180k of Grok 4.7 (xhigh). Three Claude models generated the most output tokens (~202k to ~562k per task) and score 2.8% to 6.4%.

    Generating more output tokens doesn’t necessarily translate to a higher score. GPT-6 Astra (max) scores 8.6% on ~81k output tokens per task, under half the ~180k of Grok 4.7 (xhigh). Three Claude models generated the most output tokens (~202k to ~562k per task) and score 2.8% to 6.4%.

  15. Artificial AnalysisAI score28

    Among models with a Hallucination-Gated All-Pass Rate above 0%, four set the Pareto frontier for score vs. Cost per Task: GPT-6 Luna (max), GPT-6.1 Sol (max), Muse Spark 1.3 (max) and Grok 4.7 (xhigh). Grok 4.7 (xhigh) leads at ~$9.50 per task and Muse Spark 1.3 (max) comes second at ~$4.20, while the three Claude models cost ~$18 to ~$22 per task. GPT-6 Luna (max) is the cheapest at ~$0.22 per task, scoring 3.3%.

    Among models with a Hallucination-Gated All-Pass Rate above 0%, four set the Pareto frontier for score vs. Cost per Task: GPT-6 Luna (max), GPT-6.1 Sol (max), Muse Spark 1.3 (max) and Grok 4.7 (xhigh). Grok 4.7 (xhigh) leads at ~$9.50 per task and Muse Spark 1.3 (max) comes second at ~$4.20, while the three Claude models cost ~$18 to ~$22 per task. GPT-6 Luna (max) is the cheapest at ~$0.22 per task, scoring 3.3%.

  16. Artificial AnalysisAI score34

    Artificial Analysis compares six hallucination checkers on 20 shared tasks

    Artificial Analysis compared six hallucination checkers on the same deliverables from 20 tasks across eight models. GPT-6 Sol and GPT-6 Luna generally flagged the most material hallucinations, while Claude Sonnet 5.5 and Gemini 3.8 Flash flagged far fewer, with Claude Opus 5.5 falling between Grok 4.7 and Sonnet. The counts reflect checker behavior rather than establishing accuracy or ruling out self-preference.

  17. OpenAI DevelopersAI score47

    In Codex and ChatGPT Work, access is available on Pro 500, eligible usage-based Enterprise, and credit-based Edu plans. Enterprise admins must enable access. Ultrafast for GPT-6.1 Sol is available in all supported regions, including support for data residency in the US and EU. We’ve also added support for EU data residency for GPT-6.1 Sol Fast and GPT-6 Luna Fast. https://developers.openai.com/api/docs/guides/ultrafast-mode

    In Codex and ChatGPT Work, access is available on Pro 500, eligible usage-based Enterprise, and credit-based Edu plans. Enterprise admins must enable access. Ultrafast for GPT-6.1 Sol is available in all supported regions, including support for data residency in the US and EU. We’ve also added support for EU data residency for GPT-6.1 Sol Fast and GPT-6 Luna Fast. https://developers.openai.com/api/docs/guides/ultrafast-mode

  18. OpenAI DevelopersAI score38

    API pricing for GPT-6.1 Sol in Ultrafast mode is $12 per million input tokens and $60 per million output tokens. Built for work where speed and intelligence make a difference: debugging an outage, agents navigating apps, and live experiences where every second counts.

    API pricing for GPT-6.1 Sol in Ultrafast mode is $12 per million input tokens and $60 per million output tokens. Built for work where speed and intelligence make a difference: debugging an outage, agents navigating apps, and live experiences where every second counts.

  19. TechCrunch · AIAI score58

    OpenAI's annualized revenue reportedly about $20 billion below earlier estimates

    OpenAI has reportedly told investors its annualized revenue is approaching $50 billion, about $20 billion below a previously reported $70 billion figure. The Financial Times reports the earlier number came from investor attempts to compare OpenAI with Anthropic, which counts cloud partners' sales differently. OpenAI's IPO has reportedly been pushed to early 2027.

  20. The DecoderAI score75

    Mathematicians call for OpenAI boycott after AI-generated proofs flood their field

    A group of mathematicians led by Terence Tao has called for a boycott of OpenAI after the company released more than 700 AI-generated proof files at once. Tao and other Fields Medalists argue that mass-produced solutions undermine the discipline's focus on conceptual understanding, while Scott Aaronson contrasts this batch release with Anthropic's collaborative approach. The article reports that the internal model tested about 8,000 problems with roughly a five percent success rate.

  21. TechCrunch · AIAI score65

    OpenAI's math solutions fall short of the field's standards, mathematicians say

    OpenAI released hundreds of claimed solutions to hard math problems but did not fully meet guidelines from the Advisory Group on Mathematics and Artificial Intelligence. Only 10 of 719 manuscripts included chain-of-thought releases, and just 42% of proofs were formalized. A Cambridge and King's College paper found discrepancies between a natural language proof and its Lean code for a Navier-Stokes-derived problem.

  22. The Verge · AIAI score44

    USA Today Co. sues OpenAI, seeking over $250 million for copyright infringement

    USA Today Co. and several local newspapers it owns sued OpenAI, alleging the company copied "hundreds of thousands" of articles to train its AI models without permission. The publisher seeks damages of more than $250 million, arguing the unauthorized use has caused real and continuing harm. OpenAI did not immediately respond to The Verge's request for comment.

  23. OpenAI · YouTubeAI score70

    OpenAI adds Intelligent UI to GPT-6 in ChatGPT Chat tab

    OpenAI introduced Intelligent UI for GPT-6 in ChatGPT, which lets the chatbot answer with fully interactive interfaces and quickly build tools for a task. The feature rolled out globally to Plus, Pro, Business, and Enterprise tiers in the Chat tab and expands to Free and Go tiers, with Enterprise access depending on workplace admin settings.

    This story has a top pick“OpenAI rolls out GPT-6 and Intelligent UI to all ChatGPT users”

  24. Artificial IgnoranceAI score52

    Charlie Guo maps the core primitives that make AI agents work over time

    The author argues that agent systems are converging on shared primitives grouped into doing the work, continuing the work, and delegating the work. These include instructions and skills, tools and connectors, sandboxes, sessions, compaction, schedules, and subagents. He also flags memory, proactivity, and agent identity as emerging areas still lacking settled standards.

  25. Ethan MollickAI score42

    Interesting to see, given the controversy over the OpenAI release of a series of proofs and what it means for the discipline of mathematics, that at least some of the OpenAI proofs seem to have kicked off extremely rapid iterative advances from a wide community of collaborators.

    Interesting to see, given the controversy over the OpenAI release of a series of proofs and what it means for the discipline of mathematics, that at least some of the OpenAI proofs seem to have kicked off extremely rapid iterative advances from a wide community of collaborators.