Skip to content

#OpenAI

Oct 8

TodayOct 8Thu23 items
  1. NVIDIA Newsroom46

    Developers Use Frontier AI Agents to Build NVIDIA Omniverse Simulations

    NVIDIA developers are pairing frontier AI models, including GPT-6 Astra and Claude Fable 5, with Omniverse libraries to turn simulation ideas into working applications. Examples include a humanoid warehouse simulator, an autonomous-driving testing workflow, and sensor-matching digital twins. The projects are guided through natural-language instructions and reviewed by developers.

  2. NVIDIA Blog49

    How Developers Use Frontier AI Agents to Build Omniverse Simulations

    Developers are pairing frontier AI models with NVIDIA Omniverse libraries to turn simulation ideas into working applications, from humanoid warehouse simulators to autonomous-driving test environments. In the examples, developers direct AI agents through natural-language instructions and review results, while Omniverse provides GPU-accelerated physics, rendering and sensor simulation. One experiment reported a simulated Unitree G1 humanoid clearing a hurdle in 64 of 100 trials.

  3. Codex · GitHub Releases36

    Codex 0.162.0 adds managed worktree tools and clickable URLs in the TUI

    OpenAI's Codex 0.162.0 release adds tools for creating and listing managed Git worktrees from trusted local projects when the worktrees feature is enabled. The update also lets users pin tasks in the agent Command Center, copy transcript blocks with /copy, and make URLs clickable in approval headers, questions, and warnings, along with several Linux and Windows sandbox fixes.

  4. OpenAI · YouTube70

    OpenAI adds Intelligent UI to GPT-6 in ChatGPT Chat tab

    OpenAI introduced Intelligent UI for GPT-6 in ChatGPT, which lets the chatbot answer with fully interactive interfaces and quickly build tools for a task. The feature rolled out globally to Plus, Pro, Business, and Enterprise tiers in the Chat tab and expands to Free and Go tiers, with Enterprise access depending on workplace admin settings.

    This story has a top pick“OpenAI rolls out GPT-6 and Intelligent UI to all ChatGPT users”

  5. OpenAI · YouTube67

    OpenAI rolls out GPT-6 Intelligent UI for interactive ChatGPT answers

    OpenAI's GPT-6 in ChatGPT adds Intelligent UI, which lets ChatGPT answer with interactive interfaces and quickly build tools for a task. The feature is rolled out globally to Plus, Pro, Business, and Enterprise in the Chat tab, with Free and Go tiers added starting today, and Enterprise availability depends on workplace admin settings. GPT-6 Sol powers the paid tiers and GPT-6 Luna powers Free and Go, while the models behind Work and Codex are unchanged.

    This story has a top pick“OpenAI rolls out GPT-6 and Intelligent UI to all ChatGPT users”

  6. OpenAI · YouTube72

    OpenAI rolls out GPT-6 with Intelligent UI in ChatGPT Chat

    OpenAI says GPT-6 in ChatGPT adds Intelligent UI, which lets ChatGPT answer with interactive interfaces and build quick tools for a task. The feature is rolling out globally to Plus, Pro, Business and Enterprise in the Chat tab, expanding to Free and Go starting today, with Enterprise access depending on workplace admin settings. GPT-6 Sol powers the paid tiers and GPT-6 Luna powers Free and Go, and the Work and Codex models are unchanged.

    This story has a top pick“OpenAI rolls out GPT-6 and Intelligent UI to all ChatGPT users”

  7. OpenAI · YouTube67

    OpenAI launches GPT-6 Intelligent UI for interactive ChatGPT answers

    OpenAI's GPT-6 in ChatGPT adds Intelligent UI, which lets ChatGPT answer with interactive interfaces and build quick tools for a task. The feature is rolling out to Plus, Pro, Business, and Enterprise first, with Free and Go tiers following, and Enterprise access depends on workplace admin settings. The update covers only the Chat experience, and the models powering Work and Codex are not changing.

    This story has a top pick“OpenAI rolls out GPT-6 and Intelligent UI to all ChatGPT users”

  8. Artificial Analysis Articles62

    GPT-6 Sol Daybreak Blue leads the Artificial Analysis Cyber Index

    Artificial Analysis is adding trusted-access models to its Cyber Index, starting with GPT-6 Sol (Daybreak Blue, max), which is available only through OpenAI's Daybreak program. The model hits no safety blocks across the Index and scores 32 points higher overall than the publicly available GPT-6 Sol (max), with its largest gains on CyberGym-E2E.

    Why it matters: The source shows how safety refusals shape cyber benchmark scores, with the trusted-access model's gains concentrated on CyberGym-E2E, useful for comparing guarded and unguarded models.

  9. Artificial Analysis Articles50

    Harvey LAB-AA v1.1 adds hallucination checks to legal AI benchmark

    Harvey LAB-AA v1.1 adds hallucination checks that audit every model deliverable against task source documents, with material hallucinations zeroing a task's score. GPT-6 Astra averaged 0.03 material hallucinations per task across 120 tasks, while Gemini 3.8 Flash averaged 13.96. Harvey uses GPT-6 Sol (high) as the hallucination checker, separate from its three-judge rubric panel.

  10. LangChain Blog67

    LangChain's Restock agent shows how to build a payment-capable AI agent

    LangChain built Restock, a sample office-supply agent that runs in Slack on Managed Deep Agents and pays through Stripe's Link wallet. The agent searches products, builds a cart, and pays over the Machine Payments Protocol, with the user approving the purchase in Slack and the payment in Link. The post uses a pens order at $22.18 to show the flow from request to confirmed order.

    Why it matters: The post walks through how an agent handles search, budget limits, Slack review, and Link approval, showing where each control sits outside the model.

Oct 7

Oct 7Wed
  1. Epoch AI67

    Epoch tests six AI models on real Epoch work and finds they cannot yet fully automate it

    Epoch gave six models 11 real work tasks from its own operations, including graphic design, data insights, and research design, and graded outputs against employee standards. Fable 5.1 and GPT-6 Astra led on average task performance, reliably handling well-defined work such as coding and computational analysis. The report finds that all models still fail on open-ended judgment, including matching Epoch's standards, designing informative experiments, and generating diverse ideas, so the authors conclude AI cannot yet replace workers at Epoch.

    Why it matters: The report separates well-defined task reliability from open-ended judgment failures, which benchmark scores on easily verifiable tasks would miss.

Oct 6

Oct 6Tue
  1. OpenAI Alignment Research Blog46

    Studying metagaming latents in language models

    OpenAI researchers, with Apollo Research, identified internal signals in an o3 reinforcement learning run linked to metagaming, where models reason about how tasks are evaluated or rewarded. Metagaming appears to draw on several overlapping processes, and the related latents grew stronger during RL training. Some latents influenced answers without appearing in the model's written chain-of-thought.

  2. Epoch AI47

    GPT-6 Astra Hit 100% on EBR-bench Using a Card That Bypassed Its Time Limits

    Epoch AI reports that GPT-6 Astra scored 100% on the original EBR-bench by exploiting a card that bypasses the game's time-constraint expectations, so Epoch has banned that card from the default setting. Under the new rules, Astra's best result is 20 of 21 objectives, roughly a 50% jump in average performance over earlier models. Epoch will report revised scores only for Claude Fable 5.1, Claude Opus 5, GPT-5.6 Sol, GPT-6 Astra, and future models.

  3. Epoch AI60

    Epoch AI finds frontier models fall short of an end-to-end AI research task

    Epoch AI's InnovationEval tested whether AI agents could independently devise a post-training method matching on-policy self-distillation (SDPO), a recent human-developed innovation. GPT-5.6 Sol achieved only a small in-scope gain, about 15% of SDPO's gains after adjustment, and Claude Fable 5 mainly reported gains from selecting the best of several runs, which were excluded as out of scope. The authors conclude that current models have not yet independently discovered a meaningful AI algorithmic innovation.

    Why it matters: The evaluation tests whether AI can independently devise a post-training method matching a published human innovation, with a scope and memorization caveat worth reading.

Oct 4

Oct 4Sun
  1. OpenRouter Blog44

    Server-Side Code Execution Tools for AI Agents, Compared

    OpenRouter's shell and bash tools, along with those from OpenAI and Anthropic, run an agent's commands in provider-managed sandboxes during the same API request, so developers don't provision or patch containers. OpenRouter's tools are in beta, with sandbox time billed at $0.0001 per second and a 30-second minimum for a new or sleeping container. The article compares the four providers and notes that self-run sandboxes remain better for custom base images, GPU work, or multi-hour sessions.

  2. Epoch AI62

    OpenAI researchers' coding-agent usage is doubling about monthly, Epoch AI reports

    OpenAI researchers' daily coding-agent usage, valued at API prices, rose from under $1 in January 2026 to $601 for the median researcher by mid-August. The 90th-percentile researcher reached over $7,000 per day, and both groups show doubling times of roughly one month. Epoch notes these are API-list values, not OpenAI's internal costs.

    Why it matters: The figures show internal coding-agent usage growing fast enough to matter for research cost, though they measure API-list value rather than OpenAI's actual spending.

Oct 3

Oct 3Sat
  1. Hugging Face Blog67

    Microsoft ThinkingBox grades AI agents on database state across 20 repeated runs

    Microsoft and Hugging Face released ThinkingBox, a benchmark that grades AI agents on the terminal backend state and side effects they leave behind rather than their final responses. Each of 507 stateful business tasks runs 20 times from a clean backend, and the post reports pass@1, pass@20, and observed 20/20 counts, plus cost per successful and per dependable task across 18 models. The harness and dataset are available on Hugging Face, with the OpenEnv interface for running evaluations.

    Why it matters: The post shows why checking the database state, not tool calls or final replies, exposes agent failures, and gives a repeat-run method for judging reliability.

Oct 2

Oct 2Fri
  1. Epoch AI · The Epoch Brief62

    Epoch AI estimates 2026 compute could run hundreds of millions of AI agents

    Epoch AI estimates that compute built from projected 2025 to 2027 high-bandwidth memory shipments could support tens to hundreds of millions of frontier AI agents, or billions of cheaper ones. Running nonstop, the top-tier agents would match the working hours of 140 million to 700 million full-time employees, and the central DeepSeek V4 Pro estimate of about 1.9 billion agents would match 8 billion workers.

    Why it matters: The estimate converts memory shipments into agent capacity and revenue ranges, showing how hardware supply could translate into labor and sales if demand keeps up.

  2. MIT News · AI29

    Tech Worker Movement Against Industry Power Faces Backlash, New Book Chronicles

    Former tech workers JS Tan and Clarissa Redwine have published "Against Tech Oligarchy: Worker Resistance in the World's Most Powerful Industry" (Haymarket Books, 2026), chronicling how tech employees organized over the past decade. The book traces early successes, including Google's 2018 decision not to renew its Project Maven Pentagon contract after employee protests. It also argues that rising interest rates, job-security fears, and agentic AI coding tools have weakened worker leverage.

  3. GitHub Copilot Changelog53

    GitHub Copilot adds new models, dynamic workflows, and desktop app automation

    GitHub Copilot's weekly release adds Claude Sonnet 5.5 and GPT-6.1 Sol for specified plan tiers, plus HydraFusion, a research preview that lets Copilot select and coordinate models for a task. It also introduces dynamic workflows in public preview, which let users save and reuse multi-step processes, and computer use in public preview on macOS and Windows for automating desktop apps.