Skip to content

Companies & models

NVIDIA Latest news

Follow NVIDIA’s AI chips and ecosystem: new GPUs, CUDA, robotics platforms, and the AI compute market.

8 picksPast 30 days: 6 itemsTotal: 177 items

Updated

Key moments

Since 1993
  1. CompanyNVIDIA founded
  2. ProductGeForce 256 introduced as the first GPU
  3. ProductCUDA announced
  4. ProductDGX-1 deep learning supercomputer announced
  5. ProductVolta V100 GPU announced
  6. ProductA100 GPU announced
  7. ProductH100 GPU announced
  8. CompanyReaches a $1 trillion market value
  9. ProductBlackwell architecture announced
  10. CompanyBriefly becomes the world's most valuable company
  11. ModelNemotron-4 340B released
  12. CompanyFirst company to reach a $4 trillion market value
  13. CompanyPlans to invest up to $100 billion in OpenAI
  14. CompanyFirst company to reach a $5 trillion market value

NVIDIA top picks

Oct 8

TodayOct 8ThuItems 1–8
  1. PyTorch Blog62

    NVIDIA Dynamo adds session-level IDs to route and cache agentic inference

    NVIDIA Dynamo uses a unified session-level identifier to make its inference stack aware of agent sessions, subagents, and their KV cache across turns and tool calls. On SWE-bench, two TP4 MiniMax-M2 replicas on one 8xH100 node gained roughly 12-16% throughput from program-aware scheduling over KV-aware routing alone. The post also describes experimental shared-pool indexing and a proposed KvHint interface for session-aware cache policies in vLLM and SGLang.

    Why it matters: The post explains how session identifiers let an inference stack track agent working sets, with measured throughput gains on SWE-bench and agentic RL rollouts.

Oct 7

Oct 7Wed
  1. NVIDIA Blog67

    NVIDIA and Microsoft Launch RTX Spark Laptops and DGX Station for Windows AI Agents

    NVIDIA and Microsoft announced RTX Spark laptops and compact desktops that run the full NVIDIA AI stack locally, with laptop preorders open today and sales from October 16. Microsoft also announced general availability of Microsoft Execution Containers (MXC), an OS-level infrastructure for agents to run securely in the background, while NVIDIA previewed DGX Station for Windows with 748GB of coherent memory and up to 20 petaFLOPS of FP4 compute.

    Why it matters: The announcement pairs Windows agent infrastructure with local hardware, showing how agents may move onto personal computers and enterprise desktops rather than only cloud services.

Oct 1

Oct 1Thu
  1. NVIDIA Blog62

    NVIDIA Blackwell GPUs power OpenAI's GPT-6 Astra Ultrafast mode in API

    GPT-6 Astra Ultrafast, running on NVIDIA Blackwell GPUs, is now available in the OpenAI API and to eligible ChatGPT Work and Codex users. The source says Ultrafast offers up to 8x faster token generation than Astra Standard mode, which can shorten coding agents' response times between tool calls. OpenAI also says it uses its own models to keep optimizing inference software on NVIDIA GPUs after deployment.

    Why it matters: The source ties a specific speed claim to coding agents' edit-test-debug loops, showing where faster token generation changes developer workflows.

Sep 23

Sep 23Wed
  1. Baseten Blog62

    Baseten launches NVIDIA Nemotron 3 Diarization with four latency profiles

    Baseten has made NVIDIA Nemotron 3 Diarization available as batch, streaming, and real-time diarized transcription presets. The single checkpoint serves four algorithmic latencies from 0.32 to 30.4 seconds, and the post reports DER of 9.8% on AISHELL-4 at the low profile versus 27.2% for Streaming Sortformer v2.1.

    Why it matters: The post shows one checkpoint serving four latency profiles with DER figures against named baselines, useful for judging real-time speaker labeling tradeoffs.

Sep 21

Sep 21Mon
  1. LMSYS Org65

    SGLang adds NVFP4 KV cache for longer context on Blackwell GPUs

    LMSYS Org says NVFP4 KV cache in SGLang fits about 1.78x more context into GPU memory and speeds long-context decoding by up to 78%. Built with Alibaba Qwen and NVIDIA for Blackwell, it stores KV at about 56% of FP8's per-token footprint, with decode throughput up 37%, 58%, and 78% at 32K, 160K, and 1M context. The post reports near-lossless accuracy versus FP8 on GPQA-Diamond and AIME 2025 using Qwen3.5-397B-A17B, and it can be enabled with --kv-cache-dtype nvfp4.

    Why it matters: The post gives specific memory and throughput figures for NVFP4 KV cache in SGLang, showing how the format trades cache footprint against long-context decode speed.

Sep 12

Sep 12Sat
  1. Epoch AI · The Epoch Brief60

    Epoch Brief covers Huawei chips, Nvidia's GDP effect, and GPT-6 Astra benchmarks

    Epoch AI's newsletter reports that Huawei is far behind Nvidia and is unlikely to catch up this decade due to export controls. It also finds official US GDP statistics understate growth by about 0.3 percentage points over the past year, and that GPT-6 Astra set new records on Epoch's evaluations, including the Epoch Capabilities Index.

    Why it matters: The newsletter bundles several analyses of AI chips, GDP measurement, and benchmarks, so it helps readers scan the research agenda behind each finding.

Sep 3

Sep 3Thu
  1. Sundar Pichai72

    NVIDIA to acquire Hugging Face, with Google citing strengthened open model ecosystem

    NVIDIA announced it will acquire Hugging Face, and Sundar Pichai congratulated Jensen Huang and Clement Delangue on the deal. Pichai said Google was an earlier investor in Hugging Face and remains a partner, expecting the deal to strengthen the open model ecosystem.

    Why it matters: Pichai's post confirms Google's prior investment and partnership with Hugging Face, adding context to the acquisition's effect on the open model ecosystem.

Sep 2

Sep 2Wed
  1. NVIDIA · new models on Hugging Face67

    NVIDIA releases Nemotron-3-Labs-Ultra-Math-RL for mathematical proof reasoning

    NVIDIA has published Nemotron-3-Labs-Ultra-Math-RL on Hugging Face, a 550B total, 55B active parameter model for solving difficult math problems and identifying proof mistakes. The model is part of an ensemble that reached gold-medal level at the International Mathematical Olympiad 2026, and it is available for commercial and non-commercial use under the OpenMDW-1.1 license. Deployment is designed for NVIDIA Blackwell or Hopper GPUs, with a recommended minimum of 8× B200 on a single node and a context length of up to 1M tokens.

    Why it matters: The release details the model's math-proof role, its 550B total and 55B active parameters, and its vLLM deployment requirements for teams weighing adoption.