Skip to contentSkip to stories

Updated

#Deployment/Engineering

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 9

Oct 9Fri
  1. vLLMOfficialAI score33

    vLLM says Rubin GPUs run its Blackwell kernels unmodified

    AIvLLM says a Rubin GPU has 3.5x the FP4 FLOPS, about 2.4x the HBM bandwidth, and 1.7x the NVLink bandwidth of a GB200. The faster NVLink lowers the cost of AllReduce and all-to-all operations in MoE serving. vLLM's Blackwell kernels run on Rubin unmodified, and FlashInfer 0.7.0 adds Rubin-tuned attention, GEMM, and MoE kernels.

    Image from @vllm_project's post
  2. vLLMOfficialAI score28

    vLLM uses NVIDIA locality domains to speed MoE decode up to 1.2x

    AIvLLM's locality-aware MoE shards FC1 and FC2 weights column-wise and launches one kernel per CUDA 13.4 locality domain using Green Contexts, so each SM reads only local HBM. Early results show MiniMax M3 MoE layers up to 1.2x faster in decode. This is post 4 of a 5-post thread.

    GIF from @vllm_project's post
  3. vLLM BlogOfficialAI score62

    vLLM adds support for NVIDIA Vera Rubin NVL72 with 7.8x throughput over GB200

    AIvLLM now supports NVIDIA Vera Rubin NVL72, with daily container builds and support for models from DeepSeek, Moonshot AI, Z.ai, and MiniMax. In early AgentX benchmarks, vLLM running MiniMax M3 delivered up to 7.84x the throughput per GPU of GB200 NVL72 at matched interactivity. The post is an early look, and the team expects further gains from ongoing optimizations.

  4. NVIDIA · new models on Hugging FaceOfficialAI score38

    NVIDIA releases GR00T N2 ONNX checkpoint for SSD pick-and-place tasks

    AINVIDIA publishes a GR00T N2 checkpoint, step 18200, as a native split ONNX export on Hugging Face for SSD pickup and placement. The model trained on 196 pickup and 149 placement episodes, with 310 training and 35 validation episodes, and one shared set of graphs serves both tasks through a host-side task selector. The package is not a TensorRT engine or a robot-ready policy, and NVIDIA does not assert full ONNX/eager numerical parity or robot success rates.

  5. Vercel DevelopersOfficialAI score22

    Liquid AI's d1 model now available on Vercel AI Gateway

    AIVercel says Liquid AI's d1 model is live on AI Gateway under the identifier liquid/d1. The model supports vision inputs for classifying, routing, and scoring decisions.

  6. Latent SpaceBlogAI score43

    Pushmeet Kohli and Sal Candido discuss why AlphaFold did not solve protein folding

    AIGoogle DeepMind's Pushmeet Kohli and Biohub's Sal Candido discuss in a panel moderated by Brandon Anderson whether scaling compute and data is enough for AI-driven biology. The conversation covers scaling laws in biological data, why protein structure prediction is not solved, and the limits of static structure prediction. The panel also addresses how the protein language models may hold scientific knowledge not yet unlocked.

  7. Vercel DevelopersOfficialAI score24

    Microsoft Decision-1 model now available on Vercel AI Gateway

    AIVercel says Microsoft's Decision-1 model, listed as microsoft/microsoft-decision-1, is now available on AI Gateway. The model classifies text, routes requests, and scores responses, returning structured answers with calibrated probabilities.

  8. Elon MuskXAI score22

    Musk says Grok can build a simulation of anything, citing Cannae example

    AIElon Musk posts that Grok can make a sim of anything, sharing a simulation of the Battle of Cannae and the double envelopment that nearly destroyed Rome. The post credits the Grok bot for the simulation, which is linked to the @UpdatingOnRome account.

  9. AI EraNewsAI score54

    Anthropic says Claude helped produce a full-sky ultraviolet map from NASA satellite data

    AIAnthropic has released what it calls the first complete ultraviolet all-sky map, built from data NASA's satellite had gathered over roughly ten years. According to the excerpt, Claude helped fill in the gaps over a few days, leaving no blank regions in the sky. The source is an excerpt only, so the methods and the star count are not confirmed here.

  10. AI EraNewsAI score36

    Anthropic says AI should test and fix its own code

    AIAnthropic recommends that AI coding tools like Claude Code test and revise their own output, rather than leaving debugging to developers. The source describes a developer who built a small app with Claude Code and added an AI customer-service bot, but the text provided is only the opening scenario.

  11. SGLangOfficialAI score34

    SGLang reports inference speedups for MLA, MoE, and KDA kernels

    AISGLang says restructured MLA decode kernels on Rubin, which fit a deeper pipeline in 327 KiB of shared memory, deliver a 16% speedup at batch 16 with 128K context and bit-identical output. The post also reports 20% faster full FP8 MLA at batch 1 and 20% faster KDA verify kernels after keeping weights in registers and reducing synchronization. It additionally covers fusing MoE finalization, the shared expert, 8-GPU all-reduce, and RMSNorm into one collective kernel.

  12. SGLangOfficialAI score37

    Miles runs full RL loop on Rubin GPUs with SGLang and Megatron

    AIMiles runs the full reinforcement learning loop on Nvidia Rubin, using SGLang for rollout, Megatron for training, and one container image. On a single 4-GPU tray, Qwen3-30B-A3B's GSM8K reward rises from about 45% to about 95% over 50 rollouts, matching the GB300 curve. The post also reports DeepSeek-V4-Flash end-to-end rollout and training, and Qwen3.5-35B-A3B agentic RL with 64 concurrent mini-SWE-agent sandboxes on SWE-bench Verified, where reward holds near 0.6 and median response length falls about 30%.

  13. GitHub Copilot ChangelogOfficialAI score36

    GitHub Copilot adds local sandboxing and separate accounts in weekly releases

    AIGitHub makes local sandboxing generally available in Copilot CLI, the Copilot app, and VS Code sessions using Agent Host, limiting agents' access to files, networks, and credentials at no extra cost. The Copilot app now lets users sign in with separate GitHub accounts for the Copilot license and for repositories. Copilot CLI's /model command lists local models from a running Ollama instance alongside cloud models, and VS Code 1.141 adds a side-by-side agent session grid and worktree cleanup.

  14. elvisXAI score44

    Tinker cuts long-context token prices, making agent RL rollouts cheaper

    AITinker has cut prices up to 70% on long-context prefill and sampling, which now cost the same as short context. The cut lowers the cost of agentic RL rollouts, which spend most of their tokens re-reading growing context, and of evaluating trained models on long inputs. Tinker also added GLM-5.3-Flash and DeepSeek-v4.1-Flash for cost-efficient long-context work.

  15. SGLangOfficialAI score52

    SGLang adds Rubin optimizations that speed up Kimi K3 inference

    AISGLang says it worked with NVIDIA to optimize attention, MoE, and speculative verification kernels for Kimi K3 inference on early-access Rubin hardware. It reports up to 20% faster FP8 MLA at batch 1 with 128K context, 20% faster KDA verification with bitwise-identical output, and a 5.9% end-to-end speedup from MoE tail fusion that removes 276 kernel launches per decode step. The post also says SGLang powers Miles' end-to-end RL training on Rubin, including agentic RL with 64 concurrent sandboxes on the Vera CPU.

  16. laurenXAI score32

    Lauren Tan proposes cutting software interviews to two technical rounds

    AILauren Tan (@poteto) says software engineering interviews could shrink to two technical rounds: a system design round that tests how clearly candidates articulate ideas and engineer solutions, and an onsite project building a real thing with agents that tests how well they turn intent into high-quality outcomes. She says the other questions traditionally asked are no longer necessary.

  17. SiliconANGLE · AINewsAI score22

    IBM's Bruno Aziza outlines enterprise AI orchestration ahead of TechXchange 2026

    AIIBM group vice president Bruno Aziza says enterprises need platforms that track and coordinate the agents their employees create, because the number of agents will outpace what organizations can manage. He says sovereignty requires control across data, technology, operations and regulation, not only data location. IBM TechXchange 2026 takes place Oct. 26–29 in Atlanta, with customer examples from Citigroup, CVS Health and DoorDash.

  18. Ars Technica · AINewsAI score38

    Ukrainian drones knock out AI data center of Russian firm Yandex

    AIUkrainian drones knocked out an AI data center belonging to Russia's Yandex, according to Reuters, which also reported outages at a dating app, a real estate aggregator, a self-publishing platform and a Premier League soccer club's website. The article says Iranian drones earlier in 2026 hit Amazon data centers in Bahrain and the United Arab Emirates, and that Ukraine has expanded its drone strikes on Russian military, energy and warehouse targets.

  19. ElevenLabs BlogOfficialAI score23

    What is voice activity detection and how does it work?

    AIVoice activity detection (VAD) classifies short audio frames, typically 10-30 milliseconds, as containing speech or not. It returns a yes-or-no decision that tells downstream tools such as speech-to-text, LLMs, and turn planners whether to process or wait. VAD does not transcribe words or decide when a speaker has finished, which is the job of endpointing systems.

  20. ThariqXAI score32

    Claude Opus 5.5 ports a side project to Claude Managed Agents

    AIBefore joining Anthropic, Thariq spent about two weeks building a side project with Opus 4 using the Agent SDK. That version needed a constantly running process and did not work well. A single prompt to Opus 5.5 ported it to Claude Managed Agents, which he says made it considerably more reliable.

  21. The Next PlatformNewsAI score46

    Upscale AI unveils SkyHammer scale-up switch ASIC to compete with Nvidia NVLink

    AIUpscale AI says its SkyHammer scale-up switch ASIC will deliver an aggregate bandwidth of 115.2 Tb/sec and support up to 576 accelerators in a single networking tier. The company is partnering with Nvidia on NVLink Fusion while also developing an alternative to Nvidia's NVSwitch for AI clusters. The article notes that Upscale AI has raised $500 million in total funding and has a $2 billion valuation.

  22. Andrew CurranXAI score55

    Prime Agent swarm rewrites itself in Rust, reaching input 13x faster

    AIPrime Intellect says Prime Agent used a swarm of over 2,000 agents to rewrite itself end to end in Rust over two weeks. The rewrite ran across 10,000+ sandboxes and over 200 billion GLM-5.3 tokens, and the company says usable input now arrives about 13 times faster with 83% less startup memory.

    Image from @AndrewCurran_'s post
  23. Prime Intellect BlogOfficialAI score65

    Prime Agent is rewritten in Rust by a swarm of agents

    AIPrime Intellect says it rewrote its Prime Agent coding tool in Rust, using a swarm of more than 2,000 agents over two weeks. The company reports cold start to typing about 13 times faster than the TypeScript version, and memory use over 80% lower after startup. Prime Agent remains open source and adds native Windows support in beta and Homebrew installation.

    Why it matters: The post shows how a multi-agent swarm rewrote a coding agent with parity checks, giving a concrete case of agent-driven software engineering with measured results.

  24. Bertholomus AIXAI score22

    DeepSeek TP=2 and TP=4 kernel kits released on GitHub

    AIA GitHub post announces that kernel kits for DeepSeek's TP=2 and TP=4 recipes are now live. The post links to a repository named deepseek-v4.1-tensorfold-tp2-2xgb10 and gives no further details on performance or features.

  25. TinkerOfficialAI score32

    Tinker removes extra prefill charges for 128k and 256k context

    AITinker says prefill for 128k and 256k context no longer costs extra, an effective discount of over 2x for long-context models including Kimi K2.6, gpt-oss-120b, and Inkling. Prefill is also discounted for seven Qwen and Nemotron models, and sampling is cut for Qwen3.5-9B and 9B-Base.

  26. TinkerOfficialAI score40

    Tinker adds GLM-5.3-Flash and DeepSeek-v4.1-Flash models

    AITinker adds GLM-5.3-Flash and DeepSeek-v4.1-Flash, both of which natively accept image inputs and use efficient attention architecture. GLM-5.3-Flash costs 4-5 times less on Tinker than GLM-5.3. Long-context options for Qwen3.5-4B and Qwen3.6-35B-A3B are also live.

  27. ElevenLabs BlogOfficialAI score58

    ElevenLabs releases synthetic voice detection in ElevenAgents for business calls

    AIElevenLabs is releasing synthetic voice detection in ElevenAgents, which analyzes a caller's speech in the first few seconds and labels it as human or AI generated. Businesses can then set rules, such as prioritizing verified humans, limiting AI callers to bounded exchanges, or stopping impersonation attempts before sensitive actions. The feature is available now to enterprise customers supported by its Forward Deployed Engineering team, and will reach a broader group of enterprise customers later this month as a configurable option.

  28. Soumith ChintalaXAI score22

    Tinker cuts prices up to 70% as efficiency improves

    AITinker, an API for training and fine-tuning models, is cutting prices by up to 70% after engineering efficiency gains. The company says the savings are passed on to customers, and that buying more produces greater savings. GLM-5.3-Flash and DeepSeek-v4.1-Flash are also now available on Tinker for long-context work.

  29. Rohan PaulXAI score46

    Pine launches a cloud computer built for AI agents

    AIPine has launched a cloud computer for AI agents that developers create through an SDK and direct with plain-language tasks. Running GPT-5.6 Luna, Pine reports about 1/20 the model-token cost of GPT-5.6 Sol with Codex on SaaS-Bench v1.1, with a 78.3% checkpoint score.

    Image from @rohanpaul_ai's post
  30. Hacker News · AI (150+ points)BlogAI score38

    Show HN: big-arrow-on-the-screen lets AI agents draw arrows and text on macOS

    AIbig-arrow-on-the-screen (bigarrow) is a MIT-licensed macOS command-line tool and skill for Claude Code and Codex that draws arrows, boxes and text over any window. Clicks pass through, keyboard focus stays put, and each arrow removes itself after a set duration or when its agent process ends. The tool only points; it never clicks, types or captures the screen, and it requires no macOS permission to draw.

  31. Interconnects (Nathan Lambert)BlogAI score47

    Lambert expects rapid AI infrastructure gains but not general superintelligence

    AINathan Lambert says coding agents will soon outperform top researchers at GPU engineering and infrastructure, making AI experimentation far easier. He expects inference cost to fall near-exponentially over the coming years and predicts pretraining research for architecture and data selection could be automated in 2-3 years. He argues these gains will not make models dramatically different in nature.

  32. Prime IntellectOfficialAI score46

    Prime Agent swarm of 2,000+ agents rewrites itself in Rust

    AIPrime Intellect says its Prime Agent orchestrated over 2,000 agents over two weeks to rewrite the agent in Rust. The run used more than 10,000 sandboxes, over 200B GLM-5.3 tokens, and 16,000 agent-to-agent messages. The company says the rewritten agent reaches usable input about 13 times faster and uses 83% less startup memory.

    Video from @PrimeIntellect's post