Skip to contentSkip to stories

Updated

#Product update

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 20

Sep 20Sun
  1. LMSYS OrgOfficialAI score32

    RLinf adds Cosmos3 support with SGLang, boosting evaluation throughput 3.33x

    AIRLinf, an open-source framework for embodied intelligence and AI agents, now supports Cosmos3 from fine-tuning through robot evaluation. With SGLang inference, it delivers 3.33x end-to-end evaluation throughput, batching inference for 128 parallel environments on 8 GPUs across 500 episodes of the full LIBERO-10 evaluation. RLinf also overlaps CPU simulation with GPU inference to reduce waiting between stages.

    Image from @lmsysorg's post
  2. QwenOfficialAI score34

    Qwen-Image-2.1 now supported in ComfyUI for image generation

    AIQwen-Image-2.1 is now supported in ComfyUI, and Qwen invites users to try it and share their creations. ComfyUI describes it as an open-weights 7B checkpoint that handles both generation and editing, with native 2K image generation and instruction editing from up to 10 reference images in one pass.

  3. OpenBMBOfficialAI score44

    MiniCPM-o Booking Desk: open-source real-time voice appointment agent built on MiniCPM-o 4.5

    AIDeveloper @mrgoodmantweets built MiniCPM-o Booking Desk, an open-source appointment booking agent that uses MiniCPM-o 4.5 for real-time, full-duplex voice and audio-visual interaction. The agent listens, speaks, and reads live booking status from an operator screen, while deterministic state control keeps execution reliable. An appointment is only booked after user confirmation.

    Image from @OpenBMB's post
  4. QwenOfficialAI score34

    Qwen-Image-2.1 edits three marked regions in one prompt

    AIQwen-Image-2.1 supports several ways to specify local edits, including circles marking regions on an image. In the example, the model removes a metal watch, changes hair color to black, and replaces clothing in three circled areas in a single pass.

    Image from @Alibaba_Qwen's post

Sep 19

Sep 19Sat

Sep 18

Sep 18Fri
  1. LM StudioOfficialAI score62

    LM Studio adds Qwen3.8-27B running at up to 144 tok/sec on M5 Max

    AILM Studio announced that Qwen3.8-27B runs at up to 144 tokens per second on an M5 Max MacBook Pro through its partnership with Inco Splash. The post claims up to 3× the decode speed of Ollama, 2× oMLX, and almost 4× when an agent fans out into sub-agents. Inco Splash is described as an open-source inference engine built for the model and Apple silicon, available through the linked LM Studio blog.

  2. Google GemmaOfficialAI score22

    DiffusionGemma runs as a parallel decision model, faster than autoregressive generation

    AIGoogle Gemma's account says DiffusionGemma, running in a Jev-style decision setup, denoises an open canvas in one step rather than generating tokens sequentially, taking about 0.2 seconds on a DGX Spark. It says full bidirectional attention lets every option attend to the full context at once, and that the model inherits Gemma 4's spatial vision capabilities for visual and text decisions.

    Video from @googlegemma's post
  3. Greg BrockmanXAI score44

    ChatGPT now lets users connect multiple accounts to most plugins

    AIChatGPT now supports connecting multiple accounts to most plugins, so users can bring work and personal context into one conversation. Connections are made in the plugin directory at Developers need no changes, though adding a profile tool to their MCP server lets ChatGPT label each account.

  4. Thomas DohmkeXAI score24

    Claude Code adds AGENTS.md support in version 2.1.277

    AIClaude Code version 2.1.277 now reads AGENTS.md when a folder has no CLAUDE.md, according to Anthropic engineer Thariq Shihipar's post. The behavior can be toggled in /config. Thomas Dohmke's main post jokingly says AI is finally aligned, with no further technical detail.

  5. Google AIOfficialAI score47

    Google's weekly recap: Gemini 3.8 Live, Dreambeans, CC, and more

    AIGoogle's weekly recap covers Gemini 3.8 Live and 3.8 Live Extended Thinking, described as its most advanced live dialogue audio models yet. It also notes Dreambeans, a GoogleLabs experiment curating daily personalized stories, is now generally available, and that CC has expanded into a shared agent for household coordination. Google Pics, a Workspace tool for generating and co-creating images, is now GA, alongside AlphaGenome Atlas, DeepMind's interactive genomics discovery platform.

  6. LMSYS OrgOfficialAI score52

    LMSYS blog shows DeepSeek-V4-Flash and Kimi-K3 running on consumer hardware via SSD Expert Pack

    AILMSYS Org announced a blog on running DeepSeek-V4-Flash and Kimi-K3 on consumer hardware using SSD Expert Pack, built by WiCi AI and the SGLang team. Routed experts stay on an NVMe SSD, and the runtime loads only router-selected experts into a GPU cache. On one RTX 5090, 32 GB RAM, and a 2 TB SSD, DeepSeek-V4-Flash MXFP4 decoded at 1.85–1.99 tokens/sec and Kimi-K3 community Q2_K (text-only) at about 0.29 tokens/sec.

    Image from @lmsysorg's post
  7. Google · AI blogOfficialAI score29

    Google co-builds Google Flow tools with two designers for New York Fashion Week runways

    AIGoogle's Envisioning Studio, with Google Labs, co-developed custom Google Flow tools with designers Jane Wade and Sergio Hudson ahead of New York Fashion Week. Wade's Styling Suite let her style runway looks on digital models before producing physical samples, while Hudson's Runway Visualization helped him stage his show within a tight budget. The source says the tools are built with natural language and no coding experience.

  8. KrASIA · Big TechNewsAI score47

    Huawei unveils Atlas 960E superpod linking 4,096 NPUs with near-packaged optics

    AIHuawei unveiled the Atlas 960E superpod, which links up to 4,096 NPUs using near-packaged optics (NPO) and claims eight exaflops at FP8 precision and up to one petabyte of high-bandwidth memory. Its Hi-ONE engine provides 7.2 terabits per second of transmission capacity, and Huawei says 5,500 engines replace 48,000 conventional 800G optical modules, cutting power consumption by more than 550 kilowatts. Huawei has proposed its NPO implementation agreement to the Optical Internetworking Forum, though it has not yet become a standard.

  9. InferactOfficialAI score46

    Kimi K3 serving in vLLM is now 2.2–2.8× faster

    AIInferact, with Red Hat AI, NVIDIA, and Huawei, co-led an optimization effort that makes Kimi K3 on vLLM 2.2–2.8× faster. The work spans scheduling, KDA state handling, and custom MoE kernels. The vLLM project's background post cites those throughput gains on a B300 benchmark against v0.27.1 and links a technical deep dive.

  10. Gemini API ChangelogOfficialAI score26

    Google limits Gemini 2.5 model API access to users who used them recently

    AIGoogle is restricting access to its Gemini 2.5 models to users who have actively used them in the past, to keep performance reliable. The models are not deprecated and remain available through the API until further notice. For new projects, Google recommends its latest models, 3.5 Flash-Lite or 3.8 Flash.

Sep 17

Sep 17Thu
  1. Felix RiesebergXAI score20

    Anthropic's Felix Rieseberg says built-in database and multiplayer simplify team apps

    AIFelix Rieseberg, an Anthropic-associated account holder, says a demo shows teams can build internal tools without handling the database or multiplayer features, which are built in. He calls the demo silly but says it makes building apps for teams very easy. The post names no specific product, version, or figures.

  2. Together AI BlogOfficialAI score31

    Fintech Scales Coding Agent Traffic on Together's Dedicated Model Inference

    AIA global fintech scaled its AI coding agent traffic by running the GLM-5.2 model on Together AI's Dedicated Model Inference, after capacity planning failed to keep pace with unpredictable engineering-hour bursts. The customer gained self-service endpoint provisioning, a metrics API for diagnosing queuing, and live configuration changes that shipped with zero downtime. The setup runs dozens of B200 GPUs at 256K context across multiple replicas.

  3. vLLM BlogOfficialAI score38

    vLLM Adds NVIDIA Hardware Video Decoding to Scale Multi-GPU Video Captioning

    AIvLLM now supports NVIDIA hardware video decoding through PyNvVideoCodec, moving video decoding off the CPU so multi-GPU video captioning can scale to 8 GPUs. In benchmarks on 8xH100 GPUs, GPU-based decoding provides more than double the throughput of the CPU-based decoder for Qwen/Qwen3-VL-8B-Instruct with 8 single-GPU vLLM replicas. The functionality is included in standard CUDA vLLM releases, and PyNvVideoCodec==2.0.4 is required for custom installations.

  4. xAI News (Grok)OfficialAI score42

    Grok Voice Transcribe 2.0 Doubles Accuracy of Predecessor at Same Price

    AIxAI released Grok Voice Transcribe 2.0, a speech-to-text model that is twice as accurate as Grok Voice Transcribe 1.0 at the same price, and ranks first for accuracy among 32 streaming models on the Artificial Analysis leaderboard. Batch transcription costs $0.10 per hour of audio and streaming $0.20 per hour, with diarization, timestamps, and key terms included. Existing Speech-to-Text API integrations gain the improvement with no code changes, and developers must pin grok-voice-transcribe-1.0 to stay on the older model during the transition.

  5. Sherwin WuXAI score44

    ChatGPT for Word launches, bringing ChatGPT and Codex into Microsoft Word

    AIOpenAI has released ChatGPT for Word, letting users access ChatGPT and Codex directly inside Microsoft Word. The post says ChatGPT for Excel and PowerPoint has been growing rapidly, and Word completes that set. The quoted ChatGPT post adds that the tool can draft from notes, rewrite paragraphs, proofread, suggest edits, and flag formatting issues.