Skip to contentSkip to stories

Updated

All AI news

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 21

Sep 21Mon
  1. TechNode · AINewsAI score36

    IQAX Pushes eBLs, AI, and Digital Twins to Connect Global Trade Data

    AIIQAX has surpassed one million electronic Bills of Lading (eBLs), built on the GSBN blockchain and supporting DCSA, BIMCO, and ISO standards. The company is combining AI, IoT, and digital twins to move supply chain management from visibility toward predicting risks, with its AI-powered IoT platform covering more than 200 regions and 11,000 city pairs and about 92,000 connected devices.

Sep 20

Sep 20Sun
  1. OpenBMBOfficialAI score22

    MiniCPM5-2B on a local Mac correctly reconciles naproxen medication history

    AIOpenMed reports that MiniCPM5-2B, run on a local Mac, correctly kept naproxen in medication history rather than the current-medication export after a newer note said it was stopped. Every graph connection in the run links back to its source, using fictional clinical notes.

  2. swyxXAI score22

    Jev Podcast Episode Announced by Latent Space Host swyx

    AIswyx announced a Latent Space podcast episode featuring Jev, subscribable on Apple and YouTube, and thanked guests Allen Park and Ke. A quoted post from @CompleteSkeptic claims Jev is a frontier model with 20-200x faster speed and 40-400x lower cost, but this post itself adds no verified details.

    Image from @swyx's post
  3. LMSYS OrgOfficialAI score32

    RLinf adds Cosmos3 support with SGLang, boosting evaluation throughput 3.33x

    AIRLinf, an open-source framework for embodied intelligence and AI agents, now supports Cosmos3 from fine-tuning through robot evaluation. With SGLang inference, it delivers 3.33x end-to-end evaluation throughput, batching inference for 128 parallel environments on 8 GPUs across 500 episodes of the full LIBERO-10 evaluation. RLinf also overlaps CPU simulation with GPU inference to reduce waiting between stages.

    Image from @lmsysorg's post
  4. OpenBMBOfficialAI score44

    MiniCPM-o Booking Desk: open-source real-time voice appointment agent built on MiniCPM-o 4.5

    AIDeveloper @mrgoodmantweets built MiniCPM-o Booking Desk, an open-source appointment booking agent that uses MiniCPM-o 4.5 for real-time, full-duplex voice and audio-visual interaction. The agent listens, speaks, and reads live booking status from an operator screen, while deterministic state control keeps execution reliable. An appointment is only booked after user confirmation.

    Image from @OpenBMB's post
  5. OpenBMBOfficialAI score35

    OpenBMB's Augury model improves on-device plant ID for farmers

    AIA developer's Augury plant identification model, built on an OpenBMB model, raised photo top-1 accuracy from 71.8% to 80.2% by merging duplicate species keys and adding PCA whitening. The next steps are reaching 90%+ accuracy and building a phone GUI so farmers can use it on-device.

Sep 18

Sep 18Fri
  1. LM StudioOfficialAI score22

    Splash engine released as open source on GitHub

    AIThe Splash engine, posted by LM Studio, is now available as open source on GitHub. The post provides only a link to the incoai/splash repository and includes no further technical details.

  2. LM StudioOfficialAI score20

    LM Studio adds Splash engine for running incoai models on Mac

    AILM Studio users can enable the Splash engine under Settings > Runtime > Experimental backends and download supported models by searching "incoai." The post recommends an M3 or newer Mac running macOS 26.4 with 36GB+ RAM.

    Image from @lmstudio's post
  3. LM StudioOfficialAI score62

    LM Studio adds Qwen3.8-27B running at up to 144 tok/sec on M5 Max

    AILM Studio announced that Qwen3.8-27B runs at up to 144 tokens per second on an M5 Max MacBook Pro through its partnership with Inco Splash. The post claims up to 3× the decode speed of Ollama, 2× oMLX, and almost 4× when an agent fans out into sub-agents. Inco Splash is described as an open-source inference engine built for the model and Apple silicon, available through the linked LM Studio blog.

  4. LMSYS OrgOfficialAI score52

    LMSYS blog shows DeepSeek-V4-Flash and Kimi-K3 running on consumer hardware via SSD Expert Pack

    AILMSYS Org announced a blog on running DeepSeek-V4-Flash and Kimi-K3 on consumer hardware using SSD Expert Pack, built by WiCi AI and the SGLang team. Routed experts stay on an NVMe SSD, and the runtime loads only router-selected experts into a GPU cache. On one RTX 5090, 32 GB RAM, and a 2 TB SSD, DeepSeek-V4-Flash MXFP4 decoded at 1.85–1.99 tokens/sec and Kimi-K3 community Q2_K (text-only) at about 0.29 tokens/sec.

    Image from @lmsysorg's post
  5. InferactOfficialAI score46

    Kimi K3 serving in vLLM is now 2.2–2.8× faster

    AIInferact, with Red Hat AI, NVIDIA, and Huawei, co-led an optimization effort that makes Kimi K3 on vLLM 2.2–2.8× faster. The work spans scheduling, KDA state handling, and custom MoE kernels. The vLLM project's background post cites those throughput gains on a B300 benchmark against v0.27.1 and links a technical deep dive.

Sep 17

Sep 17Thu
  1. OpenBMBOfficialAI score36

    OpenBMB's MiniCPM5-2B runs offline on-device with 128K context

    AIOpenBMB's MiniCPM5-2B is a 2.5B-parameter model with native 128K context, offering hybrid Think and No-Think modes in one checkpoint. Users can download it from Hugging Face and run it fully offline on-device, as RunAnywhere demonstrated. In a demo, the model first called a puzzle impossible, then corrected itself and wrote a working verifier.

  2. KrASIA · Big TechNewsAI score44

    Qianjue founder says robotics will have no single "ChatGPT moment"

    AIQianjue Technology founder Gao Haichuan argues that robotics will not see one breakthrough that suddenly lifts the whole industry, and he judges the company by deployment results rather than research papers. Qianjue, founded in 2023, has completed a Series A+ round worth a nine-figure RMB sum, with first orders coming from restaurant, cleaning, and hotel service robots. Gao says customers care about task completion, failure rates, and price rather than whether a predictive world model is used.

  3. LM StudioOfficialAI score44

    LM Studio adds session history search and @ session references

    AILM Studio's new Introspection feature lets its Bionic agent search its own session history, improving handling of long-term context across multiple compactions. Users can also reference other sessions directly in the composer with an @ mention.

    Video from @lmstudio's post
  4. Hacker News · Launch HN, YC launches (10+ points)BlogAI score58

    Skillsync launches tool to move AI chat sessions across coding agents

    AISkillsync, a Y Combinator W26 company, launched a tool that converts AI chat sessions between coding agents, including messages, reasoning and tool calls. The conversion engine txcript is open source, and the local-first app runs on Mac with a CLI and MCP support. Only sessions shared into team workspaces leave the user's machine.

  5. Daniel HanXAI score44

    Unsloth Desktop adds multi-user accounts and faster GRPO training

    AIUnsloth Desktop now supports multi-user accounts, alongside a revamped Docker image and custom Jupyter Notebook with custom themes, titles, and expandable cells. The update adds RDNA1+2 support, ARM64 Windows CUDA support, faster GRPO, and FP8/INT8 image diffusion support for 2x faster inference.

  6. Unsloth AIOfficialAI score60

    Unsloth Docker image lets users train and run 500+ models locally

    AIUnsloth announced that its Docker image now lets users train and run more than 500 models locally with no setup required. The image works on NVIDIA and AMD hardware and supports a new GUI or notebook workflow. The post links to an installation guide and the GitHub repository.

    Image from @UnslothAI's post
  7. OpenBMBOfficialAI score40

    OpenMed and MiniCPM5-2B demo local agentic clinical AI workflow

    AIOpenMed paired with MiniCPM5-2B to demonstrate a local clinical AI workflow combining privacy-preserving data processing with a compact model's tool use and long-context reasoning. OpenMed masks sensitive identifiers and extracts clinical context before MiniCPM5-2B calls tools, compares lab results, and generates clinical handoffs with source references. The post presents this as an example of keeping inference on local, resource-constrained hardware.

    Image from @OpenBMB's post
  8. OpenBMBOfficialAI score29

    Kahya-TTS: Turkish speech model fine-tuned from VoxCPM2 on 100 hours

    AIDeveloper Alican Kiraz fine-tuned OpenBMB's open-source VoxCPM2 voice model on nearly 100 hours of natural Turkish speech, creating Kahya-TTS for Turkish text-to-speech. The project shows how open-source voice models can be adapted to new languages and specialized datasets. The model is available on Hugging Face.

    Image from @OpenBMB's post

Sep 16

Sep 16Wed
  1. OpenBMBOfficialAI score20

    OpenBMB praises Dubedo's VoxCPM2-based voice cloning and dubbing studio

    AIOpenBMB says Dubedo is the kind of product it hoped VoxCPM2 would enable, citing speaker-aware cloning, multilingual generation, and an editing studio. The post praises @dubedostudio's work, while the background post describes Dubedo as dubbing into 30 languages with per-speaker voice cloning and a beta open for trial.

  2. Kling AIOfficialAI score38

    Fountain 0's ODYSSEUS: The Fall, fully generated with Kling 3.0, is out

    AIFountain 0's new feature film ODYSSEUS: The Fall, directed by Ash Koosha, is now available in full with every shot generated by Kling 3.0. The team previously premiered Dreams of Violets at the 2026 Tribeca Festival as the first AI feature film accepted into a major film festival.

    Video from @Kling_ai's post
  3. TinkerOfficialAI score32

    Sundial trains Inkling-Small to fix LaTeX errors in under a second

    AISundial fine-tuned Thinking Machines' Inkling-Small with RLVR on 3,978 verified TeX.StackExchange fixes, using rewards for compilation and PDF match and penalties for removed content. The trained model fixes 83.7% of LaTeX errors in under one second at $0.0013 per fix, according to the post. Sundial says it is rolling out the model in its editor, applying fixes as suggestions and rebuilding the PDF.

  4. Kilo (acq. by Anaconda)OfficialAI score40

    Kilo Mobile lets users run full AI agent loops from their phone

    AIKilo Mobile now lets users spawn Cloud Agents, start sessions on remote machines, and dictate prompts by voice from a phone. Users can also review and comment on pull requests and approve Security Agent remediations without a laptop. On iPhone, Live Activities show session status on the Lock Screen when an agent needs input.

    Image from @kilocode's post

Sep 15

Sep 15Tue
  1. Zed BlogOfficialAI score72

    Zed launches Delta public beta to replace pull requests with agent threads

    AIZed has launched the public beta of Delta, a multiplayer environment for coding with agents and reviewing their work, which replaces pull requests with shared threads. Delta is built on DeltaDB, which records edits and messages between Git commits, and it is free during the beta, with paid plans for individuals and teams to follow.

    Why it matters: The post explains how Delta replaces pull requests with shared agent threads and DeltaDB, showing a concrete alternative to the GitHub review workflow.

  2. VercelOfficialAI score22

    Delphi ships 100+ deploys daily on Vercel's Python backend

    AIDelphi, which turns experts' knowledge into digital minds, runs its Python backend on Vercel with a 10-person team and no dedicated infrastructure role. The stack uses Vercel Workflows for long-running agents and Vercel Queues for background jobs. The team reports more than 100 production deploys a day.

  3. Air Street PressBlogAI score39

    Air Street Capital leads $40 million Series A in Jack & Jill, an AI career agent platform

    AIAir Street Capital led Jack & Jill's $40 million Series A, with Madrona joining and Creandum and Entrepreneurs First investing again, following a $20 million seed round less than a year earlier. Jack & Jill uses AI agents named Jack, which helps candidates plan career moves and search job postings, and Jill, which helps companies recruit from opted-in candidates. The company says it has arranged 25,000 interviews and plans 5,000 more each month.

  4. TechNode · AINewsAI score38

    Twoo Adds AI as a Third Member to Two-Person Relationships

    AIShanghai-based Twoo is building an AI that participates in conversations between two users, sharing context, remembering experiences, and helping them plan activities together rather than serving one person. Founder Cobe Chen says the product's monetization includes in-app "shells" that unlock outfits for its octopus AI character, plus premium memberships for professional uses such as collaborative writing and teaching. The company is preparing overseas expansion into North America, Japan, and South Korea.

  5. Kilo (acq. by Anaconda)OfficialAI score22

    Kilo App launches on Product Hunt for iOS and Android

    AIKilo announces that its Kilo App is live on Product Hunt, letting users start coding agents, check sessions, and review pull requests from iOS and Android. The company asks supporters to upvote or comment on its Product Hunt listing.

    Image from @kilocode's post

Sep 14

Sep 14Mon
  1. vLLM BlogOfficialAI score53

    Novita AI open-sources Chord, a W4A16 MoE kernel for Kimi K2.x on vLLM

    AINovita AI has open-sourced Chord, a W4A16 MoE CUDA operator with BF16 activations, INT4 weights and group-32 scales, built for Kimi K2.x serving shapes. Measured per layer against public Humming, it reports 1.11–1.20x on H200 EP8 prefill, 1.17–1.33x on H200 TP8 serving, and 1.81–2.15x on B300 EP8 decode against an untuned Humming default. Integration of the grouped operators with vLLM's Humming backend is still a work in progress.

  2. InferactOfficialAI score42

    Inferact and Google Cloud partner to make TPUs first-class in vLLM

    AIInferact and Google Cloud announce a partnership to make Google TPUs a first-class platform in the vLLM open-source project. The collaboration targets production serving features, optimized kernels, a native PyTorch path via TorchTPU, and day-0 support for frontier model releases. A community program will offer shared TPU capacity and review and design help from vLLM core maintainers, with all outputs released as open source.

    Image from @inferact's post
  3. RadixArkOfficialAI score43

    SGLang-Diffusion runs MiniMax H3 video generation faster than playback

    AIRadixArk's SGLang-Diffusion, paired with VDN-H3, generates 14.4 seconds of 768p video in 9.0 seconds on 8× B200 GPUs. The 8-step denoising alone takes 6.9 seconds, which is over 2× real time, and the team reports no measured quality regression against dense 50-step MiniMax H3 across 103 test prompts.

  4. Intern Large ModelsOfficialAI score25

    Intern-S2-397B gets Day-0 support in vLLM

    AIIntern-S2-397B, a model built for long-horizon scientific research, now has Day-0 support in vLLM. The model brings multimodal, reasoning, coding, and scientific agent capabilities, and vLLM has published a run recipe for it.

Sep 13

Sep 13Sun
  1. Ian Johnson 🔬🤖XAI score34

    Flying through 30 million embeddings as a video game

    AIIan Johnson visualizes 30 million jina-v5-nano embeddings from 12 billion tokens across multilingual FineWeb, StarCoder, The Pile, and RedPajama as a flyable video game. He frames the project as making data exploration engaging rather than a chore.

    Video from @enjalot's post

Sep 11

Sep 11Fri
  1. InferactOfficialAI score34

    vLLM adds day-0 support for DeepSeek v4.1 Flash across six NVIDIA GPUs

    AIInferact says vLLM now supports DeepSeek v4.1 Flash on day zero across H100, H200, B200, B300, GB200, and GB300 GPUs. SemiAnalysis independently verified the NVIDIA support, while the post notes AMD vLLM still does not work with the model. Serving recipes are available at recipes.vllm.ai.

  2. VercelOfficialAI score26

    Tailscale's Aperture model router is built on Vercel AI Gateway

    AITailscale offers instant access to hundreds of models for any user in a secure tailnet through its customer-facing model router, Aperture. Aperture is built on Vercel's AI Gateway and offers zero data retention, zero markup with free BYOK, and cost and usage data on every response.

  3. LlamaIndex 🦙OfficialAI score22

    LlamaParse adds high-effort page-level confidence scores with explanations

    AILlamaParse's new high-effort mode provides granular page-level confidence scores with text explanations of parsing quality. The scoring also checks the original document as an extra verification step. High-effort mode costs 5 additional credits per page and is meant to be used only when needed.

    Image from @llama_index's post