Skip to contentSkip to stories

Updated

#On-device

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 9

TodayOct 9Fri
  1. OpenBMBOfficialAI score28

    MiniCPM5-2B runs at 37 tok/s on iPhone Air

    AIOpenBMB reports that its MiniCPM5-2B model runs at 37 tokens per second on an iPhone Air. The post presents this as evidence that small open multimodal models can run on mobile devices without a cloud GPU, with NobodyWho noting the model is available in its Chat app.

Oct 8

Oct 8Thu
  1. SantiagoXAI score34

    Atomic Agent Desktop launches with Linux support on day one

    AIAtomic Agent Desktop is a local AI-first agent available for Mac, Windows, and Linux. It offers a 4x larger context window on local models using TurboQuant, and Atomic Fusion orchestrates cloud and local models to reduce costs.

  2. 🚨 AI News | TestingCatalogXAI score62

    Atomic Agent Desktop, an open-source local AI agent app, is now available

    AIAtomic Agent Desktop is a free open-source app for macOS, Windows, and Linux that runs open models like Qwen and Gemma locally without an account. It connects to a cloud model only when selected, and its Fusion feature lets a cloud model plan a task while up to 8 local agents carry it out. The post's own text adds a setup wizard that checks RAM and suggests suitable models, and import from Claude Code, Codex, Hermes, and OpenClaw.

    Video from @testingcatalog's post
  3. Zhihao JiaXAI score62

    Lithos AI open-sources lithos-metal for fast local inference on Apple M5 Max

    AILithos AI says it is open-sourcing lithos-metal, which uses megakernels and DSpark speculative decoding. The post claims Qwen3.8-27B reaches a peak of over 200 tokens per second per user on a single Apple M5 Max. It says users can try the tool with any coding agent in one command, and links to the code on GitHub and a technical blog.

    Video from @JiaZhihao's post

Oct 7

Oct 7Wed
  1. Teknium 🪽XAI score36

    Community brings Hermes Gadget SDK to LilyGO, AIPI Lite, and old Android phones

    AIDevelopers are running Hermes on devices such as LilyGO watches, AIPI Lite, desk gadgets, and old Android phones after the Hermes Gadget open SDK and demo were released three days ago. The post credits @NousResearch and says more boards are landing on main through contributor PRs, with the SDK available on GitHub.

Oct 6

Oct 6Tue
  1. Kilo (acq. by Anaconda)OfficialAI score29

    Kilo launches Kilo Desktop, a unified app for 500+ AI models

    AIKilo has launched Kilo Desktop, a single app offering access to more than 500 models from major labs, including open-source and local models. It includes agents that plan, code, and debug alongside users, plus built-in notebooks, local model support, and conda environments.

    Image from @kilocode's post
  2. IThome · AINewsAI score41

    Strata engine runs 125B Qwen3.8 model on 12GB GPU at 94 tokens/s

    AIDeveloper Niko1221 has open-sourced Strata, an engine that runs a quantized 125B-parameter Qwen3.8-Flash-Next model on consumer GPUs with at least 12GB of VRAM. Strata loads the MoE model into RAM and keeps only frequently used experts in VRAM, and uses a lightweight model for speculative decoding. On an NVIDIA RTX 5070 with 12GB VRAM, the Q2_0 quantization reaches 94 tokens per second for output.

Oct 4

Oct 4Sun
  1. Teknium 🪽XAI score29

    Teknium says ESP32 hardware is now set up at home

    AITeknium announced that an ESP32 is now running at home, with no further technical details given in the post. The post is a brief update that quotes an @adolandev post about Hermes Gadget, an open SDK for a small device that speaks to a user's own Hermes model and can be tested with a desktop simulator.

Oct 1

Oct 1Thu
  1. Praan IncXAI score30

    Praan launches new HIVE air purifier, smaller and 30% more affordable

    AIPraan launched a new HIVE, its medical-grade indoor air purifier, priced at ₹49,999 (all-inclusive, $520). The new model is 62mm smaller, more capable, and autonomous, and is 30% more affordable than its predecessor. The company says the HIVE is now in more than 1,500 locations across 8 countries.

Sep 30

Sep 30Wed
  1. Hacker News · Launch HN, YC launches (10+ points)BlogAI score62

    Magnitude launches an open source inference engine that tunes kernels to local hardware

    AIMagnitude is an open source inference engine for agents that compiles and tunes its kernels on the user's device before running a model. The source claims up to 2x faster decoding than llama.cpp, citing 92% faster decode on Metal and 19% on CUDA, and says one click connects agents such as Pi, OpenCode, Codex, and Claude Code. It supports macOS, Windows, and Linux, and the source states that prompts and models stay on the user's machine.

  2. ollamaOfficialAI score30

    Ollama adds local decision models like Nimble via new API

    AIOllama now supports decision models such as Nimble locally, usable for tasks like ticket triaging, model routing, and content moderation. The model can be installed with ollama pull nimble and accessed through the new local /v1/systemone API, as shown in a real-time Ollama racer demo.

    Video from @ollama's post

Sep 28

Sep 28Mon
  1. Daniel HanXAI score29

    Unsloth Desktop serves local Laya decision models for real-time packing demo

    AIUnsloth Desktop can now serve local Laya decision models through a Jev-compatible API, shown in a real-time packing demo where suitcase items update as the user types. The demo runs through Unsloth's Decision API, and the team says more optimizations are coming to speed up local hardware performance.

    Video from @danielhanchen's post

Sep 24

Sep 24Thu
  1. OpenBMBOfficialAI score34

    FIT-GGUF enables size-targeted mixed-precision quantization of MiniCPM5-2B

    AIDeveloper @Scorp1o_117 used FIT-GGUF to build four MiniCPM5-2B GGUF variants, ranging from about 1.14 GiB to 1.46 GiB, tuned to target file sizes or fidelity tiers. Instead of fixed presets, FIT-GGUF allocates precision tensor by tensor, with Quality, Balanced, Compact, and Mini options, and its generated files matched predicted sizes. Builds are evaluated with KL Divergence and Same-top metrics and are available on Hugging Face.

    Image from @OpenBMB's post

Sep 22

Sep 22Tue
  1. Unsloth AIOfficialAI score26

    Qwen-Image-2.1 FP8 and GGUF quants now run in Unsloth Desktop

    AIUnsloth announced that Qwen-Image-2.1 FP8 and GGUF quantized versions should now run properly in Unsloth Desktop. The app supports both image generation and image editing with these quants. Further details are available on the Unsloth GitHub repository.

    Image from @UnslothAI's post
  2. OpenBMBOfficialAI score59

    VoxWeft runs real-time interpretation locally on Apple Silicon using VoxCPM2

    AIOpenBMB highlights VoxWeft, an open-source simultaneous interpretation system for Apple Silicon built by developer @HenryZ30734018 on an MLX implementation of VoxCPM2. The system turns live speech into translated speech on-device, with first audio streaming in about 170 ms on an M5 MacBook. VoxCPM2 generates speech in 30 languages, supports direct language-pair interpretation, and clones a target voice from about 5 seconds of reference audio.

    Video from @OpenBMB's post

Sep 21

Sep 21Mon

Sep 20

Sep 20Sun
  1. OpenBMBOfficialAI score22

    MiniCPM5-2B on a local Mac correctly reconciles naproxen medication history

    AIOpenMed reports that MiniCPM5-2B, run on a local Mac, correctly kept naproxen in medication history rather than the current-medication export after a newer note said it was stopped. Every graph connection in the run links back to its source, using fictional clinical notes.

  2. OpenBMBOfficialAI score35

    OpenBMB's Augury model improves on-device plant ID for farmers

    AIA developer's Augury plant identification model, built on an OpenBMB model, raised photo top-1 accuracy from 71.8% to 80.2% by merging duplicate species keys and adding PCA whitening. The next steps are reaching 90%+ accuracy and building a phone GUI so farmers can use it on-device.

Sep 18

Sep 18Fri
  1. LM StudioOfficialAI score20

    LM Studio adds Splash engine for running incoai models on Mac

    AILM Studio users can enable the Splash engine under Settings > Runtime > Experimental backends and download supported models by searching "incoai." The post recommends an M3 or newer Mac running macOS 26.4 with 36GB+ RAM.

    Image from @lmstudio's post
  2. LM StudioOfficialAI score62

    LM Studio adds Qwen3.8-27B running at up to 144 tok/sec on M5 Max

    AILM Studio announced that Qwen3.8-27B runs at up to 144 tokens per second on an M5 Max MacBook Pro through its partnership with Inco Splash. The post claims up to 3× the decode speed of Ollama, 2× oMLX, and almost 4× when an agent fans out into sub-agents. Inco Splash is described as an open-source inference engine built for the model and Apple silicon, available through the linked LM Studio blog.

  3. LMSYS OrgOfficialAI score52

    LMSYS blog shows DeepSeek-V4-Flash and Kimi-K3 running on consumer hardware via SSD Expert Pack

    AILMSYS Org announced a blog on running DeepSeek-V4-Flash and Kimi-K3 on consumer hardware using SSD Expert Pack, built by WiCi AI and the SGLang team. Routed experts stay on an NVMe SSD, and the runtime loads only router-selected experts into a GPU cache. On one RTX 5090, 32 GB RAM, and a 2 TB SSD, DeepSeek-V4-Flash MXFP4 decoded at 1.85–1.99 tokens/sec and Kimi-K3 community Q2_K (text-only) at about 0.29 tokens/sec.

    Image from @lmsysorg's post

Sep 17

Sep 17Thu
  1. OpenBMBOfficialAI score36

    OpenBMB's MiniCPM5-2B runs offline on-device with 128K context

    AIOpenBMB's MiniCPM5-2B is a 2.5B-parameter model with native 128K context, offering hybrid Think and No-Think modes in one checkpoint. Users can download it from Hugging Face and run it fully offline on-device, as RunAnywhere demonstrated. In a demo, the model first called a puzzle impossible, then corrected itself and wrote a working verifier.

  2. OpenBMBOfficialAI score40

    OpenMed and MiniCPM5-2B demo local agentic clinical AI workflow

    AIOpenMed paired with MiniCPM5-2B to demonstrate a local clinical AI workflow combining privacy-preserving data processing with a compact model's tool use and long-context reasoning. OpenMed masks sensitive identifiers and extracts clinical context before MiniCPM5-2B calls tools, compares lab results, and generates clinical handoffs with source references. The post presents this as an example of keeping inference on local, resource-constrained hardware.

    Image from @OpenBMB's post

Sep 15

Sep 15Tue