Skip to contentSkip to stories
Updated

#On-device

Oct 8

Oct 8Thu
  1. 🚨 AI News | TestingCatalogXAI score62

    Atomic Agent Desktop, an open-source local AI agent app, is now available

    AIAtomic Agent Desktop is a free open-source app for macOS, Windows, and Linux that runs open models like Qwen and Gemma locally without an account. It connects to a cloud model only when selected, and its Fusion feature lets a cloud model plan a task while up to 8 local agents carry it out. The post's own text adds a setup wizard that checks RAM and suggests suitable models, and import from Claude Code, Codex, Hermes, and OpenClaw.

    Video from @testingcatalog's post
  2. Zhihao JiaXAI score62

    Lithos AI open-sources lithos-metal for fast local inference on Apple M5 Max

    AILithos AI says it is open-sourcing lithos-metal, which uses megakernels and DSpark speculative decoding. The post claims Qwen3.8-27B reaches a peak of over 200 tokens per second per user on a single Apple M5 Max. It says users can try the tool with any coding agent in one command, and links to the code on GitHub and a technical blog.

    Video from @JiaZhihao's post

Oct 6

Oct 6Tue
  1. IThome · AINewsAI score41

    Strata engine runs 125B Qwen3.8 model on 12GB GPU at 94 tokens/s

    AIDeveloper Niko1221 has open-sourced Strata, an engine that runs a quantized 125B-parameter Qwen3.8-Flash-Next model on consumer GPUs with at least 12GB of VRAM. Strata loads the MoE model into RAM and keeps only frequently used experts in VRAM, and uses a lightweight model for speculative decoding. On an NVIDIA RTX 5070 with 12GB VRAM, the Q2_0 quantization reaches 94 tokens per second for output.

Sep 30

Sep 30Wed
  1. Hacker News · Launch HN, YC launches (10+ points)BlogAI score62

    Magnitude launches an open source inference engine that tunes kernels to local hardware

    AIMagnitude is an open source inference engine for agents that compiles and tunes its kernels on the user's device before running a model. The source claims up to 2x faster decoding than llama.cpp, citing 92% faster decode on Metal and 19% on CUDA, and says one click connects agents such as Pi, OpenCode, Codex, and Claude Code. It supports macOS, Windows, and Linux, and the source states that prompts and models stay on the user's machine.

Sep 22

Sep 22Tue
  1. OpenBMBOfficialAI score59

    VoxWeft runs real-time interpretation locally on Apple Silicon using VoxCPM2

    AIOpenBMB highlights VoxWeft, an open-source simultaneous interpretation system for Apple Silicon built by developer @HenryZ30734018 on an MLX implementation of VoxCPM2. The system turns live speech into translated speech on-device, with first audio streaming in about 170 ms on an M5 MacBook. VoxCPM2 generates speech in 30 languages, supports direct language-pair interpretation, and clones a target voice from about 5 seconds of reference audio.

    Video from @OpenBMB's post

Sep 18

Sep 18Fri
  1. LMSYS OrgOfficialAI score52

    LMSYS blog shows DeepSeek-V4-Flash and Kimi-K3 running on consumer hardware via SSD Expert Pack

    AILMSYS Org announced a blog on running DeepSeek-V4-Flash and Kimi-K3 on consumer hardware using SSD Expert Pack, built by WiCi AI and the SGLang team. Routed experts stay on an NVMe SSD, and the runtime loads only router-selected experts into a GPU cache. On one RTX 5090, 32 GB RAM, and a 2 TB SSD, DeepSeek-V4-Flash MXFP4 decoded at 1.85–1.99 tokens/sec and Kimi-K3 community Q2_K (text-only) at about 0.29 tokens/sec.

    Image from @lmsysorg's post

Sep 17

Sep 17Thu
  1. OpenBMBOfficialAI score40

    OpenMed and MiniCPM5-2B demo local agentic clinical AI workflow

    AIOpenMed paired with MiniCPM5-2B to demonstrate a local clinical AI workflow combining privacy-preserving data processing with a compact model's tool use and long-context reasoning. OpenMed masks sensitive identifiers and extracts clinical context before MiniCPM5-2B calls tools, compares lab results, and generates clinical handoffs with source references. The post presents this as an example of keeping inference on local, resource-constrained hardware.

    Image from @OpenBMB's post
That’s everything