Skip to contentSkip to stories

Updated

#On-device

Oct 9

TodayOct 9Fri1 item
  1. ModelScopeAI score63

    Google releases EmbeddingGemma 2, a lightweight multimodal embedding model for on-device search

    AIGoogle released EmbeddingGemma 2, a 740M-parameter multimodal embedding model under Apache 2.0 for private, on-device search and retrieval. It maps text, code, images, video, and audio into one shared space and reports a 9.92-point gain over EmbeddingGemma 1 on MTEB Code. The post lists about 191MB active RAM for quantized text-only weights and about 567MB for the full multimodal model on a Pixel 11 Pro.

Oct 8

Oct 8Thu
  1. TechCrunch · AIAI score38

    Google launches Google AI Edge Foresight, a local-first Mac meeting note-taker rival to Granola

    AIGoogle released Google AI Edge Foresight, a Mac app that captures meeting notes offline using the on-device EmbeddingGemma 2 model with 740 million parameters. The app offers split-screen shorthand and AI-generated notes, transcripts, and a Gemma 4-powered assistant that can answer questions from uploaded documents. Google's FAQ says it is optimized for Apple Silicon.

  2. Microsoft CopilotAI score29

    Copilot on Windows gains Hybrid Intelligence, blending cloud and local agents

    AIMicrosoft Copilot announced Hybrid Intelligence for Windows, which balances cloud and locally run agents to handle tasks like file wrangling and workflows on the PC. According to Satya Nadella, Copilot can use context on the PC with user permission and draw on local models when appropriate. The feature is described as coming soon.

  3. Testing CatalogAI score62

    Atomic Agent Desktop, an open-source local AI agent app, is now available

    AIAtomic Agent Desktop is a free open-source app for macOS, Windows, and Linux that runs open models like Qwen and Gemma locally without an account. It connects to a cloud model only when selected, and its Fusion feature lets a cloud model plan a task while up to 8 local agents carry it out. The post's own text adds a setup wizard that checks RAM and suggests suitable models, and import from Claude Code, Codex, Hermes, and OpenClaw.

  4. Comfy BlogAI score34

    How I Generated Live Video with MiniMax H3 on a Single GPU

    AIA ComfyUI developer generated 15-second 448×256 video in 15 seconds or less on one RTX 5090 using MiniMax H3 with FastVideo's FastH3 V2 checkpoint in four sampling steps. The setup combined sparse attention, a smaller ClipProj text encoder, a pruned INT8 checkpoint, and a fused FP4 MLP, cutting VRAM needs from 80GB to under 30GB. The custom ComfyUI node is open source.

  5. SiliconANGLE · AIAI score30

    Liquid AI Builds On-Device Personal AI Around Device-Level Context

    AILiquid AI is building personal AI that runs on devices such as phones, wearables, PCs, and cars, using its Liquid Context layer, which is optimized for Snapdragon processors, to sit between models, agents, and hardware. The company's agent harness uses its own models to decide which user context to retain and how to compress it within fixed compute limits. Liquid AI is also collaborating with Mercedes-Benz Group AG to bring on-device AI to its cars and plans observability and continuous improvement loops for self-improving agents.

  6. The Verge · AIAI score52

    Google's experimental AI Edge Foresight transcribes meetings fully offline on Mac

    AIGoogle has released AI Edge Foresight, a free experimental note-taking app that transcribes meetings and audio files entirely offline on macOS. It runs on the on-device EmbeddingGemma 2 model and turns shorthand notes into polished notes based on the transcript. Google says files, meeting audio, and notes never leave the computer, and the app is currently optimized only for Macs with Apple Silicon.

  7. IThome · AIAI score40

    Microsoft Confirms Copilot+ PC Brand Lives On, Runs 2 Trillion Local AI Inferences Monthly

    AIMicrosoft Windows and devices head Pavan Davuluri confirmed the Copilot+ PC brand has not been discontinued, saying more than 40% of commercial laptops are Copilot+ PCs shipping in tens of millions annually. He said these devices run over 2 trillion local inferences per month across search, image processing, and video calls. Microsoft plans to strengthen them through hybrid intelligence with local context, local actions, and local models.

Oct 7

Oct 7Wed
  1. Google Developers BlogAI score62

    Google open-sources ML Drift, a cross-platform GPU engine for on-device AI

    AIGoogle's AI Edge Team open-sourced ML Drift under Apache 2.0, a GPU compute engine for on-device AI inference across OpenGL ES, OpenCL, Metal, and WebGPU. It serves as the core GPU acceleration engine within LiteRT and succeeds the legacy TFLite GPU delegate, which will no longer receive new features. The post cites benchmarks showing up to 40% lower frame latency in YouTube Shorts and up to 30% faster on-device performance in Adobe Lightroom and Photoshop.

    Why it matters: The post explains how ML Drift unifies GPU shaders across platforms and replaces the TFLite GPU delegate, which matters for developers deploying on-device models.

  2. NVIDIA BlogAI score67

    NVIDIA and Microsoft Launch RTX Spark Laptops and DGX Station for Windows AI Agents

    AINVIDIA and Microsoft announced RTX Spark laptops and compact desktops that run the full NVIDIA AI stack locally, with laptop preorders open today and sales from October 16. Microsoft also announced general availability of Microsoft Execution Containers (MXC), an OS-level infrastructure for agents to run securely in the background, while NVIDIA previewed DGX Station for Windows with 748GB of coherent memory and up to 20 petaFLOPS of FP4 compute.

    Why it matters: The announcement pairs Windows agent infrastructure with local hardware, showing how agents may move onto personal computers and enterprise desktops rather than only cloud services.

  3. Satya NadellaAI score72

    Windows adds on-device agents, local coding models, and Hybrid Intelligence

    AIMicrosoft says Windows will bring unmetered intelligence to PCs, letting agents work securely on-device. The post lists MAI-Code-1.1 Flash, a 137B parameter coding model with a 256K context window optimized to run on PCs, and GitHub Copilot handoffs to local models. It also describes Hybrid Intelligence, which lets Copilot act on the PC and keep sensitive work local, and Code in Copilot for building software without cloud token spend, on devices such as Surface Laptop Ultra powered by NVIDIA RTX Spark.

  4. Liquid AIAI score36

    Liquid AI releases d1-omni-600M, a 600M multimodal model for on-device tasks.

    AILiquid AI has released d1-omni-600M, an experimental 600M-parameter model that handles text plus image or audio input. It combines LFM2.5-Encoder-350M with vision and audio encoders and leads the company's text benchmark comparison on toxicity detection and paraphrase identification. The post suggests uses such as voice-command routing, on-device moderation, and intent classification.

  5. Hugging Face BlogAI score49

    Liquid AI Releases Open d1-3B and d1-omni-600M Edge Decision Models

    AILiquid AI released two open-weight decision models, d1-3B and d1-omni-600M (experimental), built on its Liquid Foundation Models and available on Hugging Face. d1-3B scores 48.57 on the Decision Index 0.2.1, the highest among decision models under 10B parameters, and answers a question in 16 ms on an NVIDIA Jetson AGX Thor and under 50 ms on a Jetson Orin Nano. The models support text and images (d1-3B) or text with image or audio (d1-omni-600M).

  6. Aravind SrinivasAI score62

    Perplexity open-sources pplx-embed-v2-late multimodal embedding models

    AIPerplexity is open-sourcing pplx-embed-v2-late, multi-vector embedding models for text and images in one shared space, in 9B and 0.6B sizes. The 9B model can index multimodal data, the 0.6B model can run queries on device, and PDF pages can be searched without OCR. The author reports 92.4% on MADQA and 64% on BrowseComp+, with weights available on Hugging Face.

  7. Testing CatalogAI score41

    Google releases Foresight macOS app using Gemma 4 for voice notes

    AIGoogle released the Google AI Edge Foresight app for macOS, powered by Gemma 4 and EmbeddingGemma 2. The app can connect to Google Drive to build a knowledge graph and, when transcription is active, uses local Gemma 4 E4B or Gemma 4 12B models to transcribe voice notes into new documents. EmbeddingGemma 2 is an open-weight, Apache 2 licensed 740M-parameter multimodal embedding model with an 8K context window.

Oct 6

Oct 6Tue
  1. meng shaoAI score62

    Google DeepMind releases EmbeddingGemma 2, an open multimodal embedding model for on-device use

    AIGoogle DeepMind released EmbeddingGemma 2, an open 740M-parameter embedding model that maps text, code, images, video, and audio into one 768-dimensional space. Text-only use needs a 270M-parameter footprint, about 191MB active RAM when quantized on a Pixel 11 Pro, while loading all modalities takes about 567MB. The reported MTEB Code NDCG@10 score is 78.68, about 14% above the first generation, and MTEB Multilingual v2 is 61.36, roughly flat.

  2. Liquid AI BlogAI score62

    Liquid AI releases open d1-3B and d1-omni-600M decision models for edge devices

    AILiquid AI released two open-weight d1 decision models, d1-3B and d1-omni-600M, on Hugging Face. d1-3B scores 48.57 on the Decision Index v0.2.1 public split and answers a single question in 8 ms on an NVIDIA GeForce RTX 4090 and 50 ms on a Jetson Orin Nano. d1-omni-600M is an experimental checkpoint that handles text with images or audio and scores 15.95 on the same index.

    Why it matters: The release pairs open-weight decision models with measured latency across Apple, NVIDIA, and Jetson hardware, showing how edge deployment changes what is practical.

  3. OllamaAI score55

    Google DeepMind's EmbeddingGemma 2 is now available on Ollama

    AIOllama announced that Google DeepMind's EmbeddingGemma 2 is now available on Ollama. The author describes it as made for consumer devices and multimodal, and gives the command ollama pull embeddinggemma-2 to download it. The quoted DeepMind post says the model is a natively multimodal open model for on-device embeddings that unifies code, images, audio, and video in a shared space.