Skip to contentSkip to stories

Updated

All AI news

Items with an AI score under 20 are hidden. Show low-relevance items

Aug 12

Aug 12Wed
  1. Michael TruellAI score62

    Grok 4.6 is released with gains on agentic and knowledge-work benchmarks

    AIGrok 4.6 is released as a significant improvement over Grok 4.5 at the same price, according to the announcement. The author says it is significantly better at difficult tasks and knowledge work, combining Opus-class intelligence and polish with low cost and high speed. A comparison table shows Grok 4.6 High scoring 61 on the AA Intelligence Index, versus 56 for Grok 4.5 High, and 1753 on GDPval-AA v2, versus 1526.

Aug 11

Aug 11Tue
  1. Liquid AI BlogAI score62

    Liquid AI releases LFM2.5-VL-3B, a 3B vision-language model for edge devices

    AILiquid AI released LFM2.5-VL-3B, an open-weight 3B vision-language model that it says rivals models twice its size while running faster on CPU and GPU. Benchmarks show large gains over LFM2-VL-3B, including ScreenSpot-v2 averaging 80.7, RefCOCO precision@1 rising from 57.1 to 87.9, and ToolSandbox rising from 26.4 to 59.5. The model is available on Hugging Face and decodes 228 tokens/s on an Apple M5 Max.

    Why it matters: The post pairs benchmark gains with on-device and GPU throughput figures, showing how a 3B vision model trades size against speed and accuracy.

  2. Liquid AI · new models on Hugging FaceAI score40

    LiquidAI releases LFM2.5-VL-3B, a 3B multimodal model for on-device use

    AILiquidAI has released LFM2.5-VL-3B, a 3B-parameter multimodal model that processes text and images and is built on the LFM2.5-2.6B language model with a SigLIP2 NaFlex vision encoder. It runs at 228 tokens/s on an Apple M5 Max and 116 tokens/s on an AMD Ryzen AI Max+ 395 in under 3.3 GB of memory, with a 32,768-token context length. The model is available in native, GGUF, ONNX and MLX formats on Hugging Face.

  3. Bryan CatanzaroAI score40

    Nemotron 3.5 Lightning: NVIDIA's fast 30B MoE model for agents

    AINVIDIA's Bryan Catanzaro says Nemotron 3.5 Lightning uses the same architecture as Nemotron 3.0 Nano, adds speculative decoding, and matches the intelligence of Nemotron 3.0 Super. NVIDIA describes it as an open 30B MoE model with 3B active parameters, built for always-on agents handling high-volume, specialized tasks, with up to 4x the output speed of similar-sized models.

Aug 10

Aug 10Mon
  1. Liquid AI · new models on Hugging FaceAI score43

    LiquidAI LFM2.5-2.6B-DSpark Speeds Up LFM2.5 Decoding With Speculative Drafting

    AILiquid AI released LFM2.5-2.6B-DSpark, a 327.7M-parameter speculative-decoding draft model for its LFM2.5-2.6B target, on Hugging Face. In SGLang on a single H100 with batch size 1, mean decoding throughput rises from 323 to 864 tokens per second, about 2.67x, and on an Apple M4 Max via Metal it rises from 61 to 139 tokens per second, about 2.27x. Because the target verifies every proposed token, the output matches what LFM2.5-2.6B would generate alone.

  2. Liquid AI · new models on Hugging FaceAI score38

    Liquid AI releases LFM2.5-8B-A1B-DSpark draft model for faster LFM2.5 decoding

    AILiquid AI released LFM2.5-8B-A1B-DSpark, a 327.7M-parameter speculative-decoding draft model for its LFM2.5-8B-A1B target. In SGLang on one H100 with batch size 1, mean accepted tokens per step reached 7.21 across five benchmarks, and decoding ran about 2.6× faster. The model also runs on Apple silicon through the Metal backend, with a 1.18× mean speedup on an M4 Max.

  3. Liquid AI · new models on Hugging FaceAI score42

    LiquidAI LFM2.5-1.2B-Instruct-DSpark Drafter Speeds Up Decoding About 2x

    AILiquid AI released LFM2.5-1.2B-Instruct-DSpark, a 295.7M-parameter speculative-decoding draft model for the LFM2.5-1.2B-Instruct target on Hugging Face. On an H100 it averages 4.81 accepted tokens per step and runs about 2.10x faster across benchmarks, with about 2x speedup in SGLang and on-device Apple silicon support via Metal.

  4. Cohere · new models on Hugging FaceAI score46

    Cohere releases North Micro Vision Instruct, a 2.4B open-weight vision-language model

    AICohere has released North Micro Vision Instruct, a 2.4B-parameter open-weight vision-language model under the Apache 2.0 license, on Hugging Face. The model processes images at native resolution and handles visual question answering, captioning, grounding, OCR, and document understanding across English, German, French, Spanish, Italian, Portuguese, Hindi, Japanese, Korean, Chinese, and Arabic. It has a 128K-token language backbone context window, but its validated multimodal range is up to 8K tokens.

Aug 7

Aug 7Fri
  1. Qwen · new models on Hugging FaceAI score88

    Qwen releases open-weight Qwen3.8-2.4T-A95B, a 2.4T-parameter MoE model

    AIQwen has released the Qwen3.8-2.4T-A95B model weights on Hugging Face, with 2.4T total and 95B activated parameters in a mixture-of-experts design. The release supports reasoning_effort levels and a 262,144-token native context extensible to 1,010,000 tokens, and it is text-only with thinking mode always on. The source reports benchmark results against Opus 4.8, Fable 5, GPT 5.6 Sol, and Qwen3.7-Max, and says the official Qwen3.8-Max API adds vision input and a 1M default context.

    Why it matters: The model card gives parameters, architecture, reasoning controls, and benchmark tables against named rival models, showing what an open release of this scale actually offers.

  2. MiniMax · new models on Hugging FaceAI score44

    MiniMax Music 3 generates five-minute songs with coherent structure and vocals

    AIMiniMax Music 3 is a music generation model that creates complete songs up to five minutes long from lyrics and a music description. It pairs an 8B Global LLM for long-range structure with a 0.6B Local LLM for acoustic detail, outputting 32 kHz, 16-bit stereo WAV audio. The model is available on Hugging Face and supports SGLang-Omni, diffusers, and ComfyUI.

Aug 6

Aug 6Thu
  1. Intern Large ModelsAI score62

    Shanghai AI Lab open-sources Mobius, a Transformer alternative claiming 4x faster reasoning

    AIShanghai AI Lab open-sourced Mobius, an architecture its authors compare to the RNN-to-Transformer shift in both token and knowledge dimensions. Against Transformers, the post claims about 4x faster reasoning, the same MMLU score with 40% less data, and 2x better compositional generalization. Mobius is supported by XTuner, LMDeploy, vLLM, and SGLang, and its experimental setup and training pipeline will be released later.

    Image from @intern_lm's post
  2. InternLM (Shanghai AI Lab) · new models on Hugging FaceAI score38

    Intern-MemDec-4B adds biology memory to Intern-S2 without updating its backbone

    AIShanghai AI Lab's InternLM released Intern-MemDec-4B, a 4B-parameter memory decoder that runs alongside an Intern-S2 backbone and a token-level router to add biology knowledge. On all 21 Biology-Instructions tasks, the average score rose from 56.92 to 60.32 when paired with Intern-S2-Preview-397B. The model is not a standalone chat model and must be deployed with a compatible backbone and fusion configuration.

Aug 5

Aug 5Wed
  1. Qwen · new models on Hugging FaceAI score79

    Qwen3.8-27B releases dense vision-language model with thinking controls

    AIAlibaba's Qwen team has released Qwen3.8-27B on Hugging Face as a 27B dense model with native image and video understanding. The model card reports gains over Qwen3.6-27B on coding and agent benchmarks, including SWE-bench Pro at 61.7 versus 53.5. It adds reasoning_effort levels and preserve_thinking, and its hosted Qwen Cloud version is described as coming soon.

    Why it matters: The model card gives per-benchmark comparisons with Qwen3.6-27B and named rivals, plus reasoning_effort and preserve_thinking controls for judging cost and agent behavior.

Aug 3

Aug 3Mon
  1. Liquid AI BlogAI score72

    Liquid AI releases LFM2.5-2.6B, a 2.6B on-device agentic model

    AILiquid AI released LFM2.5-2.6B, a 2.6B-parameter agentic model that runs on-device on phones and CPUs, along with a base variant on Hugging Face. The company reports it leads on every instruction-following benchmark and nearly every tool-use benchmark it tested, and decodes 220 tokens/s on an M5 Max. The source says larger models may still suit complex agentic or coding-heavy tasks.

    Why it matters: The source reports benchmark results against several same-tier models and notes where larger models still lead, which helps judge fit for edge agent workloads.

Jul 31

Jul 31Fri
  1. DeepSeek · new models on Hugging FaceAI score75

    DeepSeek releases DeepSeek-V4-Flash-0731 with stronger agentic capabilities

    AIDeepSeek has released DeepSeek-V4-Flash-0731 as the official version superseding the preview, with substantially enhanced agentic capabilities. The source reports it outperforms DeepSeek-V4-Pro (Preview) on listed benchmarks, including Terminal Bench 2.1 at 82.7 versus 72.1, despite a far smaller activated parameter count. The model ships under the MIT License with DSpark speculative decoding supported in vLLM and SGLang.

    Why it matters: The release shows benchmark gains over the preview and a concrete vLLM and SGLang serving path, useful for teams weighing a self-hosted agentic coding model.

  2. DeepSeek API NewsAI score67

    DeepSeek-V4-Flash API enters public beta with stronger agent benchmarks

    AIDeepSeek has released the DeepSeek-V4-Flash API in public beta, and developers can use the latest version by setting the model name to deepseek-v4-flash. The source reports agent benchmark results far above V4-Pro-Preview, including 82.7 on Terminal Bench 2.1 and 70.3 on Toolathlon verified. V4-Flash natively supports the Responses API format and is adapted for Codex, while V4-Pro and the APP/WEB models are unchanged.

    Why it matters: The release lists agent benchmark results against V4-Pro-Preview and notes Responses API support for Codex, which helps developers gauge the upgrade's practical effect on their workflows.

Jul 30

Jul 30Thu
  1. MiniMax BlogAI score72

    MiniMax H3 unifies text, image, video, and audio generation in one model

    AIMiniMax launches H3, a general-purpose multimodal generation model that understands text, images, video, and audio as unified context. It generates video up to 15 seconds at 2K resolution with native stereo sound, and the company says model weights will be opened in the coming days, subject to applicable laws and regulations. MiniMax also says H3 is priced below mainstream models at 2K and 768p.

    Why it matters: The post explains how a unified multimodal design and training choices enable 2K video with native stereo sound, useful for comparing against closed video generators.

  2. Soumith ChintalaAI score57

    Thinking Machines releases Inkling-Small, a 276B-parameter model with full weights

    AIThinking Machines is releasing Inkling-Small, which it says achieves performance comparable to Inkling at a quarter of its size. The model has 276B total parameters with 12B active, and the full weights are available. Users can fine-tune it on Tinker or chat with it in text, image, and audio on Tinker Playground.

Jul 29

Jul 29Wed
  1. Air Street PressAI score75

    Poolside's Laguna S 2.1 is an open agentic coding model that runs on one DGX Spark

    AIPoolside released Laguna S 2.1, an open-weights agentic coding model with 118 billion total parameters and about 8 billion active per token, supporting up to a million tokens of context. Quantized, it fits on one NVIDIA DGX Spark, and Poolside reports 70.2% on Terminal-Bench 2.1 with thinking enabled, with its evaluation trajectories published online. The same week it shipped the Poolside Desktop Assistant for macOS, which runs Laguna locally or alongside Claude Code, Codex, and Gemini agents.

  2. Alibaba NLP (Tongyi) · new models on Hugging FaceAI score40

    Alibaba NLP releases UEmbed-9B, a unified sparse and dense multimodal embedding model

    AIAlibaba NLP has released UEmbed-9B, a decoder-only multimodal embedding model built on Qwen3.5 9B that outputs both dense and SPLADE-style sparse embeddings from one forward pass. It supports text, image, video, and mixed-modal inputs for retrieval and multimodal search, and the family also includes 2B and 4B variants. The model is available on Hugging Face, with transformers and vLLM inference support.

  3. Alibaba NLP (Tongyi) · new models on Hugging FaceAI score38

    Alibaba NLP releases UEmbed-4B, a unified sparse and dense multimodal embedding model

    AIAlibaba NLP has released UEmbed-4B, a decoder-only multimodal embedding model built on Qwen3.5 4B that outputs both dense and sparse embeddings from one forward pass. It handles text, image, video, and mixed-modal inputs for retrieval and visual-document search, and sparse activations map to vocabulary terms usable with inverted indexes. The model is available on Hugging Face in a family that also includes 2B and 9B variants.

  4. Alibaba NLP (Tongyi) · new models on Hugging FaceAI score43

    Alibaba-NLP releases UEmbed-2B, a multimodal model producing dense and sparse embeddings

    AIAlibaba-NLP's UEmbed-2B, a decoder-only multimodal embedding model built on Qwen3.5 2B, produces both dense and SPLADE-style sparse embeddings from a single forward pass. It supports text, image, video, and mixed-modal inputs for retrieval, and the 4B and 9B variants are also available. The team reports state-of-the-art results on the text and agent tracks of MMEB-v3.

Jul 28

Jul 28Tue
  1. MiniMax · new models on Hugging FaceAI score76

    MiniMax H3 releases open-weight omni-modal video model with native stereo audio

    AIMiniMax released H3, an open-weights omni-modal model that generates video with native stereo audio up to 2K and 15 seconds. The system combines H3-Context-IR preprocessing, the H3-Base generator at 768p, and H3-Regenerate-2K for 2K output, with the Context-IR and 2K modules available only through API.

    Why it matters: The source details a three-module pipeline and open weights with deployment paths, showing how a video model is served and reproduced locally.

Jul 27

Jul 27Mon
  1. Liquid AI BlogAI score49

    Liquid AI Releases LFM2.5-Encoders for Fast Long-Context Encoding on CPU

    AILiquid AI released LFM2.5-Encoder-230M and LFM2.5-Encoder-350M, bidirectional encoders built on the LFM2 hybrid architecture and available on Hugging Face. They support an 8,192-token context and are designed for fine-tuning on classification and token-level tasks. On CPU, LFM2.5-Encoder-230M is the fastest model tested from 1K tokens up, running about 3.7x faster than ModernBERT-base at 8,192 tokens.

  2. Kimi.aiAI score86

    Moonshot AI releases Kimi K3 weights and technical report

    AIMoonshot AI is releasing the model weights and technical report for Kimi K3, a 2.8T-parameter MoE model with native visual understanding and a 1M-token context window. The post says the new architecture delivers 2.5x the intelligence per unit of compute, and the company is also opening high-performance attention kernels, an MoE communication library, and infrastructure for running agent environments at scale.

    Why it matters: The source names the model size, context window, and released weights, which helps readers compare its scale and openness with other frontier releases.

    Image from @Kimi_Moonshot's post
  3. Air Street PressAI score72

    Black Forest Labs releases FLUX 3, extended to video and robot control

    AIBlack Forest Labs released FLUX 3, a multimodal model trained on images, video, and audio, and mimic built FLUX-mimic on its video backbone to control robots. In a soft-body kitting task, mimic reports a 95% success rate without single-task fine-tuning, compared with 55% for an adapted π0.5 model. FLUX 3 Video is in early access, with action prediction offered to selected partners and an open-weight backbone planned.

Jul 24

Jul 24Fri