Skip to content

#Model release

Aug 9

Aug 9Sun
  1. Fireworks AI BlogAI score60

    Meta releases Muse Glimmer 30B, available on Fireworks for always-on agents

    Meta's Muse Glimmer is a 30B dense model with a 128K+ token context window, now available on Fireworks in serverless and on-demand deployments. Meta reports it leads its size class on MCP Atlas (75.5) and DeepSearch QA (74.6) against Gemma 4 31B and Qwen 3.6 27B, with its sliding-window attention and two KV heads keeping the cache small for concurrent agent sessions.

    AIWhy it matters: The post pairs an architecture explained through KV cache size with benchmark tables against two rival models, which helps readers judge whether it fits their agent workload.

Aug 7

Aug 7Fri
  1. Qwen · new models on Hugging FaceAI score88

    Qwen releases open-weight Qwen3.8-2.4T-A95B, a 2.4T-parameter MoE model

    Qwen has released the Qwen3.8-2.4T-A95B model weights on Hugging Face, with 2.4T total and 95B activated parameters in a mixture-of-experts design. The release supports reasoning_effort levels and a 262,144-token native context extensible to 1,010,000 tokens, and it is text-only with thinking mode always on. The source reports benchmark results against Opus 4.8, Fable 5, GPT 5.6 Sol, and Qwen3.7-Max, and says the official Qwen3.8-Max API adds vision input and a 1M default context.

    AIWhy it matters: The model card gives parameters, architecture, reasoning controls, and benchmark tables against named rival models, showing what an open release of this scale actually offers.

  2. MiniMax · new models on Hugging FaceAI score44

    MiniMax Music 3 generates five-minute songs with coherent structure and vocals

    MiniMax Music 3 is a music generation model that creates complete songs up to five minutes long from lyrics and a music description. It pairs an 8B Global LLM for long-range structure with a 0.6B Local LLM for acoustic detail, outputting 32 kHz, 16-bit stereo WAV audio. The model is available on Hugging Face and supports SGLang-Omni, diffusers, and ComfyUI.

Aug 6

Aug 6Thu
  1. Intern Large ModelsAI score62

    Shanghai AI Lab open-sources Mobius, a Transformer alternative claiming 4x faster reasoning

    Shanghai AI Lab open-sourced Mobius, an architecture its authors compare to the RNN-to-Transformer shift in both token and knowledge dimensions. Against Transformers, the post claims about 4x faster reasoning, the same MMLU score with 40% less data, and 2x better compositional generalization. Mobius is supported by XTuner, LMDeploy, vLLM, and SGLang, and its experimental setup and training pipeline will be released later.

  2. InternLM (Shanghai AI Lab) · new models on Hugging FaceAI score38

    Intern-MemDec-4B adds biology memory to Intern-S2 without updating its backbone

    Shanghai AI Lab's InternLM released Intern-MemDec-4B, a 4B-parameter memory decoder that runs alongside an Intern-S2 backbone and a token-level router to add biology knowledge. On all 21 Biology-Instructions tasks, the average score rose from 56.92 to 60.32 when paired with Intern-S2-Preview-397B. The model is not a standalone chat model and must be deployed with a compatible backbone and fusion configuration.

Aug 5

Aug 5Wed
  1. Qwen · new models on Hugging FaceAI score79

    Qwen3.8-27B releases dense vision-language model with thinking controls

    Alibaba's Qwen team has released Qwen3.8-27B on Hugging Face as a 27B dense model with native image and video understanding. The model card reports gains over Qwen3.6-27B on coding and agent benchmarks, including SWE-bench Pro at 61.7 versus 53.5. It adds reasoning_effort levels and preserve_thinking, and its hosted Qwen Cloud version is described as coming soon.

    AIWhy it matters: The model card gives per-benchmark comparisons with Qwen3.6-27B and named rivals, plus reasoning_effort and preserve_thinking controls for judging cost and agent behavior.

Aug 3

Aug 3Mon
  1. Liquid AI BlogAI score72

    Liquid AI releases LFM2.5-2.6B, a 2.6B on-device agentic model

    Liquid AI released LFM2.5-2.6B, a 2.6B-parameter agentic model that runs on-device on phones and CPUs, along with a base variant on Hugging Face. The company reports it leads on every instruction-following benchmark and nearly every tool-use benchmark it tested, and decodes 220 tokens/s on an M5 Max. The source says larger models may still suit complex agentic or coding-heavy tasks.

    AIWhy it matters: The source reports benchmark results against several same-tier models and notes where larger models still lead, which helps judge fit for edge agent workloads.

Aug 1

Aug 1Sat

Jul 31

Jul 31Fri
  1. Thinking MachinesAI score44

    Releasing weights indiscriminately isn't safe. Neither is keeping capable models inside a few labs. We think there's a path between them. We haven't mapped all of it. Our new post covers the part we can see: how we assessed Inkling, and why access should widen in stages. https://thinkingmachines.ai/blog/a-safe-path-to-open-weights

    Releasing weights indiscriminately isn't safe. Neither is keeping capable models inside a few labs. We think there's a path between them. We haven't mapped all of it. Our new post covers the part we can see: how we assessed Inkling, and why access should widen in stages. https://thinkingmachines.ai/blog/a-safe-path-to-open-weights

  2. DeepSeek · new models on Hugging FaceAI score75

    DeepSeek releases DeepSeek-V4-Flash-0731 with stronger agentic capabilities

    DeepSeek has released DeepSeek-V4-Flash-0731 as the official version superseding the preview, with substantially enhanced agentic capabilities. The source reports it outperforms DeepSeek-V4-Pro (Preview) on listed benchmarks, including Terminal Bench 2.1 at 82.7 versus 72.1, despite a far smaller activated parameter count. The model ships under the MIT License with DSpark speculative decoding supported in vLLM and SGLang.

    AIWhy it matters: The release shows benchmark gains over the preview and a concrete vLLM and SGLang serving path, useful for teams weighing a self-hosted agentic coding model.

  3. DeepSeekAI score38

    ⚠️ Note 🔷 DeepSeek-V4-Flash-0731 keeps the exact same model architecture and size as the preview version. 🔷 Today's upgrade applies ONLY to the DeepSeek-V4-Flash API. The DeepSeek-V4-Pro API and App/Web models remain unchanged for now. The official release of DeepSeek-V4-Pro is coming ASAP! Stay tuned.

    ⚠️ Note 🔷 DeepSeek-V4-Flash-0731 keeps the exact same model architecture and size as the preview version. 🔷 Today's upgrade applies ONLY to the DeepSeek-V4-Flash API. The DeepSeek-V4-Pro API and App/Web models remain unchanged for now. The official release of DeepSeek-V4-Pro is coming ASAP! Stay tuned.

  4. DeepSeek API NewsAI score67

    DeepSeek-V4-Flash API enters public beta with stronger agent benchmarks

    DeepSeek has released the DeepSeek-V4-Flash API in public beta, and developers can use the latest version by setting the model name to deepseek-v4-flash. The source reports agent benchmark results far above V4-Pro-Preview, including 82.7 on Terminal Bench 2.1 and 70.3 on Toolathlon verified. V4-Flash natively supports the Responses API format and is adapted for Codex, while V4-Pro and the APP/WEB models are unchanged.

    AIWhy it matters: The release lists agent benchmark results against V4-Pro-Preview and notes Responses API support for Codex, which helps developers gauge the upgrade's practical effect on their workflows.

Jul 30

Jul 30Thu
  1. MiniMax BlogAI score72

    MiniMax H3 unifies text, image, video, and audio generation in one model

    MiniMax launches H3, a general-purpose multimodal generation model that understands text, images, video, and audio as unified context. It generates video up to 15 seconds at 2K resolution with native stereo sound, and the company says model weights will be opened in the coming days, subject to applicable laws and regulations. MiniMax also says H3 is priced below mainstream models at 2K and 768p.

    AIWhy it matters: The post explains how a unified multimodal design and training choices enable 2K video with native stereo sound, useful for comparing against closed video generators.

  2. Thinking Machines LabAI score65

    Thinking Machines proposes staged, evidence-based release path for open-weight models

    Thinking Machines argues that safe open-weight releases depend on both model safety testing and readiness of the surrounding ecosystem, and that release should proceed in iterative stages. For its Inkling and Inkling-Small models, internal evaluations, four external red-teaming groups, and adversarial fine-tuning tests led the company to conclude that releasing the weights was not likely to add material risk beyond existing open-weight models.

    AIWhy it matters: The post lays out a staged, evidence-gated path to releasing open weights, with concrete safety tests and the ecosystem measures behind each stage.

  3. Thinking MachinesAI score43

    Like Inkling, it's natively multimodal. It’s encoder-free, with audio and images processed jointly with text. It nearly matches Inkling across multimodal evals, and it can use Python to crop, zoom, and inspect images while reasoning over documents and charts.

    Like Inkling, it's natively multimodal. It’s encoder-free, with audio and images processed jointly with text. It nearly matches Inkling across multimodal evals, and it can use Python to crop, zoom, and inspect images while reasoning over documents and charts.

  4. Thinking MachinesAI score44

    Inkling-Small began training after its larger counterpart, so it benefits from everything we learned: an improved pre-training data mix, a refined ML recipe, on-policy distillation with Inkling as the teacher, and two further weeks of agentic coding RL.

    Inkling-Small began training after its larger counterpart, so it benefits from everything we learned: an improved pre-training data mix, a refined ML recipe, on-policy distillation with Inkling as the teacher, and two further weeks of agentic coding RL.

  5. Thinking MachinesAI score38

    Efficiency is the point. Across agentic tool use (Terminal-Bench 2.1), reasoning (HLE), and instruction following (IFBench), Inkling-Small delivers more performance per FLOP than Inkling. Variable thinking effort lets you pick your point on the cost/performance curve.

    Efficiency is the point. Across agentic tool use (Terminal-Bench 2.1), reasoning (HLE), and instruction following (IFBench), Inkling-Small delivers more performance per FLOP than Inkling. Variable thinking effort lets you pick your point on the cost/performance curve.

  6. IdeogramAI score43

    Today, we're introducing P-Image-Ideogram: a family of Pareto-optimal image models with the best quality-speed-cost trade-off, co-developed with @PrunaAI. Four quality modes. Native 1K and 2K generation. From $0.003 per image. Live now on the API and all our partner platforms. Give it a try: https://ideogram.ai/tools/p-image-ideogram

    Today, we're introducing P-Image-Ideogram: a family of Pareto-optimal image models with the best quality-speed-cost trade-off, co-developed with @PrunaAI. Four quality modes. Native 1K and 2K generation. From $0.003 per image. Live now on the API and all our partner platforms. Give it a try: https://ideogram.ai/tools/p-image-ideogram

Jul 29

Jul 29Wed
  1. Google LabsAI score60

    Google Launches Lyria 3.5 in Flow Music With Better Vocals and Lyrics

    Google is rolling out Lyria 3.5, its newest music generation model, in Google Flow Music today. The update improves musicality, lyric quality and prompt adherence, and vocal expressiveness and pronunciation, and gives users more control over tempo and duration.

    AIWhy it matters: The post names the specific capability changes and where users can access them, which helps readers judge fit for music creation workflows.

  2. Air Street PressAI score75

    Poolside's Laguna S 2.1 is an open agentic coding model that runs on one DGX Spark

    Poolside released Laguna S 2.1, an open-weights agentic coding model with 118 billion total parameters and about 8 billion active per token, supporting up to a million tokens of context. Quantized, it fits on one NVIDIA DGX Spark, and Poolside reports 70.2% on Terminal-Bench 2.1 with thinking enabled, with its evaluation trajectories published online. The same week it shipped the Poolside Desktop Assistant for macOS, which runs Laguna locally or alongside Claude Code, Codex, and Gemini agents.

  3. Liquid AI NewsletterAI score46

    Liquid AI Expands LFM2 Tokenizer to 128K, Speeding On-Device Thai, Vietnamese, and Hindi

    Liquid AI doubled the LFM2 tokenizer's vocabulary from 65K to 128K without retraining from scratch, extending the original BPE merges and initializing new embeddings as the mean of their sub-tokens. The expanded tokenizer needs 4.0× fewer tokens for Thai, 2.6× fewer for Vietnamese, and 2.4× fewer for Hindi, which the source says yields roughly 2.2–3.7× faster on-device decoding for these languages with no reported quality loss on previously supported languages. LFM2.5-8B-A1B and the expanded tokenizer are available on Hugging Face with open weights.

  4. Alibaba NLP (Tongyi) · new models on Hugging FaceAI score40

    Alibaba NLP releases UEmbed-9B, a unified sparse and dense multimodal embedding model

    Alibaba NLP has released UEmbed-9B, a decoder-only multimodal embedding model built on Qwen3.5 9B that outputs both dense and SPLADE-style sparse embeddings from one forward pass. It supports text, image, video, and mixed-modal inputs for retrieval and multimodal search, and the family also includes 2B and 4B variants. The model is available on Hugging Face, with transformers and vLLM inference support.

  5. Alibaba NLP (Tongyi) · new models on Hugging FaceAI score38

    Alibaba NLP releases UEmbed-4B, a unified sparse and dense multimodal embedding model

    Alibaba NLP has released UEmbed-4B, a decoder-only multimodal embedding model built on Qwen3.5 4B that outputs both dense and sparse embeddings from one forward pass. It handles text, image, video, and mixed-modal inputs for retrieval and visual-document search, and sparse activations map to vocabulary terms usable with inverted indexes. The model is available on Hugging Face in a family that also includes 2B and 9B variants.

  6. Alibaba NLP (Tongyi) · new models on Hugging FaceAI score43

    Alibaba-NLP releases UEmbed-2B, a multimodal model producing dense and sparse embeddings

    Alibaba-NLP's UEmbed-2B, a decoder-only multimodal embedding model built on Qwen3.5 2B, produces both dense and SPLADE-style sparse embeddings from a single forward pass. It supports text, image, video, and mixed-modal inputs for retrieval, and the 4B and 9B variants are also available. The team reports state-of-the-art results on the text and agent tracks of MMEB-v3.

Jul 28

Jul 28Tue
  1. Augment Code BlogAI score39

    GPT-5.6 Sol Becomes Augment Cosmos's Default Model for Token Efficiency

    Augment Code has made GPT-5.6 Sol the default model in Cosmos, choosing it as the most token-efficient model to clear its pass-rate floor for long-horizon software engineering tasks. The company ranks models by cost per task rather than list price per million tokens, since retries on failed steps add token spend. Users can still select any model, and the default will change as more token-efficient models emerge.

  2. MiniMax · new models on Hugging FaceAI score76

    MiniMax H3 releases open-weight omni-modal video model with native stereo audio

    MiniMax released H3, an open-weights omni-modal model that generates video with native stereo audio up to 2K and 15 seconds. The system combines H3-Context-IR preprocessing, the H3-Base generator at 768p, and H3-Regenerate-2K for 2K output, with the Context-IR and 2K modules available only through API.

    AIWhy it matters: The source details a three-module pipeline and open weights with deployment paths, showing how a video model is served and reproduced locally.

  3. Intern Large ModelsAI score62

    Intern Large Models introduces Visual Pretraining learned from visual documents

    Intern Large Models introduces Visual Pretraining, a pretraining paradigm for foundation models that learns directly from visual documents. The post says it outperforms text-only pretraining across backbones and benchmarks, and links the arXiv paper 2607.09657 along with Intern-S2-Preview (35B) and Intern-S2-Preview-397B on Hugging Face, the latter presented as a multimodal foundation model trained with this recipe.

Jul 27

Jul 27Mon
  1. Liquid AI BlogAI score49

    Liquid AI Releases LFM2.5-Encoders for Fast Long-Context Encoding on CPU

    Liquid AI released LFM2.5-Encoder-230M and LFM2.5-Encoder-350M, bidirectional encoders built on the LFM2 hybrid architecture and available on Hugging Face. They support an 8,192-token context and are designed for fine-tuning on classification and token-level tasks. On CPU, LFM2.5-Encoder-230M is the fastest model tested from 1K tokens up, running about 3.7x faster than ModernBERT-base at 8,192 tokens.

  2. KimiAI score65

    Kimi K3 becomes available on Nebius Token Factory via API

    Kimi K3 is now available on Nebius Token Factory, which is named a Day 0 launch partner, through an OpenAI-compatible API and console. The quoted post says Artificial Analysis scores the open-weight model at 57 on its Intelligence Index, two points behind GPT-5.6 Sol (max), and lists up to 1M tokens of context.

    AIWhy it matters: The source names the cloud access route and an Artificial Analysis score of 57, letting readers compare Kimi K3 against GPT-5.6 Sol.

  3. KimiAI score86

    Moonshot AI releases Kimi K3 weights and technical report

    Moonshot AI is releasing the model weights and technical report for Kimi K3, a 2.8T-parameter MoE model with native visual understanding and a 1M-token context window. The post says the new architecture delivers 2.5x the intelligence per unit of compute, and the company is also opening high-performance attention kernels, an MoE communication library, and infrastructure for running agent environments at scale.

    AIWhy it matters: The source names the model size, context window, and released weights, which helps readers compare its scale and openness with other frontier releases.