Skip to contentSkip to stories

Updated

#Multimodal

Showing low-relevance items too. Hide low-relevance items

Aug 30

Aug 30Sun
  1. Alibaba NLP (Tongyi) · new models on Hugging FaceOfficialAI score40

    Alibaba NLP Releases Core-Embed 8B for Compositional Multimodal Retrieval

    AIAlibaba NLP has released core-emb-8b, an MLLM-based multimodal embedding model that distills a reranker's compositional judgments to distinguish attribute-object bindings such as "a white plate and a black chair" versus "a black plate and a white chair." The 8B dense embedding model, built on the Qwen3-VL-based VL-Emb backbone, scores 0.666 total average on compositional benchmarks, 5.7 points above its backbone. It is part of a family that also includes 2B embedding and reranker models.

  2. Alibaba NLP (Tongyi) · new models on Hugging FaceOfficialAI score38

    Alibaba-NLP releases Core-Reranker-8B, a compositional multimodal reranker on Hugging Face

    AIAlibaba-NLP has published Core-Reranker-8B on Hugging Face, an 8B-parameter multimodal reranker fine-tuned from Qwen3-VL-Reranker to better distinguish attribute-object bindings in text and image relevance scoring. On compositional reasoning benchmarks COLA, SugarCrepe++, and NegBench, it reports an 82.7% total average, 10.7 points above Jina-Reranker. The model is part of the Core-Embed family, which also includes 2B and 8B embedding models, with Core-Embed-8B reporting a 0.666 total average.

  3. Alibaba NLP (Tongyi) · new models on Hugging FaceOfficialAI score40

    Alibaba NLP releases Core-Embed multimodal embedding models for compositional retrieval

    AIAlibaba NLP has released core-emb-2b and core-emb-8b, multimodal embedding models built on Qwen3-VL that distill reranker judgments to better match attribute-object bindings in text and image retrieval. The Core-Embed-8B model posts the best total average (0.666) among evaluated embedding models on compositional benchmarks, 5.7 points above its VL-Emb-8B backbone. Companion Core-Reranker-2B and 8B models are also available, with the 8B reranker reaching 82.7% total average on the same benchmarks.

  4. Alibaba NLP (Tongyi) · new models on Hugging FaceOfficialAI score36

    Alibaba's core-reranker-2b Model Targets Compositional Image-Text Relevance Scoring

    AIAlibaba NLP released core-reranker-2b, a 2B-parameter multimodal relevance-scoring model built on Qwen3-VL-Reranker to better distinguish attribute-object bindings in text and image pairs. The Core-Reranker family also includes an 8B variant, and Core-Reranker-8B reports an 82.7% total average on compositional reasoning benchmarks COLA, SugarCrepe++, and NegBench, 10.7 points above Jina-Reranker. Usage details are provided in the source, including loading through the GitHub repository wrapper classes.

Aug 27

Aug 27Thu
  1. LMSYS OrgOfficialAI score47

    MiniMax-H3 gets up to 6.24x speedup on 8×H200 GPUs

    AIMiniMax-H3 on 8×H200 GPUs reaches 1.85–1.95x lossless speedup over Diffusers without approximation, with fixed prompts, seeds, resolution, FPS, and 50 denoising steps. Adding step reuse and sparse attention raises speedup to as much as 6.24x, but quality varies by workload, with SSIM from 0.76 to 0.91. Two presets trade off the two: a conservative Cache-DiT setting gives 2.99x at 0.90–0.98 SSIM, while a faster SubBlock 0.75 plus Cache-DiT stride gives 4.90–5.93x at 0.77–0.92.

    Image from @lmsysorg's post
  2. Qwen · new models on Hugging FaceOfficialAI score62

    Qwen-Drive-1.0 releases open weights for driving VQA, perception, and planning

    AIQwen has published Qwen-Drive-1.0-4B on Hugging Face, a vision-language model for autonomous driving built on Qwen3.5-4B. The release includes a BEV perception head and two Planning Experts, planner-sft and planner-rl, with code and an inference example in the linked GitHub repository.

    Why it matters: The source gives concrete benchmark results and a runnable setup, letting readers judge how a driving VLM with planning and perception heads compares with existing systems.

Aug 26

Aug 26Wed
  1. Bryan CatanzaroXAI score46

    NVIDIA Releases DLSS 4.5 Ray Reconstruction with Better Image Quality

    AINVIDIA's DLSS 4.5 Ray Reconstruction is now available, using a second-generation joint denoiser and super-resolution model. According to the post, it delivers much better image quality at the same compute cost, pushing the trade-off between image quality and rendering cost further.

  2. LM StudioOfficialAI score57

    GLM-5.3-Flash by Z.ai is now live in LM Studio

    AILM Studio announced that Z.ai's GLM-5.3-Flash, previously previewed as Ox Alpha, is available in LM Studio Bionic. The source says the model outperforms GLM-5.2 at 9-10x lower cost, supports image input, and is served from US-based servers with ZDR enabled by default.

  3. LMSYS OrgOfficialAI score65

    Zhipu's GLM-5.3-Flash adds native vision with day-0 SGLang support

    AIZ.ai released GLM-5.3-Flash, a 320B-A18B model, with day-0 support in SGLang, after appearing earlier as ox-alpha. The post calls it the first native multimodal model in the GLM-5 series and says it outperforms GLM-5.2 at one-tenth the cost, with stable 1M-token long-context performance.

    Why it matters: The post reports GLM-5.3-Flash's native multimodal design, its efficiency claims, and day-0 SGLang support, which bear on running it in practice.

Aug 25

Aug 25Tue
  1. Google Developers BlogOfficialAI score35

    Google Brings Qwen3-Embedding-8B to Cloud TPU via vLLM with Long-Context Support

    AIGoogle Cloud has added native TPU support to vLLM and engineered optimizations to serve the Qwen3-Embedding-8B model on Cloud TPU, targeting 4K+ token text and 15K+ token multimodal inputs. The work addresses tensor alignment, lazy-loading, compilation pre-warming, and long-context pooling, with a cosine similarity pass threshold of at least 0.999 for text and 0.995 for multimodal inputs against XPU reference vectors.

  2. Z.ai Release NotesOfficialAI score62

    Z.ai releases GLM-5.3-Flash with native visual capabilities and hybrid architecture

    AIZ.ai has released GLM-5.3-Flash, a model with native visual capabilities that observe interfaces, rendering results, and interaction feedback across code, browsers, and GUIs. It uses a hybrid linear and sparse attention architecture with 320B total parameters and 18B activated, which the company says significantly reduces compute and KV-cache requirements. The release notes also describe support for office document and financial research workflows.

    Why it matters: The release notes give GLM-5.3-Flash's architecture, parameter counts, and cybersecurity findings, which make the model's scope concrete for comparison with earlier GLM releases.

  3. Stability AIOfficialAI score36

    Stability AI raises $76M Series B backed by Electronic Arts, Sony Music, Universal Music, Warner Music

    AIStability AI announced a $76M Series B round, bringing total funding to $232M under CEO Prem Akkaraju, with new investors including Electronic Arts, Sony Music Group, Universal Music Group, and Warner Music Group. The company said the capital will fund its creative production product suite, applied research, and professional services. The announcement followed the launch of Stable Audio 3.0, a family of open-weight music models trained on fully licensed data.

  4. Z.ai (GLM) · new models on Hugging FaceOfficialAI score72

    Z.ai releases GLM-5.3-Flash, a natively multimodal model with 320B parameters

    AIZ.ai released GLM-5.3-Flash on Hugging Face, the first natively multimodal model in the GLM-5 series, with 320B total parameters and 18B active parameters. The source says it outperforms GLM-5.2 across benchmarks at one-tenth the price and approaches Claude Opus 4.8 on coding and agentic benchmarks. It adopts a hybrid sparse and linear attention architecture to reduce long-context serving costs.

    Why it matters: The release shows a hybrid sparse and linear attention design aimed at cutting long-context serving costs, which is useful for comparing efficiency trade-offs.

Aug 21

Aug 21Fri
  1. DeepSeekOfficialAI score42

    DeepSeek adds vision API support via deepseek-v4-flash-vision-exp model

    AIDeepSeek's API now accepts multimodal input through the model deepseek-v4-flash-vision-exp, supporting mixed text and image requests. Each image is billed at up to 384 tokens at V4-Flash pricing, and it works with Chat Completions, Messages, and Responses endpoints. Images can be supplied as base64, external URLs, or via the Files API.

  2. DeepSeekOfficialAI score62

    DeepSeek releases experimental multimodal model V4-Flash-Vision-Exp on its API

    AIDeepSeek has made its experimental multimodal model DeepSeek-V4-Flash-Vision-Exp available on the DeepSeek API Platform. The company says it matches DeepSeek-V4-Flash on text tasks, including agents, reasoning, and world knowledge. On multimodal agent benchmarks it improves substantially over V4-Flash and approaches Opus-4.8, and DeepSeek Harness 0.1.1 was released the same day with support for the new model.

    Image from @deepseek_ai's post
  3. DeepSeek API NewsOfficialAI score60

    DeepSeek releases experimental vision model DeepSeek-V4-Flash-Vision-Exp on its API

    AIDeepSeek has made DeepSeek-V4-Flash-Vision-Exp, an experimental multimodal vision understanding model, available on its API platform via model='deepseek-v4-flash-vision-exp'. The source says its pure-text capabilities are on par with DeepSeek-V4-Flash, while it shows a significant leap on agent benchmarks requiring visual understanding, which it says brings multimodal agent capabilities close to Opus-4.8.

    Why it matters: The source gives benchmark scores and a model identifier, so readers can compare the experimental vision model against the text-only DeepSeek-V4-Flash on agent tasks.

Aug 19

Aug 19Wed
  1. TinkerOfficialAI score38

    Qwen3.8-27B is now available on Tinker

    AITinker has made Qwen3.8-27B available today. The model is natively multimodal, handling images and video, with flexible thinking control. Tinker says it performs meaningfully better at coding, professional work, research, and long-horizon agentic tasks.

  2. Google · new models on Hugging FaceOfficialAI score22

    Google releases TIPS g/14 low-res v1 vision-language model on Hugging Face

    AIGoogle has released TIPS g/14 low-res (v1) on Hugging Face, a Text-Image Pre-training with Spatial awareness vision-language model with 1.1B vision parameters and 389M text parameters. The model produces spatially rich image features aligned with text embeddings at 224 resolution, under the Apache 2.0 license. It supports image encoding, text encoding, and zero-shot classification via the transformers library.

  3. Google · new models on Hugging FaceOfficialAI score26

    Google releases TIPS g/14 v1 vision-language model on Hugging Face

    AIGoogle has released the original TIPS g/14 (v1) vision-language model on Hugging Face under Apache 2.0, with 1.1B vision parameters and 389M text parameters at 448 resolution. The TIPS family, presented at ICLR 2025, produces spatially rich image features aligned with text embeddings, and the release includes a low-res 224 variant.

  4. Google · new models on Hugging FaceOfficialAI score22

    TIPS So400m/14 v1 Vision-Language Model Released on Hugging Face

    AIGoogle released google/tipsv1-so400m14, the original v1 So400m/14 checkpoint of TIPS, a contrastive vision-language model that produces spatially rich image features aligned with text embeddings. The model has 413M vision parameters and 448M text parameters at 448 resolution, and is licensed under Apache 2.0.

  5. Google · new models on Hugging FaceOfficialAI score22

    Google releases TIPS L/14 v1 vision-language model on Hugging Face

    AIGoogle has published google/tipsv1-l14, the original v1 L/14 release of TIPS, a contrastive vision-language model that produces spatially rich image features aligned with text embeddings. The L/14 variant has 304M vision parameters and 184M text parameters at 448 resolution, with an embedding dimension of 1024, and is licensed under Apache 2.0.

  6. Google · new models on Hugging FaceOfficialAI score22

    Google releases TIPS B/14 v1 vision-language model on Hugging Face

    AIGoogle has published TIPS B/14 (v1) on Hugging Face, a contrastive vision-language model that produces spatially rich image features aligned with text embeddings. The model has 86M vision parameters and 110M text parameters at native 448 resolution, and is licensed under Apache 2.0. The release includes usage code for image and text encoding, zero-shot classification, and spatial feature visualization.

Aug 18

Aug 18Tue
  1. Stability AIOfficialAI score38

    Stability AI launches Stable Audio 3.0 plugin and enhanced web experience

    AIStability AI has released two beta tools for Stable Audio 3.0: a plugin that brings audio generation into digital audio workstations (DAWs) and an upgraded experience at StableAudio.com with more editing options. Both are powered by commercially-safe models, so users own their outputs and can distribute them freely. Some features are experimental, and the company says it will keep iterating in real time.

Aug 17

Aug 17Mon
  1. Z.ai Release NotesOfficialAI score63

    Z.ai releases GLM-5.3 with stronger coding and vulnerability discovery

    AIZ.ai's release notes announce GLM-5.3, which the company says delivers a 50% gain over GLM-5.2 on Z.ai Code Bench and reaches open-source SOTA on public benchmarks including Terminal Bench 3.0. The company also reports that GLM-5.3 matches Mythos 5 in white-box code review and vulnerability discovery, identifying 2,436 vulnerabilities in real-world targets, 1,097 of them medium- or high-severity. A separate GLM-5.3-Flash entry describes native visual capabilities and a hybrid architecture with 320B total and 18B activated parameters.

    Why it matters: The release notes show GLM-5.3's coding and cybersecurity gains, with a vulnerability count, letting readers compare it against Z.ai's prior GLM-5.x line and other coding models.

Aug 16

Aug 16Sun
  1. Ian Johnson 🔬🤖XAI score34

    Ian Johnson maps Prelinger film dataset with UMAP and Marlin-2B vision latents

    AIIan Johnson used UMAP to visualize a video dataset, adding vision latents extracted from Marlin-2B for each clip alongside the included embeddings. He built the interactive map to render smoothly in the browser, with a writeup linked in the post. The quoted post by Daniel van Strien describes indexing 370 hours of Prelinger Archives films into 23,148 timestamped searchable moments.

    Video from @enjalot's post
  2. Philipp SchmidBlogAI score58

    Controlling Android with Gemini 3.7 Flash and 150 lines of Python

    AIThe author built a Python agent that uses Gemini 3.7 Flash to control an Android emulator from raw screenshots, returning normalized 0–999 coordinates that are scaled to 1080x1920 pixels over ADB. In a test, the agent opened Chrome, closed popups, and solved one round of Wordle in two guesses without accessibility IDs or DOM access. The article presents the loop as usable for UI testing and task automation across native apps, webviews, and canvas interfaces, with code in an open-source quickstart repository.

Aug 14

Aug 14Fri
  1. Cohere · new models on Hugging FaceOfficialAI score60

    Cohere releases North Small Translate 1.0 open weights for 50-language translation

    AICohere and Cohere Labs released North Small Translate 1.0 as open weights for research, a sparse Mixture-of-Experts model with 25B active and 218B total parameters. It is specialized for machine translation across 50 languages, with a 16K input and 16K output context. The chart shows a WMT26 all-languages score of 83.60, rising to 84.36 with the agentic multi-pass workflow, and the model is licensed CC BY-NC 4.0 with an acceptable use policy.

    Why it matters: The model card lists the benchmark score, hardware needs, and license terms, which helps readers judge whether this translation model fits their use.

Aug 13

Aug 13Thu
  1. ByteDance · new models on Hugging FaceOfficialAI score52

    ByteDance releases Bernini-Diffusers-v2 video generation and editing model

    AIByteDance has released Bernini-Diffusers-v2 on Hugging Face, a video generation and editing pipeline combining a Qwen2.5-VL planner with Wan2.2 diffusion components. The model card recommends it over Bernini-R for complex requests needing stronger instruction following and multi-step semantic planning. Code and weights are available under Apache License 2.0.

Aug 11

Aug 11Tue
  1. Liquid AI BlogOfficialAI score62

    Liquid AI releases LFM2.5-VL-3B, a 3B vision-language model for edge devices

    AILiquid AI released LFM2.5-VL-3B, an open-weight 3B vision-language model that it says rivals models twice its size while running faster on CPU and GPU. Benchmarks show large gains over LFM2-VL-3B, including ScreenSpot-v2 averaging 80.7, RefCOCO precision@1 rising from 57.1 to 87.9, and ToolSandbox rising from 26.4 to 59.5. The model is available on Hugging Face and decodes 228 tokens/s on an Apple M5 Max.

    Why it matters: The post pairs benchmark gains with on-device and GPU throughput figures, showing how a 3B vision model trades size against speed and accuracy.

  2. Liquid AI · new models on Hugging FaceOfficialAI score40

    LiquidAI releases LFM2.5-VL-3B, a 3B multimodal model for on-device use

    AILiquidAI has released LFM2.5-VL-3B, a 3B-parameter multimodal model that processes text and images and is built on the LFM2.5-2.6B language model with a SigLIP2 NaFlex vision encoder. It runs at 228 tokens/s on an Apple M5 Max and 116 tokens/s on an AMD Ryzen AI Max+ 395 in under 3.3 GB of memory, with a 32,768-token context length. The model is available in native, GGUF, ONNX and MLX formats on Hugging Face.

Aug 10

Aug 10Mon
  1. Cohere · new models on Hugging FaceOfficialAI score46

    Cohere releases North Micro Vision Instruct, a 2.4B open-weight vision-language model

    AICohere has released North Micro Vision Instruct, a 2.4B-parameter open-weight vision-language model under the Apache 2.0 license, on Hugging Face. The model processes images at native resolution and handles visual question answering, captioning, grounding, OCR, and document understanding across English, German, French, Spanish, Italian, Portuguese, Hindi, Japanese, Korean, Chinese, and Arabic. It has a 128K-token language backbone context window, but its validated multimodal range is up to 8K tokens.

Aug 6

Aug 6Thu

Aug 5

Aug 5Wed
  1. Qwen · new models on Hugging FaceOfficialAI score79

    Qwen3.8-27B releases dense vision-language model with thinking controls

    AIAlibaba's Qwen team has released Qwen3.8-27B on Hugging Face as a 27B dense model with native image and video understanding. The model card reports gains over Qwen3.6-27B on coding and agent benchmarks, including SWE-bench Pro at 61.7 versus 53.5. It adds reasoning_effort levels and preserve_thinking, and its hosted Qwen Cloud version is described as coming soon.

    Why it matters: The model card gives per-benchmark comparisons with Qwen3.6-27B and named rivals, plus reasoning_effort and preserve_thinking controls for judging cost and agent behavior.