Skip to contentSkip to stories

Updated

#Hugging Face

Items with an AI score under 20 are hidden. Show low-relevance items

Jun 26

Jun 26Fri
  1. Qwen · new models on Hugging FaceAI score44

    Qwen3-ForcedAligner-0.6B-hf Adds Timestamp Alignment for Speech Transcripts

    AIQwen released Qwen3-ForcedAligner-0.6B-hf, a Transformers-format forced aligner that predicts timestamps for arbitrary units within up to 5 minutes of speech in 11 languages. The model accepts transcripts from any ASR system, and the documentation shows it paired with Qwen3-ASR-0.6B and NVIDIA Parakeet CTC. Until it ships in an official Transformers release, users must install Transformers from source.

Jun 25

Jun 25Thu

Jun 18

Jun 18Thu
  1. Cohere · new models on Hugging FaceAI score43

    Cohere Releases Open-Source 2B Arabic Speech Recognition Model Transcribe Arabic

    AICohere and Cohere Labs released Cohere Transcribe Arabic, an open-source 2B-parameter Arabic automatic speech recognition model under Apache 2.0. It is optimized for Arabic, Arabic dialects, English, and Arabic-English code-switched speech, using a Conformer encoder-decoder architecture supported natively in Transformers. The model's average WER of 25.87 and CER of 11.80 on the Open Universal Arabic ASR Leaderboard, as of 07.07.2026, is reported in the source.

Jun 12

Jun 12Fri
  1. Georgi GerganovAI score34

    Gerganov flags locate-anything.cpp, a ggml runtime for NVIDIA's LocateAnything-3B

    AIGeorgi Gerganov highlighted locate-anything.cpp, a native C++/ggml inference implementation of NVIDIA's LocateAnything-3B for open-vocabulary object detection and visual grounding. The project, from the LocalAI team, runs without Python on CPU and GPU, with quantized GGUF weights published on Hugging Face.

Jun 8

Jun 8Mon
  1. ByteDance · new models on Hugging FaceAI score46

    ByteDance Open-Sources Bernini-R 1.3B Video Diffusion Renderer on Hugging Face

    AIByteDance has open-sourced the 1.3B-parameter weights of its Bernini Renderer (Bernini-R), available on Hugging Face as ByteDance/Bernini-R-1.3B-Diffusers. Fine-tuned from Wan2.1-1.3B, the model performs close to the 14B variant on simple tasks such as style transfer, subtitle or watermark removal, and local editing, but lags on complex tasks such as human generation. The release requires a CUDA GPU, with an H100 recommended for FlashAttention-3.

May 20

May 20Wed
  1. PaddlePaddleAI score36

    PaddleOCR 3.5 adds Hugging Face Transformers as inference backend

    AIPaddleOCR 3.5 now supports Hugging Face Transformers as an inference backend, letting users run PP-OCRv5 and PaddleOCR-VL 1.5 models directly within the Transformers ecosystem. Users can select it with engine="transformers" while keeping the same PaddleOCR pipeline, which the post says eases integration for RAG and Document AI applications.

Apr 27

Apr 27Mon
  1. Xiaomi MiMo · new models on Hugging FaceAI score72

    Xiaomi releases MiMo-V2.5, an open omnimodal model with 1M context

    AIXiaomi's MiMo-V2.5 is a native omnimodal model that understands text, image, video, and audio within one architecture. It is a sparse MoE with 310B total and 15B activated parameters, and supports up to 1M tokens of context. The repository also notes a config.json and tokenizer_config.json update that users who downloaded before commit 4da2748 should re-pull.

    Why it matters: The repository documents a 310B-parameter omnimodal MoE with a hybrid attention design, useful for comparing long-context efficiency against other open multimodal models.

Apr 23

Apr 23Thu
  1. Apple · new models on Hugging FaceAI score40

    Apple releases CADD-Base-7B, a masked diffusion model for code generation

    AIApple has released CADD-Base-7B on Hugging Face, a 7B masked diffusion language model for code generation that uses Continuously Augmented Discrete Diffusion (CADD) to guide discrete denoising with a continuous flow-matching signal. The model loads through Transformers with trust_remote_code, and its diffusion_generate method supports CADD sampling modes "weighted" and "argmax" with alg options such as "entropy" and "maskgit_plus". The release builds on DiffuCoder and reuses Dream's modeling architecture and generation utilities.

Apr 6

Apr 6Mon
  1. Black Forest Labs · new models on Hugging FaceAI score41

    FLUX.2 Small Decoder offers faster, lower-VRAM drop-in replacement for FLUX.2 decoder

    AIBlack Forest Labs released FLUX.2 Small Decoder, a distilled VAE decoder that works as a drop-in replacement for the standard FLUX.2 decoder on Hugging Face. It decodes about 1.4x faster and uses about 1.4x less VRAM at decode time, with ~28M decoder parameters versus ~50M in the full decoder and minimal quality loss. It is available under the Apache 2.0 license and is compatible with FLUX.2-klein-4B, FLUX.2-klein-9B, FLUX.2-klein-9b-kv, and FLUX.2-dev.

Mar 22

Mar 22Sun
  1. FunAudioLLM (Alibaba Tongyi) · new models on Hugging FaceAI score32

    PrismAudio Adds Reinforcement Learning to Video-to-Audio Generation with Chain-of-Thought Planning

    AIPrismAudio is a framework that integrates reinforcement learning into video-to-audio generation, using a Chain-of-Thought planning mechanism. It builds on ThinkSound by splitting single-step reasoning into four CoT modules for semantic, temporal, aesthetic, and spatial dimensions, each with targeted reward functions. Code, model weights, and datasets are released for research and educational use under the MIT License, and commercial use requires explicit author authorization.

Mar 17

Mar 17Tue
  1. Apple · new models on Hugging FaceAI score43

    Apple releases SimpleSD-4B-thinking, a self-distilled Qwen model for code generation

    AIApple has published SimpleSD-4B-thinking on Hugging Face, a research checkpoint built on Qwen that improves code generation through Simple Self-Distillation without rewards, verifiers, teacher models, or reinforcement learning. On LiveCodeBench, it lifts Qwen3-4B-Thinking-2507 from 54.5% to 57.8% pass@1 on LCBv6 and from 59.6% to 63.1% pass@1 on LCBv5. The model is released as a reproducibility checkpoint under the Apple Machine Learning Research Model License, not as an optimized Qwen release.

Mar 12

Mar 12Thu
  1. Intern Large ModelsAI score47

    InternVL-U: Open-Source 4B Unified Model for Reasoning, Generation, and Editing

    AIInternVL-U is a lightweight 4B unified multimodal model that combines reasoning, generation, and editing in one framework, according to Intern Large Models. The post says it uses unified contextual modeling, modality-specific modular design, and decoupled visual representations to balance performance and efficiency. It reportedly outperforms unified baselines more than 3× its size on text rendering, scientific reasoning, and spatially grounded generation and editing, and is open-source on GitHub and Hugging Face.

    Image from @intern_lm's post

Feb 19

Feb 19Thu
  1. Guillaume Lample @ NeurIPS 2024AI score40

    Mistral releases Voxtral Realtime paper, Apache 2.0 speech model

    AIMistral has published the technical report for Voxtral Realtime, a speech transcription model released under the Apache 2.0 license. The model reportedly achieves state-of-the-art transcription performance at sub-500ms latency. Mistral also launched a Realtime playground in Mistral Studio and made the model available in Hugging Face Transformers.

    Image from @GuillaumeLample's post

Jan 21

Jan 21Wed
  1. Mistral AI · new models on Hugging FaceAI score65

    Mistral releases open-weight Voxtral Mini 4B Realtime 2602 speech model

    AIMistral AI released Voxtral Mini 4B Realtime 2602, a multilingual realtime speech-transcription model with 13 supported languages under the Apache 2.0 license. The model has a configurable transcription delay from 240ms to 2.4s, and it matches leading offline open-source models at a 480ms delay. The source says it is optimized for on-device deployment and is currently supported only in vLLM.

    Why it matters: The source specifies the 480ms delay operating point, 4B size, Apache 2.0 license, and vLLM serving path, which matter for teams weighing realtime transcription deployment.

Jan 14

Jan 14Wed
  1. Black Forest Labs · new models on Hugging FaceAI score62

    Black Forest Labs releases FLUX.2 [klein] 4B image model under Apache 2.0

    AIBlack Forest Labs released FLUX.2 [klein] 4B, a 4 billion parameter model that unifies text-to-image generation and image editing with multi-reference support. The source says it runs on consumer GPUs such as the RTX 3090 or 4070 with about 13GB VRAM, and its open weights are available under the Apache 2.0 license.

    Why it matters: The source specifies a 4 billion parameter model running on about 13GB VRAM under Apache 2.0, which helps readers judge whether local image generation fits their hardware.

  2. Black Forest Labs · new models on Hugging FaceAI score54

    Black Forest Labs releases FLUX.2 [klein] 4B Base on Hugging Face

    AIBlack Forest Labs has published FLUX.2 [klein] 4B Base, a 4 billion parameter text-to-image model that also supports multi-reference editing. The model is undistilled, is released with open weights under Apache 2.0, and is described as fitting in about 13GB VRAM on cards such as the RTX 3090 or 4070, with reference code available in its GitHub repository and support in ComfyUI and Diffusers.

  3. Black Forest Labs · new models on Hugging FaceAI score46

    FLUX.2 [klein] 9B Base Released on Hugging Face as Undistilled Open-Weight Model

    AIBlack Forest Labs has released FLUX.2 [klein] 9B Base, a 9 billion parameter undistilled rectified flow transformer with open weights for text-to-image generation and multi-reference editing. The model is intended for fine-tuning, LoRA training, and research, and fits in about 29GB VRAM on NVIDIA RTX 4090-class GPUs. A reference implementation is available on GitHub, and the model works with ComfyUI and Diffusers.

Dec 22, 2025

Dec 22, 2025Mon
  1. FunAudioLLM (Alibaba Tongyi) · new models on Hugging FaceAI score58

    Alibaba's FunAudioLLM releases Fun-Audio-Chat-8B for low-latency voice interaction

    AIFunAudioLLM has released Fun-Audio-Chat-8B, a roughly 8B-parameter large audio language model for natural, low-latency voice interaction, under Apache 2.0. It uses Dual-Resolution Speech Representations with a 5Hz frame rate, which the source says reduces GPU hours by nearly 50%, and it supports English and Chinese.

Dec 11, 2025

Dec 11, 2025Thu
  1. OpenAI · new models on Hugging FaceAI score42

    OpenAI Releases circuit-sparsity Sparse Model Weights on Hugging Face

    AIOpenAI has published weights for a sparse model from Gao et al. 2025, used for qualitative results on bracket counting and variable binding, on Hugging Face under the openai/circuit-sparsity repository. The release includes a standalone Hugging Face implementation that loads the converted model and tokenizer with trust_remote_code and runs sample generation. The project is licensed under Apache License 2.0.