Skip to contentSkip to stories

Updated

#Image generation

Items with an AI score under 20 are hidden. Show low-relevance items

Aug 28

Aug 28Fri

Aug 27

Aug 27Thu
  1. RadixArkAI score34

    RadixArk adds LoRA SFT to Miles-diffusion for targeted post-training

    AIRadixArk introduced LoRA SFT in Miles-diffusion for fast, targeted post-training of diffusion models. The company trained a rank-64 LoRA adapter for MiniMax H3 to improve physical realism, using 254 curated training windows and under 3 hours on 8 GPUs. The adapter can be exported to safetensors and served directly with SGLang without retraining the full model.

Aug 26

Aug 26Wed

Aug 21

Aug 21Fri

Aug 19

Aug 19Wed
  1. Google · new models on Hugging FaceAI score22

    Google releases TIPS B/14 v1 vision-language model on Hugging Face

    AIGoogle has published TIPS B/14 (v1) on Hugging Face, a contrastive vision-language model that produces spatially rich image features aligned with text embeddings. The model has 86M vision parameters and 110M text parameters at native 448 resolution, and is licensed under Apache 2.0. The release includes usage code for image and text encoding, zero-shot classification, and spatial feature visualization.

Aug 18

Aug 18Tue
  1. Stability AIAI score38

    Stability AI launches Stable Audio 3.0 plugin and enhanced web experience

    AIStability AI has released two beta tools for Stable Audio 3.0: a plugin that brings audio generation into digital audio workstations (DAWs) and an upgraded experience at StableAudio.com with more editing options. Both are powered by commercially-safe models, so users own their outputs and can distribute them freely. Some features are experimental, and the company says it will keep iterating in real time.

  2. Stability AIAI score40

    Stability AI launches Stable Audio plugin and enhanced web app for Stable Audio 3.0

    AIStability AI released a Stable Audio plugin that runs Stable Audio 3.0 generation inside DAWs as an instrument, available as a macOS AU and VST3 with Apple Silicon and Intel support. The enhanced StableAudio.com web app adds iterative prompting, audio-to-audio variations, per-track mixing controls, and export, and both tools are in beta and powered by commercially-safe models that users can distribute freely.

Aug 17

Aug 17Mon

Aug 13

Aug 13Thu
  1. ByteDance · new models on Hugging FaceAI score52

    ByteDance releases Bernini-Diffusers-v2 video generation and editing model

    AIByteDance has released Bernini-Diffusers-v2 on Hugging Face, a video generation and editing pipeline combining a Qwen2.5-VL planner with Wan2.2 diffusion components. The model card recommends it over Bernini-R for complex requests needing stronger instruction following and multi-step semantic planning. Code and weights are available under Apache License 2.0.

Aug 11

Aug 11Tue

Aug 4

Aug 4Tue

Jul 30

Jul 30Thu
  1. MiniMax BlogAI score72

    MiniMax H3 unifies text, image, video, and audio generation in one model

    AIMiniMax launches H3, a general-purpose multimodal generation model that understands text, images, video, and audio as unified context. It generates video up to 15 seconds at 2K resolution with native stereo sound, and the company says model weights will be opened in the coming days, subject to applicable laws and regulations. MiniMax also says H3 is priced below mainstream models at 2K and 768p.

    Why it matters: The post explains how a unified multimodal design and training choices enable 2K video with native stereo sound, useful for comparing against closed video generators.

Jul 29

Jul 29Wed

Jul 24

Jul 24Fri

Jul 7

Jul 7Tue
  1. Meta AI BlogAI score75

    Meta launches Muse Image, an agentic image model with search and code tools

    AIMeta Superintelligence Labs has released Muse Image, which can invoke search and coding tools and self-refine its generations before output. It is available today in the Meta AI app, meta.ai, Instagram Stories in the US, and WhatsApp in limited countries, with Facebook coming soon. Meta also previewed Muse Video, which is coming soon to creators and Meta AI and is reported as ranking No. 3 on Arena for text-to-video at the time of writing.

    Why it matters: The source describes how search, code execution, and self-refinement change image generation, which matters to anyone comparing agentic media models with plain prompt-to-image systems.

Jul 1

Jul 1Wed

Jun 30

Jun 30Tue

Jun 25

Jun 25Thu

Jun 18

Jun 18Thu

Jun 17

Jun 17Wed

Jun 15

Jun 15Mon
  1. ByteDance · new models on Hugging FaceAI score24

    Sa2VA-LLaVA-1.5-7B: ByteDance's SAM2-Grounded Segmentation and Chat Model

    AIByteDance has released Sa2VA-LLaVA-1.5-7B on Hugging Face, a model built on LLaVA-1.5-7B with a SAM2 grounding encoder that performs dense image and video referring segmentation alongside open-ended chat. The checkpoint is self-contained and loads with trust_remote_code=True without extra packages, and it is positioned as a LISA-comparable baseline within the Sa2VA family. Reported results include 80.3 cIoU on RefCOCO val and 54.8 J&F on MeViS (val_u).

Jun 9

Jun 9Tue
  1. ByteDance · new models on Hugging FaceAI score28

    ByteDance releases Sa2VA-Qwen3-VL-4B-SAM3 for image and video referring segmentation

    AIByteDance's Sa2VA-Qwen3-VL-4B-SAM3 is built on Qwen3-VL-4B-Instruct with a SAM3 grounding encoder and produces dense image and video referring segmentation alongside chat. It reports 83.7 cIoU on RefCOCO val, 65.3 J&F on MeViS (val_u), and 77.1 on Ref-DAVIS17. The checkpoint is self-contained and loads on Hugging Face with trust_remote_code=True, with no extra packages required.

  2. Z.ai (GLM) · new models on Hugging FaceAI score52

    Z.ai releases SCAIL-2, an open-source end-to-end character animation model

    AIZ.ai released SCAIL-2, an open-source model that animates a reference character from a driving video without skeleton maps or inpainting masks. It also supports character replacement, multi-character scenes, and animal-driving, with 512p and 704p resolutions and inputs whose height and width are both divisible by 32.

Jun 8

Jun 8Mon
  1. ByteDance · new models on Hugging FaceAI score46

    ByteDance Open-Sources Bernini-R 1.3B Video Diffusion Renderer on Hugging Face

    AIByteDance has open-sourced the 1.3B-parameter weights of its Bernini Renderer (Bernini-R), available on Hugging Face as ByteDance/Bernini-R-1.3B-Diffusers. Fine-tuned from Wan2.1-1.3B, the model performs close to the 14B variant on simple tasks such as style transfer, subtitle or watermark removal, and local editing, but lags on complex tasks such as human generation. The release requires a CUDA GPU, with an H100 recommended for FlashAttention-3.

Jun 2

Jun 2Tue
  1. ByteDance · new models on Hugging FaceAI score44

    ByteDance Releases Bernini-R Diffusers Weights for Video Generation and Editing

    AIByteDance has open-sourced the inference code and model weights of the Bernini Renderer (Bernini-R), a DiT-based renderer paired with an MLLM-based semantic planner for video generation and editing. A diffusers-format version, ByteDance/Bernini-R-Diffusers, bundles the Wan2.2 base components with the Bernini-R transformer weights for direct loading, and the framework requires a CUDA GPU with PyTorch 2.5.1+cu124.

Jun 1

Jun 1Mon
  1. PaddlePaddleAI score36

    PaddleOCR and ERNIE Image now available as official Dify plugins

    AIPaddleOCR and ERNIE Image are now available as official Dify plugins, bringing document parsing and image generation into Dify's agent workflows. PaddleOCR, powered by PP-OCRv5, PP-StructureV3, and PaddleOCR-VL, turns images, scanned PDFs, and multilingual documents into structured data for chunking, vectorization, and RAG, with private or on-prem deployment supported. ERNIE Image offers free generation, a Turbo mode with 8-step inference, and an OpenAI-style API.

May 29

May 29Fri