Skip to content

#Image generation

Oct 8

TodayOct 8Thu2 items
  1. Artificial AnalysisAI score38

    Grok Imagine Video 1.5 Lite leads in architecture, consumer, and knowledge-work use cases

    Artificial Analysis reports that Grok Imagine Video 1.5 Lite comes closest to the frontier in Architecture & Real Estate, Consumer, and Productivity & Knowledge Work use cases. It sits furthest from the frontier in Live-Action Film and Frontier use cases. Against Grok Imagine Video 1.5, Lite matches it in Social Media & Creator Content and trails it on the other nine use cases.

  2. Artificial AnalysisAI score31

    Grok Imagine Video 1.5 Lite sits on the quality and speed frontier on AA-Video-T2V-Silent v2.0 Among the 12 models on AA-Video-T2V-Silent v2.0 that we benchmark for generation speed, no model is both faster and higher quality than Grok Imagine Video 1.5 Lite. It generates a 10 second 1080p clip in a median of 60.5 seconds. Kling 3.0 1080p (Pro) scores slightly higher and takes 94 seconds for a 5 second clip. Vidu Q3 Turbo is 9 seconds faster on a 5 second 720p clip, and scores well below it.

    Grok Imagine Video 1.5 Lite sits on the quality and speed frontier on AA-Video-T2V-Silent v2.0 Among the 12 models on AA-Video-T2V-Silent v2.0 that we benchmark for generation speed, no model is both faster and higher quality than Grok Imagine Video 1.5 Lite. It generates a 10 second 1080p clip in a median of 60.5 seconds. Kling 3.0 1080p (Pro) scores slightly higher and takes 94 seconds for a 5 second clip. Vidu Q3 Turbo is 9 seconds faster on a 5 second 720p clip, and scores well below it.

Oct 7

Oct 7Wed

Oct 6

Oct 6Tue
  1. Google AI StudioAI score45

    introducing @NanoBanana 2.1 🍌 our latest image generation model outperforms our previous models across the board: - better visual design - mask-based editing - subject consistency - more natural-looking images try it out today in https://ai.studio

    introducing @NanoBanana 2.1 🍌 our latest image generation model outperforms our previous models across the board: - better visual design - mask-based editing - subject consistency - more natural-looking images try it out today in https://ai.studio

  2. Gemini API ChangelogAI score58

    Google releases Gemini Nano Banana 2.1 for general availability

    Google has made Gemini Nano Banana 2.1, identified as gemini-nano-banana-2.1, generally available as an image generation and conversational editing model. It improves visual quality, prompt adherence, multi-turn character consistency, and text rendering, and adds panoramic aspect ratios such as 1:4, 4:1, 1:8, and 8:1 at 1K, 2K, and 4K resolutions. The gemini-3.1-flash-image model is deprecated with no shutdown date announced, and developers are told to migrate to the new model.

Oct 5

Oct 5Mon
  1. Liquid AI · new models on Hugging FaceAI score44

    LiquidAI releases d1-omni-600M, a 600M decision model for text, image and audio

    LiquidAI has released d1-omni-600M on Hugging Face, a 587M-parameter model that answers named yes/no, choice and score questions over text, images or up to 30 seconds of speech in a single forward pass. It returns typed answers with zero output tokens by reading the model's distribution over options, and is built on LFM2.5-Encoder-350M with a 16,384-token context length. The model is not a chat model and does not generate text.

Oct 1

Oct 1Thu
  1. NVIDIA · new models on Hugging FaceAI score44

    NVIDIA releases PixelUMM, an encoder-free model for pixel-space image and video tasks

    NVIDIA has released PixelUMM, an encoder-free unified multimodal model with 15,199,672,064 parameters that handles text, image, and video understanding and generation directly in pixel space. It represents images as 16-by-16 RGB pixel patches on a Qwen3-8B language backbone, with iterative denoising for generation. The checkpoint is licensed for non-commercial research or evaluation only, while the source code is under Apache License 2.0.

  2. ReplicateAI score47

    FLUX 3 Image is here. The latest from @bfl_ai, generate stunning native 4K imagery, and get maximum control over every pixel. Construct images with hyper-specific layouts using bounding boxes, make multiple targeted edits at once, and stay consistent across edits. For its first week, it's 50% off.

    FLUX 3 Image is here. The latest from @bfl_ai, generate stunning native 4K imagery, and get maximum control over every pixel. Construct images with hyper-specific layouts using bounding boxes, make multiple targeted edits at once, and stay consistent across edits. For its first week, it's 50% off.

Sep 30

Sep 30Wed
  1. IdeogramAI score46

    Ideogram 4.5 comes in four quality modes, ranging from 0.8¢ to 22¢ per image, all at native 2K resolution. Available now on our launch partners: @superscale_ai @Picsart @cfabricacom @luminal_ai @arena @runware @florafaunaai @krea_ai @trymoda @LeonardoAi @runwayml @pika_labs @fal @ComfyUI @magnific @GammaApp @LumaLabsAI @DesignArena

    Ideogram 4.5 comes in four quality modes, ranging from 0.8¢ to 22¢ per image, all at native 2K resolution. Available now on our launch partners: @superscale_ai @Picsart @cfabricacom @luminal_ai @arena @runware @florafaunaai @krea_ai @trymoda @LeonardoAi @runwayml @pika_labs @fal @ComfyUI @magnific @GammaApp @LumaLabsAI @DesignArena

  2. IdeogramAI score22

    Here's a side-by-side comparison of the same edits on Ideogram 4.5 vs. GPT Image 2.5 Sunburst, Nano Banana Pro, and Nano Banana 2. GPT Image and Nano Banana outputs become unusable within a few edits, while Ideogram 4.5 stays clean edit after edit.

    Here's a side-by-side comparison of the same edits on Ideogram 4.5 vs. GPT Image 2.5 Sunburst, Nano Banana Pro, and Nano Banana 2. GPT Image and Nano Banana outputs become unusable within a few edits, while Ideogram 4.5 stays clean edit after edit.

  3. IdeogramAI score38

    Introducing Ideogram 4.5, the most precise edit model. With each edit, leading models add artifacts, pixel shifts, and color changes. Ideogram 4.5 eliminates artifact buildup, making multi-turn editing possible. Live in Ideogram, the API, and launch partners. Open weights soon.

    Introducing Ideogram 4.5, the most precise edit model. With each edit, leading models add artifacts, pixel shifts, and color changes. Ideogram 4.5 eliminates artifact buildup, making multi-turn editing possible. Live in Ideogram, the API, and launch partners. Open weights soon.

Sep 24

Sep 24Thu
  1. ModelScopeAI score38

    Qwen-Image-2.1-Fun-Controlnet-Union adds eight controls and inpainting

    ModelScope released Qwen-Image-2.1-Fun-Controlnet-Union, a single checkpoint adding eight structural controls, including Canny, Depth, Pose, and Scribble, plus inpainting to Qwen-Image 2.1. Control and inpainting share one branch with 16 injection points across every second Transformer block, keeping the base model frozen and requiring no checkpoint switching. It runs at guidance scale 1.0 with CFG-distilled sampling and prefix KV caching, and is available under the Qwen Research License with base Qwen-Image 2.1 weights required.

Sep 23

Sep 23Wed
  1. KrASIA · Big TechAI score46

    Tencent Hy Image 3.5 preview refined through its consumer and business products

    Tencent has released a preview of its Hy Image 3.5 image generation model, which product teams across Yuanbao, WorkRally, Ima, and other services are helping refine through co-design. Tencent Cloud prices the model at USD 0.024 per 2K output image, and it supports text-to-image and image-to-image generation with up to five reference images. Tencent said an internal blind evaluation found it on par with ByteDance's Seedream 5.0 Pro and slightly better than Nano-Banana Pro and Qwen-Image-3.0 Pro.

Sep 22

Sep 22Tue

Sep 20

Sep 20Sun
  1. ModelScopeAI score62

    Qwen-Image-2.1 unifies image generation and editing with native transparency

    Alibaba's ModelScope introduces Qwen-Image-2.1, a model that handles image generation and editing together, with native transparency and a compact 7B visual generation component. It adds KV cache reuse to speed up generation and editing while reducing memory use, especially with multiple reference images. The model can combine up to 10 reference images, make targeted local edits, and preserve portrait identity and product details.

  2. Qwen · new models on Hugging FaceAI score62

    Qwen releases Qwen-Image-2.1 prompt rewriter for image editing on Hugging Face

    Qwen has open-sourced Qwen-Image-2.1, a unified text-to-image generation and image editing model with 7B visual generation parameters. The Hugging Face page for Qwen-Image-2.1-PE-I2I is a fine-tuned Qwen3.5-VL 9B prompt rewriter that turns vague editing instructions and input images into precise editing prompts, supporting up to 10 reference images.

    AIWhy it matters: The model card documents usage with transformers and diffusers, letting readers see how the editing prompt rewriter connects to the generation pipeline.

  3. Qwen · new models on Hugging FaceAI score62

    Qwen releases open-source Qwen-Image-2.1 with a prompt rewriting model

    Qwen has open-sourced Qwen-Image-2.1, a unified text-to-image generation and image editing model with a 7B-parameter visual generation component. The release also includes Qwen-Image-2.1-PE-T2I, a fine-tuned Qwen3.5-VL 9B model that rewrites brief image requests in any language into detailed English prompts with a recommended aspect ratio.

    AIWhy it matters: The release pairs a 7B visual generation component with a separate prompt rewriting model, showing how a brief image request becomes a detailed English prompt before rendering.

Sep 17

Sep 17Thu
  1. inclusionAI (Ant Ling) · new models on Hugging FaceAI score46

    Ming-Image-0.1-Design-Layer splits flattened design images into RGBA layers

    inclusionAI has released Ming-Image-0.1-Design-Layer on Hugging Face, a model that decomposes a flattened design image into a requested number of RGBA layers using an image and a layer plan. The model runs at 1024 resolution (512 for faster processing) with 12 sampling steps, a CFG scale of 2.0, and BF16 precision on one CUDA GPU with 80 GiB VRAM. It is released under the MIT License.

  2. inclusionAI (Ant Ling) · new models on Hugging FaceAI score42

    inclusionAI releases Ming-Image-0.1-Design, a 6B text-to-image model for text-rich designs

    inclusionAI has released Ming-Image-0.1-Design, a 6B text-to-image model for UI, infographics, and posters that outputs RGBA images with transparent backgrounds. The model is available on Hugging Face and ModelScope under the MIT License. It runs at 2048 x 2048 with 12 sampling steps and a CFG scale of 1.0, validated on one CUDA GPU with 80 GiB VRAM.

Sep 14

Sep 14Mon
  1. Sherwin WuAI score40

    Don't forget that GPT Image 2.5 is now live in ChatGPT! It's incredible at keeping consistency while editing images, and SOTA on all image leaderboards. If you've been burned by faces slightly changing when editing with GPT Image 2, give it another try! https://openai.com/index/introducing-chatgpt-images-2-5/

    Don't forget that GPT Image 2.5 is now live in ChatGPT! It's incredible at keeping consistency while editing images, and SOTA on all image leaderboards. If you've been burned by faces slightly changing when editing with GPT Image 2, give it another try! https://openai.com/index/introducing-chatgpt-images-2-5/

Sep 13

Sep 13Sun
  1. Qwen · new models on Hugging FaceAI score67

    Qwen releases open-source Qwen-Image-2.1 for generation and editing

    Qwen has open-sourced Qwen-Image-2.1, a unified text-to-image generation and image editing model with 7B parameters in its visual generation component. The model can generate regular or transparent RGBA images, supports up to 10 reference images for editing, and is licensed under the Qwen Research License Agreement.

    AIWhy it matters: The source specifies the 7B visual component, transparent RGBA output, and up to 10 reference images, which helps readers judge its fit for generation and editing workflows.

Sep 4

Sep 4Fri
  1. Mustafa SuleymanAI score21

    Our new image model generates images 2x faster than GPT-Image-2, currently the best model in the world. It's also 72% more efficient in GPU usage, so we can provide it at an incredible price. This gives it the best price-performance score in the world. Unbelievable work from the team. So much more to come! Try MAI-Image-2.6-Flash out now!

    Our new image model generates images 2x faster than GPT-Image-2, currently the best model in the world. It's also 72% more efficient in GPU usage, so we can provide it at an incredible price. This gives it the best price-performance score in the world. Unbelievable work from the team. So much more to come! Try MAI-Image-2.6-Flash out now!

  2. Google AI StudioAI score44

    Lyria 3.5, our best-sounding music generation model, is now available in AI Studio, via the Gemini API, and in the Gemini app this model brings more expressive vocals and richer musical arrangements, allowing you to craft tracks with higher fidelity try it today: http://ai.studio

    Lyria 3.5, our best-sounding music generation model, is now available in AI Studio, via the Gemini API, and in the Gemini app this model brings more expressive vocals and richer musical arrangements, allowing you to craft tracks with higher fidelity try it today: http://ai.studio