Skip to contentSkip to stories

Updated

#Multimodal

Showing low-relevance items too. Hide low-relevance items

Sep 20

Sep 20Sun
  1. ModelScopeOfficialAI score62

    Qwen-Image-2.1 unifies image generation and editing with native transparency

    AIAlibaba's ModelScope introduces Qwen-Image-2.1, a model that handles image generation and editing together, with native transparency and a compact 7B visual generation component. It adds KV cache reuse to speed up generation and editing while reducing memory use, especially with multiple reference images. The model can combine up to 10 reference images, make targeted local edits, and preserve portrait identity and product details.

    Image from @ModelScope2022's post
  2. Qwen · new models on Hugging FaceOfficialAI score62

    Qwen releases Qwen-Image-2.1 prompt rewriter for image editing on Hugging Face

    AIQwen has open-sourced Qwen-Image-2.1, a unified text-to-image generation and image editing model with 7B visual generation parameters. The Hugging Face page for Qwen-Image-2.1-PE-I2I is a fine-tuned Qwen3.5-VL 9B prompt rewriter that turns vague editing instructions and input images into precise editing prompts, supporting up to 10 reference images.

    Why it matters: The model card documents usage with transformers and diffusers, letting readers see how the editing prompt rewriter connects to the generation pipeline.

Sep 19

Sep 19Sat
  1. StepFunOfficialAI score62

    StepFun Launches Step 5 Preview, a 600B MoE Model for Agentic Work

    AIStepFun has released Step 5 Preview, a flagship model for agentic work that it says delivers frontier-level performance in software engineering and professional knowledge work, with particular strength in finance. The model is a 600B total, 27B active mixture-of-experts design with a 1M context window and vision support. StepFun says it offers substantially lower task cost at comparable intelligence, and open weights are scheduled for October 15.

    Image from @StepFun_ai's post

Sep 18

Sep 18Fri
  1. One Useful Thing (Ethan Mollick)BlogAI score50

    Mollick says AI already does weeks of human work when guided, citing Zork and Eco library demos

    AIEthan Mollick says GPT-6 Astra and Fable 5.1 already enable transformative impact and can reliably handle weeks of human work when properly guided. He cites GPT-6 Astra turning the 1977 text adventure Zork into a 3D action-adventure game and Fable 5.1 reconstructing Umberto Eco's Milan library in 3D from videos, photos, and catalogues.

  2. Mustafa SuleymanXAI score22

    Mustafa Suleyman claims best image generation quality-price performance

    AIMustafa Suleyman, owner of the source account associated with Microsoft and Copilot, says the post claims the best image generation quality-price performance in the world. The post itself gives no specific model, price, or benchmark figures. Background from Artificial Analysis says Muse Image, MAI-Image-2.6, and GPT Images 2.5 recently shifted text-to-image price and speed frontiers.

  3. MiniMax Design (H3)OfficialAI score18

    Hailuo AI speeds up storyboarding with 3x3 panel-to-video generation

    AIHailuo AI's post suggests that a single 3x3 storyboard image can be converted into a video using the MiniMax H3 Max r2v model at 480p for 15 seconds. The quoted post describes settings with Quality prompt tuning and standard reference strength, and asks for a 2D animation with panel-to-panel cuts while excluding multiple panels and BGM.

  4. Liquid AI · new models on Hugging FaceOfficialAI score55

    Liquid AI releases LFM2.5-VL-3B-DSpark drafter for faster vision-language decoding

    AILiquid AI released LFM2.5-VL-3B-DSpark, a speculative-decoding draft model for its LFM2.5-VL-3B vision-language model. The source reports decoding up to 2.66× faster on a single H100 with SGLang, up to 3.13× on Apple M5 Max with MLX-VLM, and up to 2.14× on Apple M3 Ultra with llama.cpp, with output unchanged under greedy decoding.

Sep 17

Sep 17Thu
  1. WanOfficialAI score13

    Qwen's Wan3.0 creator advice: tell emotional stories, not chase visuals

    AIAlibaba Wan's post advises beginners making AI video to design a viewing experience and tell an emotion-triggering story rather than chasing the prettiest frame. It notes a single 30-second shot can be enough, citing the filmmaker behind Soulscape and Johnny Mai from Alibaba Cloud. The post promotes bringing Wan3.0 to teams via a sign-up form.

    Video from @Alibaba_Wan's post
  2. vLLM BlogOfficialAI score38

    vLLM Adds NVIDIA Hardware Video Decoding to Scale Multi-GPU Video Captioning

    AIvLLM now supports NVIDIA hardware video decoding through PyNvVideoCodec, moving video decoding off the CPU so multi-GPU video captioning can scale to 8 GPUs. In benchmarks on 8xH100 GPUs, GPU-based decoding provides more than double the throughput of the CPU-based decoder for Qwen/Qwen3-VL-8B-Instruct with 8 single-GPU vLLM replicas. The functionality is included in standard CUDA vLLM releases, and PyNvVideoCodec==2.0.4 is required for custom installations.

  3. Mike KriegerXAI score12

    Anthropic's Krieger builds interactive map of Iron Tangle level

    AIMike Krieger, Anthropic's account owner, used an interactive Claude artifact to visualize the Iron Tangle level from Dungeon Crawler Carl, saying it helped him finally understand the layout. The post shares a link to the artifact, with no further details about its features.

  4. Gemini NotebookOfficialAI score37

    Mariposa Museum exhibit shows town in 1859, a decade after Gold Rush

    AIA Mariposa Museum exhibit photographed by writer Steven Johnson, shared by Gemini Notebook, documents the Sierra Nevada town in 1859, ten years after the Gold Rush began. Johnson says he used the Gemini Notebook mobile app's camera feature to generate a detailed report from photos of the display, which he says was 99% accurate on fact-checking.

  5. SenseTimeOfficialAI score44

    SenseNova U1.5 open-sources 8B unified model for understanding and generation

    AISenseTime released its SenseNova U1.5 technical report, describing an open-source 8B native MoT unified model that connects understanding and generation through shared attention. The model reports 68.2% on VBVR-Pro-Bench, ahead of Nano-Banana-Pro (56.4%) and GPT-Image-2 (50.7%), and its full training recipes, including SFT, RL, and multi-expert on-policy distillation, are open-sourced.

    Image from @SenseTime_AI's post
  6. KrASIA · Big TechNewsAI score50

    SenseTime's Lin Dahua Says Multimodal AI Breakthrough Could Come Within Two Years

    AISenseTime chief scientist Lin Dahua argues that native multimodal AI, which processes language, vision and other information in one shared model, is essential for AI to move beyond coding into industries and the physical world. SenseTime released the open-source SenseNova U1 in April and U1.5 Lite nearly four months later, and reported first-half 2026 revenue of RMB 2.91 billion, up 23.4% year-on-year. Lin's claim that a breakthrough could come within two years is the source's prediction, not a confirmed result.

Sep 16

Sep 16Wed
  1. Google for DevelopersOfficialAI score38

    Three companies use Gemini agentic video understanding to cut token costs

    AIMosaic, Ponder Studio, and Revyl used early access to Google's Gemini Flash models to test agentic video understanding on long footage. Mosaic reports a 97% cut in median token usage and nearly double the ability to handle complex edits, while Ponder Studio reports a 0.967 F1 score and about 72% lower token costs for B-roll selection. Revyl says the approach improved mobile UI bug-catching accuracy by 65%. The capability is available now for video uploads and YouTube videos via the Gemini API.

  2. catXAI score60

    Claude merges Cowork and chat into one product with automatic routing

    AIAnthropic is merging Claude Cowork and chat into one Claude, and Claude Design is integrated so users can ask for slides, designs, or docs without switching apps. Claude decides from the prompt whether to give a quick answer or do deeper agentic work, and users can still stop, redirect, or adjust its effort. The change rolls out to Pro and Max over the next few weeks.

  3. BAAI · new models on Hugging FaceOfficialAI score34

    BAAI and Peking University release Brainμ-Spike spike camera image reconstruction model

    AIPeking University's Yu Zhaofei team and the Beijing Academy of Artificial Intelligence (BAAI) released Brainμ-Spike, a small convolutional network for spike camera image reconstruction that is paired with the Brainμ model. The package includes weights, inference scripts, and evaluation tools, but the base large model and LoRA weights are not yet released, so the full generation pipeline cannot run from this repository alone.

  4. inclusionAI (Ant Ling) · new models on Hugging FaceOfficialAI score55

    inclusionAI releases Realtime-Venus full-duplex audio-visual models on Hugging Face

    AIinclusionAI has published Realtime-Venus on Hugging Face with two 9B checkpoints: Realtime-Venus-Omni for audio-visual interaction and Realtime-Venus-Audio for audio-only conversation. Both are built on MiniCPM-o 4.5 with a Qwen3-8B backbone and support full-duplex dialogue, proactive responses, and training-free long-video memory. The asynchronous Realtime-Venus-Harness runtime is hosted in a separate GitHub repository.

Sep 15

Sep 15Tue
  1. Tencent · new models on Hugging FaceOfficialAI score37

    Tencent Releases WeVisDoc-2B and WeVisDoc-4B Document Parsing Models on Hugging Face

    AITencent's WeVisDoc-4B, fine-tuned from Qwen3-VL-4B-Instruct, scores 95.38 Overall on OmniDocBench v1.6 and 75.54 mean Overall across three PureDocBench tracks. The end-to-end parser converts page images into structured Markdown with LaTeX formulas and HTML tables, and the 2B variant is also available. The repository provides vLLM serving scripts with a 32768-token default context and a Python client for batch processing.

  2. Jazzyear · InsightsNewsAI score67

    HiDream's vivago R1 agent targets five-minute AI video delivery

    AIHiDream.ai launched vivago R1, a content creation agent, globally, with a domestic version upgrade. The company says R1 can output five-minute high-quality videos through agent planning, with a claimed 85% usable-output rate and support for multi-round extensions. It also released HiDream-O1-Video-1.0, a native omni-modal video model supporting single shots of 5 to 20 seconds at 1080p.

  3. Google AI StudioOfficialAI score72

    Google releases Gemini 3.8 Live and 3.5 Transcribe for real-time voice apps

    AIGoogle AI Studio released Gemini 3.8 Live, a native speech-to-speech model with an Extended Thinking variant, and made it available through the Live API. Gemini 3.5 Transcribe, released last month, supports 85+ languages with a reported 4.0% streaming and 2.6% non-streaming Word Error Rate, and accepts a custom vocabulary of up to 1,000 terms. Live API audio pricing is listed at $0.005/min for input and $0.018/min for output.

    Why it matters: The post lists concrete Live API capabilities, per-minute audio pricing, and transcription accuracy figures, helping developers weigh voice agent options against their own cascaded pipelines.

  4. Josh WoodwardXAI score31

    Gemini Notebook adds spoken Q&A and lecture audio notes for students

    AIGoogle's Gemini Notebook now offers live spoken Q&A over class materials in about 100 languages, and lets students record lectures on the go with audio notes saved automatically to a chosen notebook. University students in 140+ countries can also still get a free Google AI Plan for higher limits and access to more Google products.

  5. Google AI StudioOfficialAI score46

    Google launches Gemini 3.8 Live and Extended Thinking dialogue models

    AIGoogle introduced two live dialogue models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, available through AI Studio and the Gemini API. Gemini 3.8 Live is built for scale and cost efficiency, combining conversational intelligence with fluid dialogue and visual grounding. The Extended Thinking variant targets high-complexity tasks with increased intelligence and multi-step reasoning.

    Video from @GoogleAIStudio's post
  6. Google AIOfficialAI score62

    Google releases Gemini 3.8 Live and 3.8 Live Extended Thinking audio models

    AIGoogle AI announces Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking as its most advanced Gemini Audio models. Gemini 3.8 Live is built for scale, speed, and cost efficiency, handling mid-sentence interruptions, transitions across 97 languages, and visual context through Search Live. Gemini 3.8 Live Extended Thinking reasons and speaks in parallel, narrating its progress on multi-step tasks such as event planning.

    Video from @GoogleAI's post
  7. Google DeepMindOfficialAI score72

    Google DeepMind releases Gemini 3.8 Live models for real-time voice agents

    AIGoogle DeepMind introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two live dialogue models for voice agents. Extended Thinking scores 82.6 on Artificial Analysis' Speech to Speech Quality Index, 68.6% on τ-Voice, and 97.7% on Big Bench Audio. Gemini 3.8 Live is rolling out now in the Gemini API, Google AI Studio, and Search Live, with enterprise access in private preview.

    Why it matters: The release covers a voice model's benchmark results and availability across developer, enterprise, and consumer products, useful for judging voice agent options.

  8. NVIDIA · new models on Hugging FaceOfficialAI score34

    NVIDIA Releases RT-DETR Hand Detection v1.0 for Real-Time RGB Hand Localization

    AINVIDIA's RT-DETR Hand Detection v1.0 detects and localizes left and right hands in RGB images, outputting 2D bounding boxes with per-hand confidence scores in a single pass. The model, built on RT-DETRv2-S with HGNetv2-S backbone and about 20M parameters, is intended as a region-of-interest stage for downstream 3D hand pose estimation and is exported to ONNX. The source describes it as for demonstration purposes rather than production use, runs on NVIDIA Lovelace GPUs under Linux, and is licensed under the NVIDIA Software and Model Evaluation License.

  9. Google · Innovation & AIOfficialAI score52

    Google says its language technology now covers over 300 languages with new speech, data, and on-device tools

    AIGoogle reports that its technologies and products now power everyday interactions in more than 300 languages used by over 7 billion people, about 86% of the global population. The post describes new speech models, including Gemini 3.5 Live Translate and Gemini 3.5 Transcribe, plus the TranslateGemma open translation models trained across 55 languages.

  10. Sebastian RaschkaXAI score28

    GPT-5.6 Astra and Qwen3.8 Max take different Paint approaches

    AIIn a Paint recreation test, GPT-5.6 Astra built the image from layered geometric shapes, while Qwen3.8 Max worked pixel by pixel. Qwen's output looks closer to the original, but Raschka argues this single example does not show either model generalizes better or has stronger computer-use or visual understanding, and it illustrates how benchmarks comparing only final results can be misleading.

    Video from @rasbt's post
  11. MiniMax Design (H3)OfficialAI score26

    MiniMax Design canvas runs Astra agent and Blender to produce full scene

    AIMiniMax Design lets a single brief drive a full production workflow on one canvas, with the Astra agent working in Blender through an official connector to build the scene and camera direction. The final video is generated with MiniMax H3 from the same canvas, with outputs syncing directly onto the canvas.

Sep 14

Sep 14Mon
  1. NVIDIA · new models on Hugging FaceOfficialAI score40

    NVIDIA releases FoundationStereo small stereo depth model on Hugging Face

    AINVIDIA Research released FoundationStereo-small, a zero-shot stereo depth model that takes an RGB stereo pair and outputs a disparity map, on Hugging Face. The model has about 6.3×10^7 parameters and ships as ONNX files at fixed 576x960 and 320x736 resolutions, with TensorRT and ONNX runtime support. It is licensed under the NVIDIA Open Model License and is ready for commercial use.

  2. Google · new models on Hugging FaceOfficialAI score62

    Google releases EmbeddingGemma 2, an open multimodal embedding model

    AIGoogle DeepMind released EmbeddingGemma 2, an open model under Apache 2.0 that maps text, images, video, and audio into one shared 768-dimensional vector space. The model has 740M total parameters and supports 8,192-token context, with Matryoshka truncation to 128d, 256d, and 512d. The source reports 14% better code-task performance than EmbeddingGemma 1 and says it is designed for consumer hardware such as phones and laptops.

    Why it matters: The release combines text, image, video, and audio retrieval in one 768-dimensional space at 740M parameters, a useful reference for on-device multimodal search design.

  3. Intern Large ModelsOfficialAI score25

    Intern-S2-397B gets Day-0 support in vLLM

    AIIntern-S2-397B, a model built for long-horizon scientific research, now has Day-0 support in vLLM. The model brings multimodal, reasoning, coding, and scientific agent capabilities, and vLLM has published a run recipe for it.

  4. Intern Large ModelsOfficialAI score62

    Intern-S2-397B: Shanghai AI Lab releases open multimodal model for scientific research

    AIIntern Large Models introduces Intern-S2-397B, a multimodal foundation model built for long-horizon scientific research and scientific agents. The post reports leading open-source results on IMO-Proof and AdvancedMathBench, and says the model reaches the level of Gemini 3.1 Pro on those tasks. It is now supported by vLLM and SGLang, with weights on Hugging Face and ModelScope and a chat demo available.

    Image from @intern_lm's post
  5. SenseTimeOfficialAI score22

    SenseTime Outlines Three AI Paradigm Shifts Toward Agentic Intelligence

    AIAt Guotai Junan Securities' 2026 Autumn Conference, SenseTime's Head of Capital Markets Philip Wong laid out three shifts reshaping AI: from single-modal to native multimodal, from token consumption to task delivery, and from single-point models to system-level full-stack capabilities. The post presents SenseTime's "One Model + One Token Factory + One Agent Harness" framework as built for these shifts.

    Image from @SenseTime_AI's post