Skip to contentSkip to stories

Updated

#Multimodal

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 23

Sep 23Wed
  1. Philipp SchmidAI score62

    Gemini 3.8 Flash TTS guide shows how to create and reuse your own voice

    AIGemini 3.8 Flash TTS and Flash-Lite TTS are now available in the Gemini API and AI Studio, with a new feature to replicate your own voice or create one from a sentence. The guide shows recording two clips, one of 15-20 seconds of natural speech and one reading a required consent sentence, then creating a reusable voice ID. It also explains that input text is now spoken word for word, so delivery belongs in speech_metadata.style and short sounds inline.

  2. AnthropicAI score62

    Claude finds a previously unknown enzyme system in bacteriophage DNA

    AIClaude has identified a previously unknown enzyme system in bacteriophage DNA, located beside a long array of repeating DNA that somewhat resembles CRISPR. Anthropic says its function is not yet understood, but only a handful of known systems share its features, all of which can cut, copy, and paste DNA. The source notes that programmable systems like CRISPR have been important to medicine, but more work is needed to learn what this system does and whether it can be used similarly.

  3. Google AI StudioAI score62

    Google releases Gemini 3.8 Flash TTS and Flash-Lite TTS text-to-speech models

    AIGoogle introduces Gemini 3.8 Flash TTS for creative voice design and Gemini 3.8 Flash-Lite TTS for high-volume, cost-efficient speech generation. Flash TTS supports voice creation from natural language prompts across more than 100 languages and dialects, and both models are rolling out today in the Gemini API and Google AI Studio, with enterprise access coming soon via Gemini Enterprise.

  4. Google DeepMindAI score60

    Google DeepMind launches Gemini 3.8 Flash TTS and Flash-Lite TTS models

    AIGoogle DeepMind introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, text-to-speech models offering custom voice design, line-by-line performance control, and multilingual support across more than 100 languages. Flash TTS is rolling out to developers in the Gemini API and Google AI Studio and to everyone in Gemini Notebook, while Flash-Lite TTS is available to developers and in Google Vids. Voice replication requires consent verification, and generated audio carries SynthID watermarking.

    Why it matters: The source details the voice design, performance direction, and consent safeguards, showing how the model covers creative and high-volume use cases with access across several Google products.

  5. ModelScopeAI score44

    NVIDIA releases Nemotron 3 Diarization for live speaker attribution

    AINVIDIA's Nemotron 3 Diarization is now available on ModelScope, labeling speakers and timestamps in streaming audio for up to eight speaker slots per conversation. The 99.2M-parameter model uses an end-to-end streaming architecture built on NVIDIA's Streaming Sortformer, running on Ampere, Hopper, and Blackwell GPUs via NeMo Speech C++. It is designed to pair with existing ASR systems such as Nemotron ASR, Parakeet, Canary, or Whisper to produce speaker-attributed transcripts.

  6. ModelScopeAI score40

    TeleOCR: 1.2B vision-language model parses documents, tops OmniDocBench v1.6

    AITeleOCR, a lightweight 1.2B vision-language model released under Apache 2.0, parses digital PDFs and warped phone photos without a separate dewarping model. It scores 96.87 overall on OmniDocBench v1.6, the highest among listed specialized VLMs, and ranks #1 in the ICDAR 2026 Sci-ImageMiner Challenge. It supports structured parsing of text, tables, formulas, layouts, and reading order, with synchronous or asynchronous vLLM inference.

  7. ModelScopeAI score62

    Xiaomi MiMo-V2.6 open-sourced as a multimodal agent model family under MIT License

    AIXiaomi has released MiMo-V2.6 as an open model family under the MIT License, designed for large-scale reinforcement learning. MiMo-V2.6-Pro scores 46 on the Artificial Analysis Intelligence Index, with 71.9 on DeepSWE v1.1, 89.9 on Terminal-Bench 2.1, and 82.0 on OSWorld-Verified. The 1.02T-parameter MoE activates 42B parameters and supports text, image, video, and audio input with a 1M-token context.

Sep 22

Sep 22Tue

Sep 21

Sep 21Mon
  1. Xiaomi MiMoAI score36

    Xiaomi MiMo-V2.6 unifies code, design, and tool use across creative outputs

    AIXiaomi's MiMo-V2.6 combines code, design, and tool use to build frontend interfaces, presentations, Figma-linked visual assets, and videos. The post says MiMo-V2.5-TTS supports narration in video production, and that the model can compose music, including an orchestral piece for around ten instruments that can be converted to MIDI. On Design Arena, the Pro version reportedly performs comparably to Claude Opus 5 and GPT-5.6 Sol.

  2. Xiaomi MiMoAI score44

    MiMo-V2.6 builds and interacts with 3D worlds from text, images, or video

    AIXiaomi's MiMo-V2.6 combines 3D spatial reasoning, multimodal perception, and computer use to turn text, images, or video into playable 3D worlds. The model coordinates agents to build scenes, write interaction logic, and refine results, and can create Blender objects for animation, 3D printing, and games. It also controls a Franka Panda arm in simulation via visual feedback and uses desktop tools to process data, inspecting results to adjust its next actions.

  3. Xiaomi MiMoAI score78

    Xiaomi releases open-weight MiMo-V2.6 Pro and Flash omnimodal models

    AIXiaomi MiMo has launched MiMo-V2.6 Pro and Flash, two omnimodal models with open model weights, a technical report, RL environments, and training code. The post says Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks and scores 46 on the Artificial Analysis Intelligence Index, the highest among open-source models. A benchmark table compares Pro and Flash with MiMo-V2.5 Pro and frontier models across code agent, general agent, cybersecurity, and visual agent tests.

    Why it matters: The source pairs open-weight release details with a benchmark table against Claude Opus 5 and GPT-5.6 Sol, letting readers compare Pro and Flash across agent tasks.

  4. Apple · new models on Hugging FaceAI score46

    Apple releases LensVLM-9B, a vision-language model for compressed text images

    AIApple has released LensVLM-9B on Hugging Face, a 9B-parameter Vision Language Model that scans compressed images of text and selectively expands relevant pages to their uncompressed form. The repository provides a demo script and supports compression settings of 5x, 10x, and 15x. Model files are under the Apple Machine Learning Research Model License, and the accompanying source code is distributed separately under the Apple Sample Code License.

  5. Xiaomi MiMo · new models on Hugging FaceAI score67

    Xiaomi releases MiMo-V2.6-Flash-RL, a 309B sparse MoE model with 1M context

    AIXiaomi released MiMo-V2.6-Flash-RL, an efficiency-balanced checkpoint in its MiMo-V2.6 series, on Hugging Face. The model is a sparse MoE with 309B total and 15B activated parameters, supports text, image, video, and audio input, and offers a 1M-token context. The technical report says it was trained with a single mixed reinforcement learning run across coding, agent, visual, and cybersecurity tasks.

    Why it matters: The report pairs its benchmark tables with the RL training method, which helps readers judge how the checkpoint's scores relate to its training approach.

  6. Xiaomi MiMo · new models on Hugging FaceAI score74

    Xiaomi MiMo-V2.6-Pro-RL released as 1.02T-parameter omnimodal model

    AIXiaomi MiMo released MiMo-V2.6-Pro-RL on Hugging Face, a sparse MoE model with 1.02T total and 42B activated parameters and a 1M-token context. The technical report says it accepts text, image, video, and audio, and was trained with a single mixed reinforcement learning run across coding, agent, visual, and cybersecurity tasks.

    Why it matters: The report pairs a 1.02T-parameter MoE model with an RL-based self-improvement method, useful for judging how reinforcement learning is scaled in frontier open models.

Sep 20

Sep 20Sun
  1. OpenBMBAI score44

    MiniCPM-o Booking Desk: open-source real-time voice appointment agent built on MiniCPM-o 4.5

    AIDeveloper @mrgoodmantweets built MiniCPM-o Booking Desk, an open-source appointment booking agent that uses MiniCPM-o 4.5 for real-time, full-duplex voice and audio-visual interaction. The agent listens, speaks, and reads live booking status from an operator screen, while deterministic state control keeps execution reliable. An appointment is only booked after user confirmation.

  2. ModelScopeAI score62

    Qwen-Image-2.1 unifies image generation and editing with native transparency

    AIAlibaba's ModelScope introduces Qwen-Image-2.1, a model that handles image generation and editing together, with native transparency and a compact 7B visual generation component. It adds KV cache reuse to speed up generation and editing while reducing memory use, especially with multiple reference images. The model can combine up to 10 reference images, make targeted local edits, and preserve portrait identity and product details.

  3. Qwen · new models on Hugging FaceAI score62

    Qwen releases Qwen-Image-2.1 prompt rewriter for image editing on Hugging Face

    AIQwen has open-sourced Qwen-Image-2.1, a unified text-to-image generation and image editing model with 7B visual generation parameters. The Hugging Face page for Qwen-Image-2.1-PE-I2I is a fine-tuned Qwen3.5-VL 9B prompt rewriter that turns vague editing instructions and input images into precise editing prompts, supporting up to 10 reference images.

    Why it matters: The model card documents usage with transformers and diffusers, letting readers see how the editing prompt rewriter connects to the generation pipeline.

Sep 19

Sep 19Sat
  1. StepFunAI score62

    StepFun Launches Step 5 Preview, a 600B MoE Model for Agentic Work

    AIStepFun has released Step 5 Preview, a flagship model for agentic work that it says delivers frontier-level performance in software engineering and professional knowledge work, with particular strength in finance. The model is a 600B total, 27B active mixture-of-experts design with a 1M context window and vision support. StepFun says it offers substantially lower task cost at comparable intelligence, and open weights are scheduled for October 15.

Sep 18

Sep 18Fri
  1. One Useful Thing (Ethan Mollick)AI score50

    Mollick says AI already does weeks of human work when guided, citing Zork and Eco library demos

    AIEthan Mollick says GPT-6 Astra and Fable 5.1 already enable transformative impact and can reliably handle weeks of human work when properly guided. He cites GPT-6 Astra turning the 1977 text adventure Zork into a 3D action-adventure game and Fable 5.1 reconstructing Umberto Eco's Milan library in 3D from videos, photos, and catalogues.

  2. Mustafa SuleymanAI score22

    Mustafa Suleyman claims best image generation quality-price performance

    AIMustafa Suleyman, owner of the source account associated with Microsoft and Copilot, says the post claims the best image generation quality-price performance in the world. The post itself gives no specific model, price, or benchmark figures. Background from Artificial Analysis says Muse Image, MAI-Image-2.6, and GPT Images 2.5 recently shifted text-to-image price and speed frontiers.

  3. Liquid AI · new models on Hugging FaceAI score55

    Liquid AI releases LFM2.5-VL-3B-DSpark drafter for faster vision-language decoding

    AILiquid AI released LFM2.5-VL-3B-DSpark, a speculative-decoding draft model for its LFM2.5-VL-3B vision-language model. The source reports decoding up to 2.66× faster on a single H100 with SGLang, up to 3.13× on Apple M5 Max with MLX-VLM, and up to 2.14× on Apple M3 Ultra with llama.cpp, with output unchanged under greedy decoding.

Sep 17

Sep 17Thu
  1. vLLM BlogAI score38

    vLLM Adds NVIDIA Hardware Video Decoding to Scale Multi-GPU Video Captioning

    AIvLLM now supports NVIDIA hardware video decoding through PyNvVideoCodec, moving video decoding off the CPU so multi-GPU video captioning can scale to 8 GPUs. In benchmarks on 8xH100 GPUs, GPU-based decoding provides more than double the throughput of the CPU-based decoder for Qwen/Qwen3-VL-8B-Instruct with 8 single-GPU vLLM replicas. The functionality is included in standard CUDA vLLM releases, and PyNvVideoCodec==2.0.4 is required for custom installations.

  2. SenseTimeAI score44

    SenseNova U1.5 open-sources 8B unified model for understanding and generation

    AISenseTime released its SenseNova U1.5 technical report, describing an open-source 8B native MoT unified model that connects understanding and generation through shared attention. The model reports 68.2% on VBVR-Pro-Bench, ahead of Nano-Banana-Pro (56.4%) and GPT-Image-2 (50.7%), and its full training recipes, including SFT, RL, and multi-expert on-policy distillation, are open-sourced.

  3. KrASIA · Big TechAI score50

    SenseTime's Lin Dahua Says Multimodal AI Breakthrough Could Come Within Two Years

    AISenseTime chief scientist Lin Dahua argues that native multimodal AI, which processes language, vision and other information in one shared model, is essential for AI to move beyond coding into industries and the physical world. SenseTime released the open-source SenseNova U1 in April and U1.5 Lite nearly four months later, and reported first-half 2026 revenue of RMB 2.91 billion, up 23.4% year-on-year. Lin's claim that a breakthrough could come within two years is the source's prediction, not a confirmed result.

Sep 16

Sep 16Wed
  1. Google for DevelopersAI score38

    Three companies use Gemini agentic video understanding to cut token costs

    AIMosaic, Ponder Studio, and Revyl used early access to Google's Gemini Flash models to test agentic video understanding on long footage. Mosaic reports a 97% cut in median token usage and nearly double the ability to handle complex edits, while Ponder Studio reports a 0.967 F1 score and about 72% lower token costs for B-roll selection. Revyl says the approach improved mobile UI bug-catching accuracy by 65%. The capability is available now for video uploads and YouTube videos via the Gemini API.

  2. Cat WuAI score60

    Claude merges Cowork and chat into one product with automatic routing

    AIAnthropic is merging Claude Cowork and chat into one Claude, and Claude Design is integrated so users can ask for slides, designs, or docs without switching apps. Claude decides from the prompt whether to give a quick answer or do deeper agentic work, and users can still stop, redirect, or adjust its effort. The change rolls out to Pro and Max over the next few weeks.

  3. BAAI · new models on Hugging FaceAI score34

    BAAI and Peking University release Brainμ-Spike spike camera image reconstruction model

    AIPeking University's Yu Zhaofei team and the Beijing Academy of Artificial Intelligence (BAAI) released Brainμ-Spike, a small convolutional network for spike camera image reconstruction that is paired with the Brainμ model. The package includes weights, inference scripts, and evaluation tools, but the base large model and LoRA weights are not yet released, so the full generation pipeline cannot run from this repository alone.

  4. inclusionAI (Ant Ling) · new models on Hugging FaceAI score55

    inclusionAI releases Realtime-Venus full-duplex audio-visual models on Hugging Face

    AIinclusionAI has published Realtime-Venus on Hugging Face with two 9B checkpoints: Realtime-Venus-Omni for audio-visual interaction and Realtime-Venus-Audio for audio-only conversation. Both are built on MiniCPM-o 4.5 with a Qwen3-8B backbone and support full-duplex dialogue, proactive responses, and training-free long-video memory. The asynchronous Realtime-Venus-Harness runtime is hosted in a separate GitHub repository.

Sep 15

Sep 15Tue
  1. Tencent · new models on Hugging FaceAI score37

    Tencent Releases WeVisDoc-2B and WeVisDoc-4B Document Parsing Models on Hugging Face

    AITencent's WeVisDoc-4B, fine-tuned from Qwen3-VL-4B-Instruct, scores 95.38 Overall on OmniDocBench v1.6 and 75.54 mean Overall across three PureDocBench tracks. The end-to-end parser converts page images into structured Markdown with LaTeX formulas and HTML tables, and the 2B variant is also available. The repository provides vLLM serving scripts with a 32768-token default context and a Python client for batch processing.

  2. Jazzyear · InsightsAI score67

    HiDream's vivago R1 agent targets five-minute AI video delivery

    AIHiDream.ai launched vivago R1, a content creation agent, globally, with a domestic version upgrade. The company says R1 can output five-minute high-quality videos through agent planning, with a claimed 85% usable-output rate and support for multi-round extensions. It also released HiDream-O1-Video-1.0, a native omni-modal video model supporting single shots of 5 to 20 seconds at 1080p.