Skip to contentSkip to stories

Updated

#Multimodal

Oct 1

Oct 1Thu
  1. One Useful Thing (Ethan Mollick)AI score62

    Ethan Mollick Says Agent Coordination Is Easier Than Expected

    AIEthan Mollick says he was wrong to think coordinating AI agents would require careful human-designed management structures. He points to personal agents like dots and Muse, and to a swarm of thousands of OpenAI agents that solved a Navier-Stokes problem in 88 hours with thin coordination. He argues many management problems stem from human limits, which agents lack, so people should mainly guide direction while agents handle organizing.

  2. Manus BlogAI score45

    Manus 2.0 Adds Video Editor for Creating and Editing Publishable Videos

    AIManus 2.0 introduces Video Editor, which lets users refine videos Manus generates, including changes to music, captions, and cut timing, without regenerating the entire video. The article describes Manus creating explainers, launch films, and animations from a single prompt, drawing on web search, video models such as Seedance 2.5, and code for motion graphics.

Sep 30

Sep 30Wed
  1. Google DeepMindAI score88

    Google DeepMind releases Gemini 4 Argon to trusted cyber defenders first

    AIGoogle DeepMind announced Gemini 4 Argon, rolling out first to trusted cyber defenders through its Fairwind Program. Argon will launch at an introductory price of $2 per million input tokens and $10 per million output tokens, with output limits raised to 1M tokens. The post cites a 77.9% score on DeepSWE v1.1 and 91.7% on LVBench, and says broad availability will follow safeguard testing.

    Why it matters: The post pairs Argon's benchmark claims with the phased release, pricing, and safeguard details, helping readers weigh its frontier-level capabilities against its access limits.

  2. LlamaIndexAI score14

    LlamaIndex hosts document-processing events for AI agents in New York and San Francisco

    AILlamaIndex held Tuesday-night events in New York and San Francisco on document processing for AI agents, with the New York room filling a waitlist and San Francisco drawing almost 600 attendees. The talks focused on the problem that agents often receive document text without its layout, so they must guess which figures, such as a monthly rate versus a total on an invoice, mean what.

  3. IdeogramAI score38

    Ideogram 4.5 launches as a precise image edit model

    AIIdeogram released Ideogram 4.5, which it calls the most precise edit model, claiming it avoids the artifacts, pixel shifts, and color changes that leading models add with each edit. The company says this eliminates artifact buildup and makes multi-turn editing possible. It is live in Ideogram, via the API, and with launch partners, with open weights promised soon.

  4. SenseTimeAI score23

    SenseTime previews Dynamic Design, animating static images with SenseNova 6.8 Flash

    AISenseTime previewed Dynamic Design, powered by SenseNova 6.8 Flash, which turns static images into animated visuals. The system decides which elements stay static and which to animate, chooses HTML/CSS, SVG, transparent images, or video for each element, and choreographs text reveals and subject motion. SenseNova 6.8 Flash is coming soon, and SenseTime is offering a limited beta.

  5. Kling AIAI score35

    Kling 4.0 full-powered version showcased in a short film demo

    AIKling AI showcased a short film generated by the full-powered KLING 4.0, which the source describes as the full-powered version. The background post from @hq4ai says the film was made with all-round reference generation, runs a native 30 seconds in 21:9 cinematic format, and offers clearer visuals and sound with more precise lip-sync than Flash. The same post states Kling 4.0 will launch in October.

  6. ModelScopeAI score62

    InSpatio-World 1.5 turns images and videos into real-time explorable 4D worlds

    AIInSpatio-World 1.5 from InSpatio_AI turns a single image, four images, a panorama, or a video into a navigable scene with wide viewpoint changes. The 1.3B model scores 68.72 on WorldScore-Dynamic, ranking first among evaluated real-time and interactive methods, with speeds up to 24 FPS. The post says the code is released under Apache 2.0 and that dependencies keep their own licenses.

  7. Kling AI BlogAI score49

    Kling 4.0 Extends Native Video to 30 Seconds With Up to 10 Keyframes

    AIKling 4.0 extends native single-pass video generation from 15 to 30 seconds and adds Multiple Keyframes supporting up to 10 keyframe images, versus Start & End Frames in Kling 3.0. It also expands reference inputs to up to 15 combined assets, including up to 5 videos totaling 30 seconds, and adds 10-bit HDR at 1080p and 4K. The all-new Kling 4.0 will officially launch in October, and Kling 4.0 Flash became available to a limited group of early-access users on September 28.

Sep 29

Sep 29Tue
  1. Jerry LiuAI score22

    GPT-6.1 Sol Improves Table Parsing and Reading Order in OCR Benchmarks

    AIJerry Liu benchmarked gpt-6.1 sol on document OCR tasks and found a sizable increase in table parsing and reading order over gpt-6 sol from a week earlier. Its table parsing is similar to gpt-6 astra. He noted frontier models still cost roughly an order of magnitude more than cost-effective document parsing solutions, leaving room to improve the premium end above 1c per page.

  2. Google ResearchAI score35

    Google Research unveils Diffusion Controller for steering AI image generation

    AIGoogle Research introduced Diffusion Controller, a framework that treats image generation as a continuous control problem rather than separate inference-time guidance and fine-tuning fixes. Its lightweight add-on "steering damper" network keeps the base model frozen and works on black-box or gray-box models, and it outperformed the industry standard on human preference matching. In a Stable Diffusion v1.4 test, the fully unlocked version achieved a 90% win rate over the baseline.

  3. Microsoft ResearchAI score34

    Microsoft Research unveils Quine, an early multimodal world model of biology

    AIMicrosoft Research has introduced Quine, an early-stage research effort to build a multimodal world model of biology that connects insights across biological scales and modalities. The system is designed to help scientists computationally search a space far larger than intuition allows and prioritize hypotheses before lab testing. Experimental results are meant to feed back into the model and sharpen future research directions.

  4. ModelScopeAI score44

    Intern-Decision multimodal models scale structured decisions at 0.8B–4B

    AIShanghai AI Laboratory's Intern-Decision family of 0.8B, 2B, and 4B multimodal models averages 79.38, 84.68, and 90.02 across seven decision benchmarks. Intern-Decision-4B scores 88.74, surpassing Jev while achieving better probability calibration. Reported mean latency is 33.98, 33.28, and 44.16 ms, versus 109.70 ms for Jev in the same local HF setup.

  5. Luma AI NewsAI score22

    AI Photo Editing Prompt Formula Preserves Color, Light, and Skin in Campaign Edits

    AIThe article presents a four-part prompt structure (action verb, target element, desired result, protection instructions) for AI photo editing, saying it preserves approved work across platforms. It identifies three common failure causes: unmatched light direction, stacked edits in one prompt, and vague visual language. It states that simple skin retouching takes 2-3 minutes versus 15-30 minutes manually.

Sep 28

Sep 28Mon
  1. LlamaIndexAI score30

    LlamaIndex says frontier VLMs still struggle parsing tax and W-series forms

    AILlamaIndex argues that frontier vision-language models still fail on real forms such as W-2s, 1040s, W-9s, and scanned W-4s, because forms require detecting every field, preserving section hierarchy, linking values to their exact boxes, and reading handwriting and checkmarks. The company's blog post details these failure modes and presents a custom cookbook for LlamaParse as a cheaper way to handle such forms.

  2. Google · Gemini appAI score38

    See what 4 builders are making with Gemini 3.8 Flash

    AIGoogle says Gemini 3.8 Flash, its most intelligent workhorse model, improves on 3.7 Flash in software engineering, agentic tasks, and multistep reasoning by running extra reasoning steps and calling tools iteratively. The post highlights four community builds, including a model rocket simulation, an animated ink-painting effect, a 3D dinosaur skeleton, and an interactive automatic transmission simulation. Developers can try the model through Google Antigravity and Google AI Studio.

  3. Kling AIAI score42

    Kling 4.0 Flash launches now for Ultra Yearly subscribers; Kling 4.0 arrives October

    AIKling AI says its Kling 4.0 Flash is live now for Ultra Yearly subscribers, with the full Kling 4.0 coming this October. The update advertises up to 4K resolution, 10-bit HDR output, stereo audio, and native 30-second generation. It also adds Omni Reference supporting up to 15 multimodal references and multi-keyframe control with up to 10 keyframes.

  4. TechNode · AIAI score60

    Sanxingdui: Future Past, China's AI-produced theatrical film, releases October 23

    AIBona Film Group announced that Sanxingdui: Future Past, a 100-minute film using AI throughout production, will screen nationwide on October 23, 2026. The production team says AI handled tasks like image generation while over 100 professionals kept creative control, and it took two years and over 1.2 million source images to maintain consistency across the film.

  5. ModelScopeAI score43

    Jina-OCR-v1 parses full pages into Markdown at 2.57 pages per second

    AIJina-OCR-v1, a 3.4B-parameter MoE model that activates 570M parameters per token, converts entire document pages into structured Markdown at 2.57 pages per second. It scores 91.14 on OmniDocBench v1.6 and 83.4 on olmOCR-Bench, 7.4 points above DeepSeek-OCR on the latter, and delivers the highest throughput among 14 evaluated systems at concurrency 32. The model is released under CC BY-NC 4.0, so commercial use requires permission.

Sep 27

Sep 27Sun
  1. DeedyAI score34

    Deedy urges explainer videos for every open source repo, citing SQLite example

    AIDeedy argues every open source repository should have a roughly seven-minute explainer video like the one made for SQLite, covering its purpose, a high-level code map, a query's path through the codebase, core abstractions, and a real execution trace including join-order query planning. He says the video was generated with Opus 5.5 and Gemini 3.8 TTS, and he expresses amazement at how coherent and capable the model is.

  2. Exponential ViewAI score44

    DeepMind Essay Argues AGI Will Emerge Through Collective Cooperation Among AI Agents

    AIDeepMind has published an essay arguing that AGI will emerge through "cooperative interactions among models, tools, institutions, and human participants" rather than from a single winning AI. The commentary supports the collective framing but rejects treating AI agents as having their own theory of mind, arguing that creating new moral subjects should remain humanity's remit.

Sep 26

Sep 26Sat
  1. InternLM (Shanghai AI Lab) · new models on Hugging FaceAI score44

    Intern-Decision-2B: Structured Multi-Question Decision Model Fine-Tuned from Qwen3.5-2B

    AIShanghai AI Lab's InternLM released Intern-Decision-2B, a multimodal structured decision model fine-tuned from Qwen3.5-2B that returns calibrated answer distributions for multiple questions in one forward pass. It averages 84.68 across listed benchmarks with a 0.437 Brier score and 33.28 ms mean latency on a single RTX 4090. Model weights, a Python DecisionEngine API, and GitHub code are available, with support for up to 16 questions and eight images.

Sep 25

Sep 25Fri
  1. Google AIAI score57

    Google AI lists weekly releases including Gemini 3.8 TTS, Live Avatar, and Project Suncatcher

    AIGoogle AI's weekly roundup lists Gemini 3.8 Flash TTS and Flash-Lite TTS as expressive audio generation models. It also announces Gemini 3.8 Live with Live Avatar for near real-time visual conversation and a Live Chat voice feature on the Gemini Notebook mobile app across about 100 languages. Project Suncatcher will launch a prototype satellite to test Google TPUs in orbit and explore solar-powered AI compute in space.

  2. Meituan LongCatAI score62

    Meituan LongCat-2.5-Preview Launches with 1.6T Parameters and 1M-Token Context

    AIMeituan's LongCat team has released LongCat-2.5-Preview, a natively multimodal model with 1.6T total parameters, about 48B active, and a 1M-token context window. The model is built for long-horizon tasks spanning terminals, browsers, GUIs, spreadsheets, and design tools. It is available now through an API on the LongCat platform and a chat interface.