Wan 3.0 is half off on Replicate this weekend only
AIReplicate is offering Wan 3.0 at half price this weekend only. The post links to the Wan 3 model page on Replicate under Alibaba.

Updated
Updated
Showing low-relevance items too. Hide low-relevance items
AIReplicate is offering Wan 3.0 at half price this weekend only. The post links to the Wan 3 model page on Replicate under Alibaba.

AIRadixArk introduced LoRA SFT in Miles-diffusion for fast, targeted post-training of diffusion models. The company trained a rank-64 LoRA adapter for MiniMax H3 to improve physical realism, using 254 curated training windows and under 3 hours on 8 GPUs. The adapter can be exported to safetensors and served directly with SGLang without retraining the full model.

AIReplicate announced that Gemini Omni 1.1 Flash from Google DeepMind is now live on its platform. The update adds scene extension, control of a shot's starting and ending frames, video input references, upscaling to 4K, and 360p fast prototyping.
AIMiniMax-H3 on 8×H200 GPUs reaches 1.85–1.95x lossless speedup over Diffusers without approximation, with fixed prompts, seeds, resolution, FPS, and 50 denoising steps. Adding step reuse and sparse attention raises speedup to as much as 6.24x, but quality varies by workload, with SSIM from 0.76 to 0.91. Two presets trade off the two: a conservative Cache-DiT setting gives 2.99x at 0.90–0.98 SSIM, while a faster SubBlock 0.75 plus Cache-DiT stride gives 4.90–5.93x at 0.77–0.92.

AIGoogle is rolling out Gemini Omni 1.1 Flash updates in Google Flow, adding start and end frame controls to keep characters and narrative consistent between transitions. Creators can draft at 360p on lower credit cost, then export final clips in 720p, 1080p, or 4K for digital, social, or broadcast workflows.
AIReplicate has made Wan 3.0 Prime, the accelerated variant of Alibaba's Wan 3.0, available for text-to-video, image-to-video, and reference-driven workflows. The model generates clips up to 30 seconds long in a single shot with integrated audio-visual generation.
AIReplicate is offering a 30% discount on Alibaba's Wan 3.0 video model, available at The background post says Wan 3.0 generates native single-take videos up to 30 seconds long with synchronized audio.
AIReplicate has made Alibaba's Wan 3 available to try through a hosted demo page. The post links to the model page at and provides no further details on capabilities, specifications, or pricing.
AIGoogle shows Gemini 3.7 Flash, paired with Nano Banana and Omni, generating a fully animated parallax landing page from a single prompt. The model calls tools to write the copy, create image references, and render parallax-ready Omni videos.
AIByteDance has released Bernini-Diffusers-v2 on Hugging Face, a video generation and editing pipeline combining a Qwen2.5-VL planner with Wan2.2 diffusion components. The model card recommends it over Bernini-R for complex requests needing stronger instruction following and multi-step semantic planning. Code and weights are available under Apache License 2.0.
AIMeituan LongCat showed a demo in which a single prompt produced an immersive web landing page from a cinematic concept. The post is a short video teaser and gives no model name, version, or technical details.
AIMeituan LongCat turns one natural-language prompt into a full CG animation sequence, covering code, motion, and rendering. The post is a short demonstration video with no further technical details on the model or its performance.
AIMiniMax launches H3, a general-purpose multimodal generation model that understands text, images, video, and audio as unified context. It generates video up to 15 seconds at 2K resolution with native stereo sound, and the company says model weights will be opened in the coming days, subject to applicable laws and regulations. MiniMax also says H3 is priced below mainstream models at 2K and 768p.
Why it matters: The post explains how a unified multimodal design and training choices enable 2K video with native stereo sound, useful for comparing against closed video generators.
AIGoogle has released short films created by its fourth Flow Sessions cohort, a six-week creative partner program for artists. The program uses Google Flow, Google's creative studio featuring its advanced generative models, to support co-creation and passion projects.
AIMiniMax released H3, an open-weights omni-modal model that generates video with native stereo audio up to 2K and 15 seconds. The system combines H3-Context-IR preprocessing, the H3-Base generator at 768p, and H3-Regenerate-2K for 2K output, with the Context-IR and 2K modules available only through API.
Why it matters: The source details a three-module pipeline and open weights with deployment paths, showing how a video model is served and reproduced locally.
AISkywork Video is presented as an AI studio that covers the full workflow from idea and script through storyboard, video clips, editing, and final export. The post shares an advertisement video created with Skywork Video and includes a coupon code, "Bishal."
AIBlack Forest Labs released FLUX 3, a multimodal model trained on images, video, and audio, and mimic built FLUX-mimic on its video backbone to control robots. In a soft-body kitting task, mimic reports a 95% success rate without single-task fine-tuning, compared with 55% for an adapted π0.5 model. FLUX 3 Video is in early access, with action prediction offered to selected partners and an open-weight backbone planned.
AISkywork has released a video creation feature that produces videos in different styles. The post gives no further details on supported styles, models, or availability.
AIEthan Mollick's guide says ChatGPT and Claude are the main choices for real work, since their agent modes can act on a computer. He separates agent modes that run on the company's computers from those that access the user's own computer. He recommends keeping approval settings on for sending, spending, or deleting, because of prompt injection risk. He also notes that Gemini currently lags for agentic work, though its Notebook and video tools are useful.
AISkywork's X post directs readers to a video creation page at with no further product details given. The post contains only a link and a call to create content, so no features, models, or specifications can be confirmed.
AISkywork Video generated a continuous single-take video with no cuts, according to the post. The post gives no further details on duration, resolution, or access.
AISkywork announced its Skywork Video tool as the way to start your next video, linking to its video page. The post gives no details on features, models, pricing, or capabilities.
AISkywork announced an upgrade to Skywork Video that adds an infinite canvas for mapping ideas, a storyboard for organizing scenes, and a video editor for refining the final video. The company says these tools keep planning, sequencing, and editing in one workflow.
AIMeta Superintelligence Labs has released Muse Image, which can invoke search and coding tools and self-refine its generations before output. It is available today in the Meta AI app, meta.ai, Instagram Stories in the US, and WhatsApp in limited countries, with Facebook coming soon. Meta also previewed Muse Video, which is coming soon to creators and Meta AI and is reported as ranking No. 3 on Arena for text-to-video at the time of writing.
Why it matters: The source describes how search, code execution, and self-refinement change image generation, which matters to anyone comparing agentic media models with plain prompt-to-image systems.
AIGoogle Labs has expanded Project Genie access to Google AI Ultra 5X subscribers globally, its latest subscription tier. The earlier background post says Project Genie is now fully available to all Google AI Ultra subscribers aged 18 and older worldwide.
AIByteDance open-sourced inference code and model weights for Bernini-Diffusers, a full video generation and editing pipeline with an MLLM-based semantic planner and a DiT-based renderer. The release bundles a Qwen2.5-VL planner and Wan2.2 diffusion components in one self-contained directory, and the source recommends it over the renderer-only Bernini-R for complex instruction following.
AIByteDance's Sa2VA-Qwen3-VL-4B-SAM3 is built on Qwen3-VL-4B-Instruct with a SAM3 grounding encoder and produces dense image and video referring segmentation alongside chat. It reports 83.7 cIoU on RefCOCO val, 65.3 J&F on MeViS (val_u), and 77.1 on Ref-DAVIS17. The checkpoint is self-contained and loads on Hugging Face with trust_remote_code=True, with no extra packages required.
AIZ.ai released SCAIL-2, an open-source model that animates a reference character from a driving video without skeleton maps or inpainting masks. It also supports character replacement, multi-character scenes, and animal-driving, with 512p and 704p resolutions and inputs whose height and width are both divisible by 32.
AIByteDance has open-sourced the 1.3B-parameter weights of its Bernini Renderer (Bernini-R), available on Hugging Face as ByteDance/Bernini-R-1.3B-Diffusers. Fine-tuned from Wan2.1-1.3B, the model performs close to the 14B variant on simple tasks such as style transfer, subtitle or watermark removal, and local editing, but lags on complex tasks such as human generation. The release requires a CUDA GPU, with an H100 recommended for FlashAttention-3.
AIAt Google I/O, creators behind Flow, Project Genie, and Google Flow Music said human imagination, not the technology itself, shapes new storytelling. Designers Khyati Trehan and Kaloyan blend traditional design knowledge with vibe-coding to build Google Flow Tools, with Trehan saying that if the right tool doesn't exist, she can make it. Google Labs points users to Google Flow, Project Genie, and Google Flow Music at labs.google.

AISoumith Chintala, a Thinky-linked voice, said the company is at step one of a plan to increase human-AI bandwidth and raise the ceiling of joint intelligence. He shared a preview of interaction models, described as real-time collaborative tools that talk, listen, watch, and think alongside people. A linked Thinking Machines post describes the approach and early results.
AINVIDIA highlighted the developers who won the Cosmos Cookoff and how they used Cosmos Reason 2 to build projects spanning disaster-response drones, explainable visual AI, and intelligent security systems. The post links to a YouTube showcase and a LinkedIn recap of the event.

AIPrismAudio is a framework that integrates reinforcement learning into video-to-audio generation, using a Chain-of-Thought planning mechanism. It builds on ThinkSound by splitting single-step reasoning into four CoT modules for semantic, temporal, aesthetic, and spatial dimensions, each with targeted reward functions. Code, model weights, and datasets are released for research and educational use under the MIT License, and commercial use requires explicit author authorization.
AIXiaomi introduces MiMo-V2-Omni, a single model that fuses image, video, and audio encoders into a shared backbone with native tool calling and UI grounding. The company reports benchmark results against Gemini 3 Pro, Claude Opus 4.6, and GPT 5.2, and demonstrates browser-based shopping and video-publishing workflows run through the OpenClaw agent scaffold. It also states the model supports over 10 hours of continuous audio understanding.
Why it matters: The page gives benchmark comparisons, a driving-risk demo, and browser-task walkthroughs, letting readers check how far the omni-modal claims extend into agent use.
AIChip Huyen praised projects at last weekend's Agentic Hackathon, which hosted by MongoDB and Cerebral Valley, where she served as a judge. Teams tackled long-running tasks such as memory management, recovery from mid-task failures, and consistency across steps and sub-agents, along with adaptive retrieval across databases, search indices, and websites. Finalist demos are scheduled in San Francisco tomorrow, with talks by Douglas Eck.

AIRunway announced GWM-1, its first general world model family, built on Gen-4.5 and generating frames autoregressively in real time under interactive control. It comes in three variants: GWM Worlds for explorable environments, GWM Avatars for conversational characters, and GWM Robotics for robotic manipulation. Runway also says it is working toward unifying these domains under a single base world model, and GWM Robotics includes a Python SDK.
Why it matters: The post separates three GWM-1 variants and ties each to a concrete use, which clarifies where a general world model would fit compared with a single model.
AIRunway announced Gen-4.5, a video generation model that it says holds the top position on the Artificial Analysis Text-to-Video benchmark with 1,247 Elo points. The model is available across all paid Runway plans at comparable pricing, and the post lists limitations including causal reasoning errors, object permanence failures, and success bias.
Why it matters: The post separates Runway's own ranking claim from the listed limitations, such as causal reasoning and object permanence errors, which helps judge where the model is reliable.