MiniMax H3 unifies text, image, video, and audio generation in one model
Original titleMiniMax H3: An Open Model Breaking the Boundaries Between Tasks and Modalities
AISummary
MiniMax launches H3, a general-purpose multimodal generation model that understands text, images, video, and audio as unified context.
It generates video up to 15 seconds at 2K resolution with native stereo sound, and the company says model weights will be opened in the coming days, subject to applicable laws and regulations.
MiniMax also says H3 is priced below mainstream models at 2K and 768p.
AIWhy it matters
The post explains how a unified multimodal design and training choices enable 2K video with native stereo sound, useful for comparing against closed video generators.
Source: MiniMax Blog · minimax.ioPublished · added here