MiniMax Open-Sources H3 Universal Video Generation Model
MiniMax has released H3, a universal multimodal model capable of generating videos at up to 2K resolution with native stereo audio and maximum duration of 15 seconds. The model processes multimodal context composed of text, images, videos, and audio, offering native support for 11 languages and flexibility across multiple aspect ratios and screen resolutions. The architecture features specialized components for multimodal context interpretation and high-resolution regeneration, enabling outputs with greater visual fidelity and precise details.