MiniMax H3 is an omni-modal generative system capable of understanding and generating content across text, images, video, and audio. It generates video with native stereo audio at resolutions up to 2K and durations of up to 15 seconds. The system's pre-trained design enables broad multimodal context understanding and strong task generalization. Read original