
Minimax H3 released with advanced capabilities
MiniMax has released H3, a general-purpose multimodal generation model that unifies text, images, video, and audio into a single context. The system produces clips up to 15 seconds long at 2K resolution with native stereo audio and supports advanced features like video-to-video motion transfer.
Published by Jin · 2 min read · 4 AUG 2026
MiniMax Blog product demonstration
MiniMax Blog product demonstration
MiniMax Blog product demonstration
- Up to 2K
- 4 to 15 seconds
- Native 32 kHz stereo
- 33B parameters
MiniMax has officially introduced H3, a next-generation multimodal generation model designed to process unified contexts across text, images, video, and audio. Succeeding the previous Hailuo generations, H3 moves beyond siloed specialized tasks by integrating diverse data types and modalities directly from the pretraining stage.
Architecture and Technical Design
The H3 architecture relies on several core components designed to enhance both capability and efficiency. The system incorporates the H3-Encoder, utilizing pretrained Qwen3-VL-32B hidden states, alongside dedicated H3-VisualVAE and H3-AudioVAE components to manage visual and stereo audio latents. The central processing unit, the H3-Omni-Transformer, is a 33B-parameter dense single-stream transformer that jointly predicts video and audio latents while employing three-dimensional Multimodal Rotary Position Embeddings.
To address computational efficiency, the H3-VAE completely overhauls prior tokenizers, achieving a high compression ratio that delivers a fourfold gain in effective sequence length. Furthermore, H3-Regenerate-2K avoids conventional bolt-on super-resolution modules by having the base model regenerate its low-resolution output in-context, preserving fine details like small text and brand marks.
Input Specifications and Performance
H3 supports video generation between 4 and 15 seconds at a fixed 24 frames per second across various aspect ratios, including 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16. The platform supports multiple input configurations, including text-to-video, first-and-last-frame conditioning, and an omni-reference mode accepting up to 9 images, 3 video clips, and 3 audio files.

Source — MiniMax Blog ↗
Worth a read?