MiniMax H3: An Open Model Breaking the Boundaries Between Tasks and Modalities

AIH3MultimodalVideo Generation

Today, we're officially launching MiniMax H3, a general-purpose omni-modal generation model. H3 can jointly understand multimodal contexts spanning text, images, video, and audio. It generates video with native stereo audio at up to 2K resolution and 15 seconds in length.

Early testing shows that MiniMax H3 delivers production-ready content generation across a wide range of use cases. It excels at instruction following, accurate text and brand presentation, and V2V Motion Transfer. With precise, controllable multimodal generation and editing, H3 is built for advertising, branding, e-commerce, product design, UI/UX, gaming, and more.

Powered by technologies including Contextual Omni Representation, H3-VAE, H3-Omni Transformer, and In-Context Regeneration, H3 delivers industry-leading price-performance. We offer 2K resolution by default. At 2K, H3's per-second price is less than a third of mainstream models, and at 768p, it's less than half the price of mainstream models' 720p.

Closed-source models have long dominated video generation, with slower iteration and a less open ecosystem than fields like large language models. To support the open-source community, accelerate compatibility with a broader range of AI hardware, and make it easier for users to build their own customized versions, we plan to open up the model weights in the coming days, subject to applicable laws and regulations. Hardware compatibility has been a key consideration since the earliest stages of H3's design.

Multimodal context understanding

Real creative work requires blending complex information across modalities, pulling in images, audio, video, and more as input sources. For example, the prompt for the shot below is: "Reference the Hitchcock camera movement from Video 1, have the character in Image 2 sing, with the vocals matching Audio 3." Just describe the relationship between the context and the target video in words. H3 handles the complex, full-modality understanding on its own.

Video 1

H3 en-v2 image 1

Image 2

Audio 3

H3-Generated Video:

H3 Use Cases

H3's Design Philosophy

Breaking the Boundaries Between Tasks: From Specialized to General-Purpose

We previously developed two generations of models: Hailuo 01 built the system from the ground up, and Hailuo 02 focused on improving core components like architectural efficiency, data quality, and scale.

In designing H3, we recognized the limitations of prior generative models around task boundaries: image generation was typically split into separate expert models for T2I, editing, subject reference, motion reference, and style reference; voice, sound effects, and music in audio generation were largely studied as separate domains; and video generation was further fragmented into text-to-video, image-to-video, first-and-last-frame, subject reference, motion reference, voice reference, video editing, and more, with clear boundaries between image, video, and audio as well.

These siloed tasks, capabilities, and modalities constrained creative freedom in practice. And, on the training side, capped the model's ability to generalize. Both point to major room for change, in application paradigm and training paradigm alike. So the first principle guiding H3's development was unifying and generalizing across tasks.

Based on this, here's a brief overview of H3's pretraining paradigm:

These choices all point to the same goal: giving H3 broad multimodal context understanding and generation capabilities from the pretraining stage onward.

Earlier we talked about the shift in training paradigm, so how is the application landscape for video models changing as well?

We're seeing creators describe their intent directly in natural language, rather than just entering a simple visual prompt. As multimodal understanding, instruction following, and complex task execution keep improving, video models will be able to grasp fuller creative intent, handle more complex content needs, and gradually move from "generating a clip" to genuinely participating in the entire content production process.

H3's Technical Choices

H3 is built on a simple design philosophy, but bringing it to life was extraordinarily complex. From Hailuo 01 and Hailuo 02 to H3, each generation has been an order of magnitude more complex to build than the last.

A brief look at some of H3's core technologies:

One of the most important things we did in engineering H3 was strengthening its captioning capability.

H3-Omni Transformer: An Architecture Built for Task Generalization

We'll be sharing the full H3 Technical Report soon. We look forward to hearing your feedback.

Our Vision & What's Next

Language, images, video, and audio are fundamental modalities of human experience. They are deeply interconnected and serve as our primary interfaces with the world. Together, they form multimodal contexts that can communicate a vast range of information efficiently, and the ability to express and share information is itself a form of productivity.

For multimodal understanding and generation, we view language as a generalizable, scalable computational system. That is why we believe multimodal intelligence should be deeply grounded in language.

H3 still has room to grow, and here are our priorities for future versions:

Intelligence with Everyone.