MiniMax released H3 on July 31, 2026 — a single omni-modal model replacing what used to be separate Text2Video, Image2Video, Video Editing, and Video Reference pipelines. MiniMax takes text, images, video, and audio as combined input and generates video + native stereo audio, up to 2K, up to 15s, 24fps.
What is different?
Design Choice
Prior video pipelines were siloed: separate models per task (T2I, editing, subject/motion/style reference) and per modality (voice, SFX, music, video treated independently). H3 collapses these into one model trained end-to-end on all of it, with task relationships expressed in natural language rather than a fixed task menu.
Architecture
DiT (MiniMaxH3DiTModel) — flow-matching diffusion transformer, not autoregressive (no KV cache; full transformer re-evaluated per denoising step). It has ~33B Params (61.7GB size in bf16 precision).
Video and audio latents are flattened into one packed sequence and denoised jointly — cross-modal attention over that shared sequence is what keeps motion and sound in sync, not a bolted-on audio model. A 1344×768, 124-frame (~15s) generation runs ~91K target tokens.
Two VAEs, denoised together: Video VAE: 24 latent channels, tiled decode, 17-frame temporal clips. Audio VAE: 32 latent channels → 32kHz stereo.
Conditioning encoder: Qwen3-VL (62.1GB bf16), used purely as a feature extractor — H3 reads the unnormalized hidden state after decoder layer 50, LM head unused.
Contextual Omni Representation
The captioning problem shifts from “describe the target” to “describe the relationship between multiple context inputs and the target.” MiniMax’s pipeline distills ~100K tokens of raw multimodal source material into ~4K tokens of structured representation per training sample — dense enough to preserve cross-modal relationships (e.g., “use this video’s camera move, this image’s character, lip-synced to this audio”).
In-Context Regeneration (the 2K path)
No separate super-resolution module. The base model regenerates its own low-res output in-context, with access back to the original conditioning — so fine detail (text, texture) is recovered from source references rather than guessed at by an upscaler that’s already lost access to them.
Weights and licensing
Released on Hugging Face / ModelScope (MiniMaxAI) under the MiniMax H3 Community License. Not a permissive Apache/MIT license — includes deployment-region restrictions and terms extending to generated outputs, not just weights.
Self-hosting reality: full bf16 realistically needs multiple 80GB-class GPUs. Consumer-card deployment is possible via int8 + CPU offload, but expect ~75GB host RAM and a real throughput hit.
Training data — not disclosed
MiniMax has not published dataset composition, size, or source mix — standard for the video-model space, unlike LLM dataset-card norms. Disclosed: methodology only (the 100K→4K captioning pipeline, natural-data-only philosophy). A full technical report is reportedly forthcoming.
Where this fits for us
We’ve started running H3 through early evals at Fabeo — specifically the omni-reference (Ref2VA) path, since combined image+audio+video conditioning maps closely to how agencies actually brief a video (brand assets + voice reference + a motion reference, not just a text prompt). If it holds up on brand consistency and cost at batch scale, it’s a strong candidate for the pipeline alongside our existing model mix. More on this once we have real output data, not just spec sheets.
Part 2 covers competitive positioning vs. Veo 3.1, Sora 2, Kling 3.0, Seedance 2.0, and Wan.
References
As of early August 2026. Pending MiniMax’s full technical report.