The technology behind MiniMax H3 Max
How the model works, based on public documentation and sourced evidence.
MiniMax H3 Max is fal's post-trained variant of MiniMax H3, co-optimized with fal's own inference stack and served only through fal. It generates video from a text prompt or animates a still image, and it produces a synchronized audio track in the same pass: room tone, foley and ambience arrive cut to what is on screen. Output runs 5 to 15 seconds at 480P or 768P, and 768P resolves to 1344x768 at 24 fps.
The headline is speed. A 5 second 768P clip comes back in 2.87 seconds of inference for text-to-video and 1.87 seconds for image-to-video, with the full round trip including queue and download landing near 9 seconds. That is faster than the clip itself plays back, and it is the reason to reach for this model over the base H3 that also sits on Vidney.
The trade is resolution. Base MiniMax H3 offers a 2K tier; H3 Max stops at 768P. Everything above 768P still belongs to the base model. In exchange H3 Max leads the public quality boards on overall quality, prompt understanding and aesthetics, and it holds art direction across shots rather than drifting.
The audio track is not optional and there is no mute switch, because sound is generated inside the same pass as the picture. Steer it from the prompt instead by naming what should be heard. Image-to-video also accepts a final keyframe, so you can pin both ends of the shot and let the model fill the motion between them.