New video model:H3 Max.Describe the shot. Get it back with sound, in seconds MiniMax released H3 Max, a post-trained video model variant that generates 5–15 second clips with synchronized audio and directed camera motion, prioritizing prompt adherence and fast iteration at up to 768P resolution. The model, available on MiniMax's platform, charges 5 credits per second at 480P and 8 at 768P, with a 5-second 768P clip rendering in about 3 seconds. Standard H3 remains available for 2K and 4K output. Camera & Motion Direct every camera move Describe a tracking shot, pan, push-in, or fast follow. H3 Max keeps camera and subject movement aligned with the shot you wrote. MiniMax H3 Max, ready to run Turn a prompt or an image into a 5–15 second video with synchronized audio, directed camera motion, and optional first and last frames. Choose H3 Max for fast 480P or 768P iteration, or standard H3 for 2K and 4K output. Creating an account is free. Generation is charged per second of output, and a render that fails returns its credits. H3 Max model inference renders a 5-second 768P clip in about 3 seconds; queues and file processing may add time. The first frame of the video. The last frame. H3 generates the motion between the two. H3 can rewrite your prompt before rendering. Off keeps your wording exactly. 5 credits per second at 480P Your video will appear here. Turn a written prompt into a complete video, or use beginning and ending images to guide the shot's composition, motion, timing, and transition. H3 is MiniMax’s video model. H3 Max is a post-trained variant tuned for stronger prompt adherence, audiovisual quality and aesthetics. It creates synchronized audio with the picture and follows a written shot more closely than stock H3 at the same resolution, which is why it is the default here. H3 Max prioritizes prompt adherence, audiovisual quality, aesthetics and fast iteration. Standard H3 raises the output ceiling to 2K and 4K. The current credit rates below update automatically. | Model | H3 MaxDefault50% off | H3 | |---|---|---| | Best for | Prompt fidelity and rapid iteration | 2K and 4K delivery | | Maximum output | 768P | 4K | | Inputs | Text or image, with an optional end frame | Text or image, with an optional end frame | | Synchronized audio | Included | Included | | Current credits per second | 480P 5 · 768P 8 | 480P 10 · 768P 12 · 2K 26 · 4K 32 | By the H3 Max team Hover to preview; click to watch larger. Direct the shot, keep the look consistent, and bring motion and sound together in one generation. Camera & Motion Describe a tracking shot, pan, push-in, or fast follow. H3 Max keeps camera and subject movement aligned with the shot you wrote. First & Last Frame Start from an opening image and add an optional closing frame. H3 Max builds continuous motion between the two moments. Character Consistency Faces, clothing, proportions, and identity can stay coherent as the setting, light, and camera angle change. Native Audio Create ambience, effects, dialogue, and music in the same pass, timed to what happens on screen. Art Direction Carry a chosen palette, texture, linework, or cinematic treatment across every beat of the sequence. Prompt Adherence Describe the action step by step. H3 Max can follow the sequence and preserve requested visual details, including on-screen text when applicable. H3 Max turns ideas into motion in seconds. Model inference for a 5-second 768P video completes in about 3 seconds, with queue and file-processing time varying. Duration, resolution, aspect ratio, prompt expansion, seed and end frame — the same fields the model takes. See the exact cost before generating. Use a subscription or top up when you need more; failed renders return their credits. H3 reads both. The interface switches with one click and remembers the choice. Sign up, use your free credits to create a video, and upgrade only when you’re ready.