# New video model:H3 Max.Describe the shot. Get it back with sound, in seconds

> Source: <https://h3max.pro/>
> Published: 2026-09-03 05:30:15+00:00

Camera & Motion

### Direct every camera move

Describe a tracking shot, pan, push-in, or fast follow. H3 Max keeps camera and subject movement aligned with the shot you wrote.

MiniMax H3 Max, ready to run

Turn a prompt or an image into a 5–15 second video with synchronized audio, directed camera motion, and optional first and last frames. Choose H3 Max for fast 480P or 768P iteration, or standard H3 for 2K and 4K output.

Creating an account is free. Generation is charged per second of output, and a render that fails returns its credits.

H3 Max model inference renders a 5-second 768P clip in about 3 seconds; queues and file processing may add time.

The first frame of the video.

The last frame. H3 generates the motion between the two.

H3 can rewrite your prompt before rendering. Off keeps your wording exactly.

5 credits per second at 480P

Your video will appear here.

Turn a written prompt into a complete video, or use beginning and ending images to guide the shot's composition, motion, timing, and transition.

H3 is MiniMax’s video model. H3 Max is a post-trained variant tuned for stronger prompt adherence, audiovisual quality and aesthetics. It creates synchronized audio with the picture and follows a written shot more closely than stock H3 at the same resolution, which is why it is the default here.

H3 Max prioritizes prompt adherence, audiovisual quality, aesthetics and fast iteration. Standard H3 raises the output ceiling to 2K and 4K. The current credit rates below update automatically.

| Model | H3 MaxDefault50% off | H3 |
|---|---|---|
| Best for | Prompt fidelity and rapid iteration | 2K and 4K delivery |
| Maximum output | 768P | 4K |
| Inputs | Text or image, with an optional end frame | Text or image, with an optional end frame |
| Synchronized audio | Included | Included |
| Current credits per second | 480P 5 · 768P 8 | 480P 10 · 768P 12 · 2K 26 · 4K 32 |

By the H3 Max team

Hover to preview; click to watch larger.

Direct the shot, keep the look consistent, and bring motion and sound together in one generation.

Camera & Motion

Describe a tracking shot, pan, push-in, or fast follow. H3 Max keeps camera and subject movement aligned with the shot you wrote.

First & Last Frame

Start from an opening image and add an optional closing frame. H3 Max builds continuous motion between the two moments.

Character Consistency

Faces, clothing, proportions, and identity can stay coherent as the setting, light, and camera angle change.

Native Audio

Create ambience, effects, dialogue, and music in the same pass, timed to what happens on screen.

Art Direction

Carry a chosen palette, texture, linework, or cinematic treatment across every beat of the sequence.

Prompt Adherence

Describe the action step by step. H3 Max can follow the sequence and preserve requested visual details, including on-screen text when applicable.

H3 Max turns ideas into motion in seconds. Model inference for a 5-second 768P video completes in about 3 seconds, with queue and file-processing time varying.

Duration, resolution, aspect ratio, prompt expansion, seed and end frame — the same fields the model takes.

See the exact cost before generating. Use a subscription or top up when you need more; failed renders return their credits.

H3 reads both. The interface switches with one click and remembers the choice.

Sign up, use your free credits to create a video, and upgrade only when you’re ready.
