#
Sol-H3- Spark Accelerating MiniMax-H3 768p Video Generation on a Single NVIDIA DGX Spark in 1 Minute
NVIDIA Research, Efficient AI Team & Singapore Lab.
768p resolution
5s duration
56s e2e latency
A two-stage pipeline specialized for a single NVIDIA DGX Spark generates a 384p draft and refines it to 768p.
Sol-H3 introduction 1344×768 / 24 FPS
Less than one minute #
56 seconds from a fresh prompt to a playable 768p video with T2VA or FL2VA. Ref2VA is slightly slower with more tokens.
Timing notes & breakdown
The two hot measurements reflect production serving: from prompt arrival, including fresh Qwen encoding, to a completed MP4 with audio.
Sol-H3 showcase #
Selected generations across T2VA, FL2VA and Ref2VA.
View exact prompt #
subject_definitions: <Subject 1> is the original grey felt raccoon in <Picture 1>, with the same wool face, dark eye patches, small paws and teal apron. <Picture 1> also defines the handmade cream-and-pistachio miniature pastry shop, counter and tart topped with one red strawberry. Preserve these character, material, prop and environment references in a new animated moment. detailed_description: Core concept: a tiny thief's innocent expression is betrayed by the evidence still in its paw. Create a five-second tactile stop-motion scene in a wide landscape composition. Show visible felt fibers, soft miniature lighting and deliberately stepped yet coherent puppet movement. Shot description: one continuous, locked medium-wide view containing the raccoon, tart and counter. From 0.0 to 1.5 seconds, <Subject 1> leans slightly toward the tart, glances sideways and slowly extends one paw toward the single strawberry. From 1.5 to 3.2 seconds, the paw gently closes around the berry and lifts it clear of the tart in one readable motion. Keep the strawberry's red body and green leaves visible, and leave the tart intact on the counter. From 3.2 to 5.0 seconds, a small offscreen wooden creak makes the raccoon straighten and look directly toward the camera with exaggerated innocence. It briefly holds that pose while one ear twitches and the apron settles; the stolen berry remains plainly visible in its raised paw. Maintain one raccoon and one strawberry throughout. No cuts, floating props, readable text, logos, subtitles, reference-image flashes or freeze frames. overall_soundscape: Tiny felt-and-fabric rustles, a faint countertop tap and one quiet wooden creak followed by soft shop room tone. No speech. non_diegetic_music: None.
View exact prompt #
subject_definitions: <Subject 1> is the Chinese mother in her early sixties in <Picture 1>, preserving her kind face, short salt-and-pepper hair, sage-green cardigan and ivory blouse. <Picture 1> also supplies the single small walnut tabletop radio with cream speaker fabric and two knobs, the oak table and warm quiet home. <Subject 2> is her adult daughter around thirty in <Picture 2>, preserving her face, straight dark shoulder-length hair, terracotta overshirt and off-white top. Compose these two women sitting together at that one table. Use the pictures as separate identity references, not as cutaway shots. detailed_description: Core concept: a repaired old radio creates a warm exchange between a mother and her adult daughter. A five-second photorealistic film, one continuous steady eye-level medium two-shot, both women's faces and the single radio visible, soft afternoon window light, clean composition and natural skin texture. In the first two seconds, the daughter looks at her mother with curiosity and clearly asks in standard Mandarin Chinese, “妈,修好了?” The mother looks up from the radio and answers warmly in standard Mandarin Chinese, “好了,听听看。” Her hand lightly rests by a knob; after finishing her reply, she turns it once with a small natural finger motion. They exchange a small smile. Maintain each speaker's identity and clothing, one radio on the table, no other people or distracting props. Each woman moves her lips only during her own line. Complete both short lines before the final moment, without overlap. No cuts, subtitles, text, logos, reference-picture flashes or freeze frames. overall_soundscape: Clear foreground Mandarin dialogue with two distinct female voices. Daughter: “妈,修好了?” Mother: “好了,听听看。” Speak only these Chinese words, with natural Mandarin pronunciation and no English, no narrator or extra chatter. Keep room ambience very quiet beneath speech. After the reply, a soft radio-knob click and a faint radio hiss; no broadcast speech or song. non_diegetic_music: None.
View exact prompt #
integrated_multimodal_description: [Shot 1] Photoreal wildlife documentary in crisp mountain daylight, a medium-wide view frames one majestic snow leopard with thick spotted fur moving across a jagged, snow-covered ridge beneath towering Himalayan peaks and a clear blue sky. During one continuous five-second shot, the camera pans slowly to follow its low stalking gait as each paw sinks slightly into fresh powder. The leopard s near the end, its long bushy tail twitching for balance while its green eyes scan the valley; frost-dusted whiskers and individual hairs remain sharply visible. Preserve the animal's anatomy, markings, scale, paw contact, and ridge geometry throughout, with no cuts, text, or logos. overall_soundscape: Soft mountain wind passes over the ridge while paws compress fresh snow with muted crunches. A faint tail brush and the leopard's quiet breathing are audible in the cold open air. non_diegetic_music: N/A
View exact prompt #
integrated_multimodal_description: [Shot 1] Cinematic live-action portrait at twilight, a young person wearing a delicate crown of wildflowers sits on a deep green mossy forest floor under cool blue moonlight. In one continuous five-second shot, the camera makes a slow small-amplitude arc around them, revealing soft petals, damp moss, and fireflies blinking warmly in the background. Tiny points of bioluminescent light play across their face as they slowly turn their head to watch one firefly settle on their shoulder, ending with the same quiet expression of wonder. Preserve the person's appearance, flower crown, seated pose, and forest layout; no cuts, captions, or logos. overall_soundscape: Quiet evening forest ambience continues beneath a light breeze through leaves and a soft rustle of clothing against moss. Sparse insects chirp in the distance as the person takes one calm breath. non_diegetic_music: N/A
View exact prompt #
integrated_multimodal_description: A naturalistic close medium shot of an elderly couple sitting together at a small seaside cafe table in warm morning sunlight. The woman has silver curls and a blue linen shirt; the man has a neat white beard and a cream cardigan. Two small ceramic coffee cups rest on the table. Over one continuous five-second take, she looks at him and asks in clear conversational English, "Same time tomorrow?" He turns toward her and replies in clear English, "Wouldn't miss it." They then share a brief unguarded laugh. The woman speaks first, the man speaks only after she finishes; both lines are fully audible and finish before the last second. Natural visible lip movement follows each speaker's own words, while the listener keeps their mouth relaxed. Their faces and clothing remain consistent. A sea breeze moves the loose edge of a striped awning, with a softly focused turquoise harbor behind them. Subtle handheld camera, lifelike skin texture, warm intimate documentary feeling, no captions or logos. overall_soundscape: Foreground intelligible English dialogue: the elderly woman warmly asks "Same time tomorrow?" and the elderly man affectionately replies "Wouldn't miss it." Natural female and male voices with distinct nonoverlapping turns, normal conversational volume, no whispers or extra words. Gentle shared laughter follows, with distant gulls, soft harbor waves and quiet cafe room tone well below the dialogue. non_diegetic_music: N/A
View exact prompt #
subject_definitions: <Subject 1> is the original young adult woman in <Picture 1>, preserving her short black bob, navy cape, cream dress and simple expressive facial design. <Picture 1> also defines the white folded-paper boat, pale stone cloud-city terrace, enormous soft clouds and peach-gold sky. Use these as character, costume, prop and environment references in original hand-painted 2D cel animation. detailed_description: Core concept: a small paper boat begins a journey into a vast welcoming sky. Create a five-second animated scene in a wide landscape composition, preserving crisp drawn silhouettes, lightly textured painted backgrounds and gentle cel shading. Shot description: one continuous, measured camera movement with no cuts. From 0.0 to 1.6 seconds, frame <Subject 1> at the terrace edge in a medium-wide view, holding the single white paper boat lightly before her. The aerial city's pale towers and bridges remain visible across the cloud sea. Her cape lifts gently in the breeze. From 1.6 to 3.0 seconds, she opens her fingers and lets the wind carry the boat smoothly away from her palms. The boat stays clearly folded paper, keeping its shape and a readable silhouette. From 3.0 to 5.0 seconds, the camera follows its modest outward drift, revealing more of the luminous cloud harbor while keeping the woman at the frame's edge. She watches it with a quiet hopeful smile. Maintain one boat and one woman, with continuous cape and cloud movement. No transformations, readable text, logos, subtitles, reference-image flashes or freeze frames. overall_soundscape: A gentle high-air breeze, a small paper flutter and a soft cape rustle. No dialogue. non_diegetic_music: None.
View exact prompt #
integrated_multimodal_description: [Shot 1] Photoreal macro product film in soft window daylight, an ornate gold nib on a handcrafted fountain pen moves across textured cream parchment while its polished mahogany barrel shows fine wood grain. In one continuous five-second shot, the camera tracks tightly with the nib as dark blue ink flows smoothly from the tip, follows a short curved stroke, and soaks into individual paper fibers. The metal flexes subtly under pressure and the fresh ink retains a wet sheen before beginning to dry. Keep the same pen, nib, ink line, and parchment surface coherent throughout, with no cuts, captions, or logos. overall_soundscape: The gold nib makes a delicate textured scratch across paper while the pen barrel shifts softly in the writer's grip. Quiet room tone and a faint rustle of parchment remain underneath. non_diegetic_music: N/A
View exact prompt #
integrated_multimodal_description: [Shot 1] Cinematic live-action inside a cramped spacecraft cockpit during atmospheric reentry, a close view holds on an astronaut's face behind a clear glass visor that reflects blinking red and green controls. In one continuous five-second shot, the camera shakes slightly with the vibrating cabin while beads of sweat travel down the astronaut's temple and their wide determined eyes scan the instrument panel. Fierce orange friction light flickers through a small porthole and moves dynamically across the detailed flight suit, helmet, and visor. Preserve the astronaut's appearance, helmet geometry, reflections, and cockpit layout; no cuts, captions, or logos. overall_soundscape: Deep mechanical vibration and rapid hull rattles fill the cockpit beneath intermittent electronic beeps. The astronaut's controlled breathing remains close inside the helmet as a low reentry roar rises outside. non_diegetic_music: N/A
View exact prompt #
subject_definitions: <Subject 1> is the adult Mediterranean woman mage in <Picture 1>, preserving her face, dark braided hair, deep indigo cloak and soft silver dress. The same picture defines the single grapefruit-sized transparent water sphere floating above her palm. Use <Picture 2> for the empty pale stone bridge, calm mountain lake and moonlit mountain setting. Place this one mage and her water sphere on that bridge. detailed_description: Core concept: a small tide of water rises gently at a mage's gesture. Create a five-second photorealistic fantasy scene with tangible water refraction, silver-blue light and generous open lake space. One continuous locked medium-wide shot keeps her face, open hand and sphere visible together. She gently raises her open palm, and the same sphere follows upward by only a few centimeters. A single moon reflection rolls over its refracting surface as she turns her eyes toward it with restrained wonder. Her cloak edge moves lightly in the breeze and small lake ripples catch the moonlight. Preserve the sphere's size and round shape, the natural hand anatomy and the quiet horizon. No transformation, cuts, text or logos. overall_soundscape: Gentle lake water, a soft night breeze and a delicate water tremble accompanying the sphere. No speech. non_diegetic_music: None.
View exact prompt #
subject_definitions: <Subject 1> is the elderly East Asian projectionist in <Picture 1>, with silver hair, round glasses and a burgundy cardigan. <Picture 1> also defines the vintage 35mm projector, close wooden booth and warm, dusty material palette. Use it as identity, wardrobe, prop and environment reference; compose a fresh moving scene from these elements. detailed_description: Core concept: a familiar machine returning to life brings a quiet private smile to its keeper. Create a five-second photorealistic film in a wide landscape composition, with believable aged skin, worn metal and soft tungsten illumination. Shot description: one continuous, restrained camera move. From 0.0 to 1.5 seconds, frame <Subject 1> and the projector together in a medium three-quarter view. He gently presses the projector's existing switch with one hand; his other hand rests naturally beside the machine. From 1.5 to 3.2 seconds, the visible reel begins turning and a narrow projection beam illuminates suspended dust across the booth. The camera edges closer along his visible side, retaining the projector in the composition. From 3.2 to 5.0 seconds, the warm reflected light reaches his face; he looks along the beam and forms a small, unhurried smile. Keep his facial features, glasses and cardigan consistent, with physically continuous hands and machinery. End on his living expression, with reel and dust still moving. No cuts, reference-image flashes, freeze frames, readable text, logos or subtitles. overall_soundscape: A soft switch click, the projector motor gathering into a gentle steady whirr, delicate film chatter and subdued booth room tone. No speech. non_diegetic_music: None.
View exact prompt #
subject_definitions: <Subject 1> is the elderly European woman in <Picture 1>, with the same silver bob and dark green knitwear. <Subject 2> is the younger adult Black man in <Picture 1>, preserving his face and tan jacket. The picture supplies their wooden night-train compartment, facing seats and the chessboard between them. Use it as identity, wardrobe, prop and spatial reference for a new moving shot. detailed_description: Core concept: one confident chess move earns a warm, wordless surrender between fellow travelers. Create a five-second photorealistic scene in a wide landscape composition. Use soft amber carriage light with cool streaks passing outside the window. Shot description: one continuous medium-wide shot across the small table, keeping both three-quarter faces and the board visible throughout. From 0.0 to 1.8 seconds, <Subject 1> calmly moves one existing chess piece to a nearby open square and releases it with a light tap. Her other hand stays resting beside her. From 1.8 to 3.3 seconds, she withdraws her hand and looks across at <Subject 2> with a slight knowing smile. He glances at the piece and then meets her eyes. From 3.3 to 5.0 seconds, he smiles broadly and raises both open hands a modest distance in playful surrender, while she remains pleasantly composed. Maintain their seated positions and the same two people, with subtle carriage sway and continuous passing window light. Keep all other chess pieces stationary. No cuts, dramatic gestures, readable text, logos, subtitles, reference-image flashes or freeze frames. overall_soundscape: Gentle wheel rhythm, soft carriage vibration, the single chess-piece tap and a small amused breath from the man. No dialogue. non_diegetic_music: None.
View exact prompt #
integrated_multimodal_description: A cinematic science-fiction shot inside a sunlit orbital greenhouse. An adult astronaut with short dark hair, wearing a practical navy flight suit without a helmet, stands beside neat rows of vivid green plants. In a continuous five-second shot, she gently brushes a broad leaf with two fingertips, then looks up toward the immense curved blue Earth visible through the observation window. Her face and clothing remain stable, her feet stay on the floor, and the plants remain rooted. The camera slowly moves sideways, revealing warm sunlight through the leaves, fine condensation on glass, and elegant white structural ribs. Grounded realistic materials and quiet wonder, no labels, captions or logos. overall_soundscape: Soft ventilation, a delicate leaf rustle, a quiet breath and a distant low machinery hum. non_diegetic_music: N/A
Coarse-to-fine pipeline #
Generate the scene on a compact H3 latent canvas, upscale in H3 space, then transfer to LTX for three-step refinement and full Conv VAE decoding.
MiniMax H3 · 384p
672 × 384 · 124 frames
NVFP4 Qwen · FP8 DiT · 4 steps LoRA
VAE Adapter
H3 space → LTX space
24s saved
VAE-free latent mapping
LTX · 768p
1344 × 768 · 121 frames
LTX-2.5 · Sol-Attn · 3 steps · no text encoder
Stage 1 supports compatible few-step LoRAs and 384p, 480p or 512p drafts, with matching latent upscaling to 768p.
The “black magic”: How Sol-H3 gets faster #
Two models need more than faster kernels: they need room to stay loaded. Our estimated naive two-stage configuration exceeds Spark’s shared memory budget.
Two changes make residency practical: (1) a learned VAE Adapter removes inter-stage video decode/re-encode; (2) a cached generic refinement prompt removes online Gemma and connector processing. Quantization then lets the retained components stay resident. The latent carries the scene condition; the generic prompt describes refinement quality, not scene content.
Technical details 5 optimizations
Sparsify on the fly
Query-dependent · one-pass · training-free
Select blocks while attention runs.
Each query block sets its own threshold over lightweight query–key scores. Selected blocks use exact attention; skipped blocks receive an approximate correction from pooled K/V summaries.
Routing, sparse attention and approximate correction share one online-softmax pass. Sol-Attn writes no separate full score map or routing-index tensor and needs no retraining.
Direct latent transfer
Direct H3 → LTX mapping VAE-free transfer
Skip the pixel-space handoff.
A learned VAE Adapter maps the upscaled H3 latent directly into LTX space. The pipeline never decodes the Stage 1 latent to frames and re-encodes it for Stage 2.
Removing the H3 VAE decoder and LTX VAE encoder eliminates two large resident components and both inter-stage forwards. Only the final LTX VAE decoder remains.
Cache the refiner prompt
Generic prompt · encoded once
Let the latent carry the scene.
The 384p draft latent already gives the refiner a strong content and motion prior. A content-agnostic quality prompt can be encoded once through Gemma and the connector.
The cached post-connector features remove runtime LTX text encoding and let Gemma stay out of the resident pipeline. Stage 1 still runs a fresh Qwen encode for every request.
Keep the runtime resident
Quantize · fit · stay warm
Fit both stages into one shared pool.
NVFP4 Qwen, an FP8 H3 DiT, the VAE Adapter, cached Stage 2 context and low-bit video components reduce the resident footprint while preserving the fixed two-stage recipe.
A resident pipeline avoids per-request model construction, checkpoint parsing and weight switching. Each request moves directly from prompt encoding through both video stages.
Fuse the execution path
BSA · layout · encode
Remove redundant work end to end.
Stage 1 VSA runs selected blocks through cuDNN BSA with fused routing and merge. The final Conv VAE emits its preferred NHWC layout directly, while chunked encode and audio mux avoid extra staging.
Eliminating dense fallbacks, standalone transposes and unnecessary host synchronization reduces memory traffic and keeps delivery work inside the measured request path.
Acknowledgements #
We thank the teams and open projects that helped turn Sol-H3 into one end-to-end Spark pipeline.
| Core contributors | Haopeng Li Junsong Chen Yitong Li Jincheng Yu Jingyu Xin Haocheng Xi Song Han Enze Xie |
|---|---|
| MiniMax | MiniMax-H3 base weights, used with FP8 quantization in Stage 1, and native video-and-audio generation. |
| Sol-Engine & Sol-Attn | The inference framework, SuperH3 workflow and sparse refinement attention. |
| Comfy-Org | Pre-quantized MiniMax-H3 FP8 checkpoints used in our earlier Spark experiments. |
| Lightricks | LTX-2.5 refinement and the official convolutional video VAE. |
| LBH-123-AI | The learned H3 latent upscaler used before the VAE adapter. |
| LightX2V | Open H3 acceleration work supporting the original Sol-Engine workflow. |
| Video DeltaNet · UC Berkeley | Complementary hybrid-attention work for MiniMax-H3. |
| Humanize +KDA | Operator and kernel optimization alongside Sol-Engine. |
| FastH3 · Hao AI Lab @ UCSD | The four-step FastH3 VSA model used for our 384p draft stage. |
Citations #
If this project helps your work, please cite the Spark release and the underlying Sol-Engine and Sol-Attn papers.
@misc{solh3spark2026,
title = {Sol-H3: Speed-of-Light MiniMax-H3 on Single NVIDIA DGX Spark Blackwell Superchip},
author = {{Sol-H3 Team}},
year = {2026},
howpublished = {\url{https://nvlabs.github.io/Sana/Sol-Engine/Sol-H3-Spark/}},
note = {Single-Spark project release}
}
@misc{li2026solvideoinferenceengine,
title = {Sol Video Inference Engine: Agent-Native Full-Stack Acceleration Framework for Efficient Video Generation},
author = {Yitong Li and Junsong Chen and Haopeng Li and Haozhe Liu and Jincheng Yu and Ligeng Zhu and Ping Luo and Song Han and Enze Xie},
year = {2026},
eprint = {2606.23743},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
doi = {10.48550/arXiv.2606.23743},
url = {https://arxiv.org/abs/2606.23743}
}
@misc{li2026solattn,
title = {Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification},
author = {Haopeng Li and Yitong Li and Junsong Chen and Tian Ye and Haozhe Liu and Jincheng Yu and Duomin Wang and Ruihua Zhang and Zeke Xie and Enze Xie and Song Han},
year = {2026},
eprint = {2607.24027},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
doi = {10.48550/arXiv.2607.24027},
url = {https://arxiv.org/abs/2607.24027}
}