{"slug": "fasth3-on-consumer-hardware-rtx-gpus-dgx-spark-and-macs-plus-fasth3-trim", "title": "FastH3 on Consumer Hardware: RTX GPUs, DGX Spark and Macs, Plus FastH3 Trim", "summary": "FastH3 V2, an eight-step video-and-audio generation model whose weights previously required 138 GiB and data-center GPUs, now runs on a single consumer machine — an RTX 5090, RTX 4090, RTX PRO 6000, DGX Spark, or Apple Silicon Mac — according to its developer. The release ships in NVFP4 for Blackwell GPUs and DGX Spark, FP8 for the RTX 4090 and lower-memory GPUs, and INT6 for Apple Silicon, with an RTX PRO 6000 rendering a 5-second 832×480 clip in 13.5 seconds and an RTX 5090 in 14.8 seconds. A companion experimental model, FastH3 Trim, removes 8 of the 50 transformer blocks, is 4.2× smaller than base H3, and runs in as little as 8 GB of GPU memory.", "body_md": "FastH3 V2 generates video with synchronized audio in eight steps, and at those eight steps it already beats base H3 on quality. Until now it needed data-center GPUs: its weights take 138 GiB. Today FastH3 V2 runs on a single consumer machine: an RTX 5090, RTX 4090 or RTX PRO 6000, a DGX Spark, or an Apple Silicon Mac. We are also releasing **FastH3 Trim**, an experimental smaller version that removes 8 of the 50 transformer blocks to run faster still.\n\nThis post continues [FastH3 Goes Local](https://haoailab.com/blogs/fasth3-local/), which brought FastH3 to the DGX Spark and the Mac and named the RTX family as the next target.\n\n## TL;DR[#](#tldr)\n\n- **FastH3 V2 now runs on one consumer GPU, a DGX Spark or a Mac.** We ship it in NVFP4 for Blackwell GPUs and DGX Spark, FP8 for the RTX 4090 and GPUs with less memory, and INT6 for Apple Silicon.\n- **Quantization keeps the quality.** Every FP4 layer uses activation scales calibrated on 1,000 prompts, so large activations are not clipped.\n- **FastH3 Trim is an experiment in making the model smaller.** It is 4.2× smaller than base H3, faster than V2 on every device and runs in as little as 8 GB of GPU memory, at minimal cost in quality.\n- **There is room to improve, and we will keep working on it.** Both models are eight-step distillations; we expect better quality and smaller models in future releases.\n\n## Same prompt, every machine[#](#same-prompt-every-machine)\n\nEach row is one machine. The left clip is FastH3 V2 and the right clip is FastH3 Trim, from the same prompt and seed at 832×480 for 5 s with audio. Turn the audio on. Three more prompts are under [More samples](#more-samples).\n\n**RTX 5090** 32 GB · NVFP4\n\n**RTX PRO 6000** 96 GB · NVFP4\n\n**RTX 4090** 24 GB · FP8\n\n**DGX Spark** 128 GB unified · NVFP4\n\n**Mac (M4 Max)** 36 GB unified · INT6\n\n## How fast it runs[#](#how-fast-it-runs)\n\nWe report two numbers per machine: a 5 s clip at 832×480 and a 5 s clip at 1344×768, both with audio. Times are end to end on a warm server, from prompt to finished MP4: text encoding, eight denoising steps, video and audio decoding, and export. Each is the median of two runs on each of two prompts.\n\n| Machine | Memory | V2, 480p | Trim, 480p | V2, 768p | Trim, 768p | \n|---|---|---|---|---|---|\n| RTX PRO 6000 | 96 GB | 13.5 s | 12.0 s | 36.5 s | 32.5 s | \n| RTX 5090 | 32 GB | 14.8 s | 13.4 s | 38.6 s | 35.4 s | \n| RTX 4090 | 24 GB | 54.6 s | 43.9 s | 154.6 s | 132.8 s | \n| RTX 4090, 16 GB limit | 16 GB | 79.9 s | 72.1 s | 153.7 s | 139.9 s | \n| RTX 4090, 12 GB limit | 12 GB | 91.2 s | 74.1 s | 170.7 s | 147.1 s | \n| RTX 4090, 8 GB limit | 8 GB | — | 82.0 s | — | — | \n| DGX Spark | 128 GB unified | 141.4 s | 125.8 s | — | 340.1 s | \n| 2× DGX Spark | 128 GB each | 87.2 s | 78.3 s | — | — | \n| Mac, M4 Max | 36 GB unified | — | 925.2 s | — | — | \n\n## Fitting H3 on one machine[#](#fitting-h3-on-one-machine)\n\nH3 is three networks. A Qwen3-VL text encoder reads the prompt, a 50-block diffusion transformer (DiT) denoises the video and audio latents together, and two VAEs decode the latents into frames and sound. In BF16 these weights total 137.7 GiB, and a 32 GB RTX 5090 cannot hold even the DiT. We reduced each network separately.\n\n- **Text encoder: 62.1 → 15.3 GiB.** H3 reads hidden state 50 of a 64-layer Qwen3-VL and never generates text, so we remove the last 14 layers and the language-model head without changing the features H3 uses. The remaining linear layers are stored in NVFP4.\n- **Transformer: 65.3 GiB in BF16.** For FastH3 V2 we store the attention, MLP and sparse-attention gate weights in NVFP4 (FP8 on GPUs without FP4 support). FastH3 Trim also removes 8 of the 50 blocks and replaces each block’s timestep projection with a rank-16 factorization, which brings it to 11.1 GiB in NVFP4.\n- **VAEs: 10.3 → 6.5 GiB.** We decode video with the[LynnReal lightweight video VAE](https://huggingface.co/stdstu123/LynnReal-Onmi-light-vae) , a distilled 26-block decoder with the same latent interface as the H3 VAE, loaded with[Kijai’s INT8 weights](https://huggingface.co/Kijai/MiniMax-H3-experimental) . It uses 2.3 GiB of GPU memory, and every device in this post uses the same one.\n\nSize matters for speed. On a 32 GB GPU, the question is whether the DiT can stay in GPU memory between requests. When it can, each request only moves the text encoder in and out.\n\n### RTX 5090[#](#rtx-5090)\n\nOur first FP4 export quantized only the MLPs, the setting we use on data-center GPUs. On a 5090 that left a 20 GB DiT, because the BF16 attention projections alone take 9.7 GB, and the text encoder no longer fit beside it. Every request moved the DiT to host memory and back: a 480p clip took 26.4 s, and a 768p clip did not fit at all.\n\nWith attention and the gate also in NVFP4, the DiT is 11.1 GiB and stays on the GPU, and the text encoder streams in one layer at a time. The same 480p clip now takes 13.4 s, and 768p fits.\n\nPyTorch’s pinned-memory allocator rounds each block up to a power of two, so 2.87 GiB of FP4 weights took 5.06 GiB of host RAM, enough to get the process killed in a 60 GB cloud container. We pin one exact-size buffer per module with `cudaHostRegister` instead, which uses 2.90 GiB.\n\n### RTX 4090 and smaller memory[#](#rtx-4090-and-smaller-memory)\n\nThe RTX 4090 has no FP4 tensor cores, so it uses FP8: 8-bit weights with one scale per output channel and 8-bit activations with one scale per token. PyTorch’s FP8 matrix multiply with these scales runs at about 70 TFLOPS on a 4090, slower than BF16 at about 160 TFLOPS. The per-tensor FP8 kernel runs at 220–305 TFLOPS, so we call it with unit scales and apply both scale vectors to the output in one fused pass. The result matches per-token, per-channel scaling and costs 5–10% more than per-tensor scaling.\n\nThree more changes bring the 4090 to 43.9 s for a 5 s clip:\n\n- **Sparse attention:** queries and keys are quantized to INT8 for the score computation, while values stay in BF16. The fine attention kernel runs 1.6× faster with about 0.6% relative error.\n- **Text encoder:** it streams to the GPU one layer at a time through exact-size pinned buffers, and one fused kernel expands its NVFP4 weights.\n- **VAE:** the same INT8 lightweight VAE as every other device, with a fused dequantization step and one shared quantized input for the Q, K and V projections. Decoded frames are bit-identical to the unoptimized path.\n\nFor the 16 GB, 12 GB and 8 GB rows, we cap GPU memory on the same 4090. Both models fit in 12 GB, and FastH3 Trim fits in 8 GB; a real card with less memory will be slower.\n\n### DGX Spark and Apple Silicon[#](#dgx-spark-and-apple-silicon)\n\nA DGX Spark keeps the text encoder, the transformer and both VAEs in its 128 GB of unified memory, so nothing moves between requests. It uses the same NVFP4 checkpoints and lightweight VAE as the 5090, and two Sparks split each request with sequence parallelism. In [FastH3 Goes Local](https://haoailab.com/blogs/fasth3-local/), a 5 s clip took 243 s on one Spark with the four-step preview model; the eight-step models now do twice as many denoising steps and still finish faster.\n\nOn Apple Silicon, both models run in MLX with INT6 weights, the NVFP4 text encoder and the lightweight VAE.\n\n## Four-bit weights without clipping[#](#four-bit-weights-without-clipping)\n\nNVFP4 stores each group of 16 values as 4-bit floats with a shared 8-bit scale, plus one scale per tensor that sets the overall range. For weights, we compute that per-tensor scale from the weights themselves. For activations it must be fixed before the data arrives. The simplest choice, a unit scale, covers magnitudes up to 6 × 448 = 2,688, and H3 activations are much larger.\n\nIn 40 of the 42 blocks, the input to the MLP output projection exceeds 2,688. In block 37 it reaches 368,640, 137 times the limit, and a unit scale clips these values on every forward pass. So we calibrate: we ran 1,000 prompts through the full eight-step sampler, recorded the largest input to each linear layer, and stored one static scale per layer in the checkpoint. FastH3 Trim has 294 calibrated layers, covering the attention projections, the MLPs and the sparse-attention gate. This extends the FastH3 V2 NVFP4 recipe from the MLPs to every quantized layer.\n\n## FastH3 Trim: an experiment in pruning[#](#fasth3-trim-an-experiment-in-pruning)\n\nPruning is our path to smaller models. We remove the blocks we measured as least important, which makes the model smaller and faster but takes some of what it learned with them. FastH3 Trim is our first step.\n\n### Choosing the blocks[#](#choosing-the-blocks)\n\nOur first pruned model chose blocks by their activations, which clearly beat removing blocks at even intervals. We recovered that model with teacher guidance and then used DMD to reduce its sampling steps. Motion coherence, fine detail and prompt adherence stayed weak. Quantization-aware distillation (QAD) did not beat post-training quantization in our side-by-side comparisons.\n\nFor FastH3 Trim we measured each block directly. Starting from base H3, we skipped one block at a time and recorded how much the video and audio predictions changed. We tested four examples (motion, speech, music and sound events) at three noise levels, for 600 measurements in total. Each block was ranked by the largest change it caused under any condition, so a block that matters to either video or audio is kept. The first and last blocks changed the output the most. The eight blocks we removed are all in the first half of the network.\n\n### Compressing the timestep conditioning[#](#compressing-the-timestep-conditioning)\n\nH3 conditions each block on the diffusion timestep through an AdaLN projection, which maps a 2,688-dimensional time embedding to six modulation vectors. These projections take 24 GiB in BF16 across the 50 blocks. Their input, however, is a smooth function of a single number, the timestep, so it uses very few of its 2,688 dimensions. FastH3 Trim replaces the projections with one shared 2,688→16 basis and a small projection per block. We store the factorized weights in FP16, because BF16 gives about 1.7× larger reconstruction error.\n\n### Training[#](#training)\n\nWe trained the new 42-block model directly with eight-step DMD2, using the FastH3 V2 objective. Base H3 initializes both the frozen teacher and the trainable critic, and attention is 80% sparse. The model samples at timesteps 999, 874, 749, 624, 500, 375, 250 and 125.\n\n## Where the time goes[#](#where-the-time-goes)\n\nFigure 5 splits one 5 s, 480p FastH3 Trim clip by stage on each machine. On the RTX PRO 6000 every model stays in GPU memory, and denoising is 78% of the 12.0 s. The RTX 5090 denoises just as fast (9.1 s) and finishes in 13.4 s, because only the text encoder and decoders move between host and GPU. The RTX 4090 has no FP4 tensor cores, so it runs FP8 and streams part of the transformer from host memory each step; denoising is 33.1 s of its 43.9 s. On a DGX Spark, video decoding takes a third of the time (42.6 s), and it is the next stage we will optimize. On a Mac, denoising is 91% of the time because attention still runs on a reference path rather than a tuned Metal kernel.\n\n## Limitations and what comes next[#](#limitations-and-what-comes-next)\n\n- **FastH3 Trim is experimental.** Busy scenes can show less detail than V2; use V2 when quality matters most.\n- **Memory tiers are emulated.** The 16 GB, 12 GB and 8 GB numbers cap memory on a 4090; real cards with that memory will be slower.\n\n## Get the models[#](#get-the-models)\n\n| Hardware / runtime | FastH3 V2 | FastH3 Trim | \n|---|---|---|\n| RTX 5090, RTX PRO 6000, DGX Spark (NVFP4) | [`FastVideo-FastH3-8-Step-V2-NVFP4-Consumer`](https://huggingface.co/FastVideo/FastVideo-FastH3-8-Step-V2-NVFP4-Consumer) | [`FastVideo-FastH3-Trim-8-Step-NVFP4`](https://huggingface.co/FastVideo/FastVideo-FastH3-Trim-8-Step-NVFP4) | \n| RTX 4090 and GPUs down to 8 GB (FP8) | [`FastVideo-FastH3-8-Step-V2-FP8`](https://huggingface.co/FastVideo/FastVideo-FastH3-8-Step-V2-FP8) | [`FastVideo-FastH3-Trim-8-Step-FP8`](https://huggingface.co/FastVideo/FastVideo-FastH3-Trim-8-Step-FP8) | \n| Apple Silicon (MLX INT6) | [`FastVideo-FastH3-8-Step-V2-MLX-INT6`](https://huggingface.co/FastVideo/FastVideo-FastH3-8-Step-V2-MLX-INT6) | [`FastVideo-FastH3-Trim-8-Step-MLX-INT6`](https://huggingface.co/FastVideo/FastVideo-FastH3-Trim-8-Step-MLX-INT6) | \n| ComfyUI checkpoints | [`FastVideo-FastH3-8-Step-V2-NVFP4-Comfy`](https://huggingface.co/FastVideo/FastVideo-FastH3-8-Step-V2-NVFP4-Comfy) | [`FastVideo-FastH3-Trim-Comfy`](https://huggingface.co/FastVideo/FastVideo-FastH3-Trim-Comfy) | \n| Source weights (BF16) | [`FastVideo-FastH3-8-Step-V2`](https://huggingface.co/FastVideo/FastVideo-FastH3-8-Step-V2) | [`FastVideo-FastH3-Trim-8-Step`](https://huggingface.co/FastVideo/FastVideo-FastH3-Trim-8-Step) | \n\nAll repositories are under the [FastVideo](https://huggingface.co/FastVideo) organization. Multi-GPU data-center serving keeps using [`FastVideo-FastH3-8-Step-V2-NVFP4`](https://huggingface.co/FastVideo/FastVideo-FastH3-8-Step-V2-NVFP4). Each FastVideo-format consumer repository includes the NVFP4 text encoder, the lightweight VAE and a `fastvideo_inference.json` file with the sampling schedule, which FastVideo reads automatically. On one RTX 5090:\n\n``` python\nimport os\n\nfrom huggingface_hub import hf_hub_download\n\nrepo = \"FastVideo/FastVideo-FastH3-8-Step-V2-NVFP4-Consumer\"\nos.environ[\"FASTVIDEO_H3_PARK_MODULES\"] = \"vae,audio_vae\"  # keep the transformer on the GPU during text encoding\nos.environ[\"FASTVIDEO_H3_ENCODER_LAYERWISE\"] = \"1\"  # stream the text encoder layer by layer\nos.environ[\"FASTVIDEO_H3_ADALN_TABLE\"] = hf_hub_download(repo, \"transformer/adaln_tables.pt\")  # skip 26 GB of AdaLN weights\n\nfrom fastvideo import VideoGenerator\n\ngenerator = VideoGenerator.from_config({\n    \"model_path\": repo,\n    \"engine\": {\n        \"num_gpus\": 1,\n        \"quantization\": {\"transformer_quant\": \"NVFP4\", \"layer_profile\": \"h3_dit_vsa\"},\n        \"offload\": {\"text_encoder\": True, \"pin_cpu_memory\": True},\n    },\n    \"pipeline\": {\"experimental\": {\"attention_backend\": \"VIDEO_SPARSE_ATTN_H3\", \"h3_sequential_load\": True}},\n})\ngenerator.generate_video(\n    prompt=\"A red fox leaps into deep snow at sunrise and pops back up with snow on its face.\",\n    height=480, width=832, num_frames=124, guidance_scale=1.0, output_path=\"out.mp4\",\n)\n```\n\nSwap in `FastVideo/FastVideo-FastH3-Trim-8-Step-NVFP4` for the faster experimental model, without the AdaLN table line.\n\n## Acknowledgements[#](#acknowledgements)\n\nFastH3 builds on [MiniMax H3](https://huggingface.co/MiniMaxAI/MiniMax-H3), and we thank the MiniMax team for releasing its weights and code.\n\nThe lightweight decoder is the [LynnReal Lightweight Video VAE](https://huggingface.co/stdstu123/LynnReal-Onmi-light-vae) ([paper](https://arxiv.org/abs/2609.15863), [code](https://github.com/LynnReal-AI/LynnReal-Omni)), loaded with the INT8 weights from [Kijai](https://huggingface.co/Kijai)’s [MiniMax-H3-experimental](https://huggingface.co/Kijai/MiniMax-H3-experimental). We thank both.\n\nWe thank the NVIDIA Enterprise Products team (Pengcheng Li and Cliff Woolley) for the Video Sparse Attention kernel, and the FlashInfer and NVIDIA Model Optimizer teams for the FP4 kernels and calibration tools. The FP4 sparse attention on RTX GPUs builds on [SageAttention](https://github.com/thu-ml/SageAttention).\n\nThe FastVideo team worked closely with [Nuva Lab](https://nuvalab.ai/), [NVIDIA FastGen](https://github.com/NVlabs/FastGen) (Julius Berner, Chao Liu, Arash Vahdat) and the NVIDIA Enterprise Products team on [FastH3](https://haoailab.com/blogs/fasth3-preview/). We also thank the [vLLM project](https://vllm.ai/), [NVIDIA](https://www.nvidia.com/en-us/) and [MBZUAI](https://mbzuai.ac.ae/) for their continued sponsorship and support of FastVideo.\n\n## FastVideo team[#](#fastvideo-team)\n\n**Contributor:** [Aryan Kumar](https://github.com/aryan5v)\n**Tech lead:** [Will Lin](https://github.com/SolitaryThinker)\n**Advisor:** [Hao Zhang](https://github.com/zhisbug)\n\n## More samples[#](#more-samples)\n\nFastH3 V2 (top) and FastH3 Trim (bottom) on one RTX 5090, 832×480, 5 s with audio.", "url": "https://wpnews.pro/news/fasth3-on-consumer-hardware-rtx-gpus-dgx-spark-and-macs-plus-fasth3-trim", "canonical_source": "https://haoailab.com/blogs/fasth3-rtx/", "published_at": "2026-10-06 07:00:00+00:00", "updated_at": "2026-10-06 20:17:29.281312+00:00", "lang": "en", "topics": ["generative-ai", "ai-research", "ai-infrastructure"], "entities": ["FastH3 V2", "FastH3 Trim", "RTX 5090", "RTX 4090", "RTX PRO 6000", "DGX Spark", "Apple Silicon", "Qwen3-VL"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/fasth3-on-consumer-hardware-rtx-gpus-dgx-spark-and-macs-plus-fasth3-trim", "markdown": "https://wpnews.pro/news/fasth3-on-consumer-hardware-rtx-gpus-dgx-spark-and-macs-plus-fasth3-trim.md", "text": "https://wpnews.pro/news/fasth3-on-consumer-hardware-rtx-gpus-dgx-spark-and-macs-plus-fasth3-trim.txt", "jsonld": "https://wpnews.pro/news/fasth3-on-consumer-hardware-rtx-gpus-dgx-spark-and-macs-plus-fasth3-trim.jsonld"}}