cd /news/generative-ai/cuda-error-out-of-memory-minimax-onl… · home topics generative-ai article
[ARTICLE · art-91584] src=discuss.huggingface.co ↗ pub= topic=generative-ai verified=true sentiment=↓ negative

CUDA error: out of memory - (Minimax only)

A user reports that the MiniMax H3 video generation model fails with a CUDA out-of-memory error while Wan 2.2 runs normally, citing insufficient GPU memory during VAE encoding. The error occurs in the MiniMax H3 pipeline when adding image conditions, despite memory management attempts to pin model data to RAM.

read2 min views1 publishedAug 11, 2026

Hi!

i only get the following error message with minimax

Wan 2.2 is running normal without any memory error

  • To create a public link, set share=True

in launch() .

Model ‘ckpts\MiniMax-H3-FL2VA-pruned_rank8_int8_convrot.safetensors’ …

Text Encoder ‘ckpts\Qwen3-VL-32B-Instruct\Qwen3-VL-32B-Instruct-layer50_quanto_bf16_int8.safetensors’ … MiniMax H3 Video VAE ‘ckpts\MiniMax-H3-video_vae_fp16.safetensors’…

************ Memory Management for the GPU Poor (mmgp 3.7.12) by DeepBeepMeep ************

Switching to partial pinning since full requirements for pinned models is 51105.5 MB while estimated available reservable RAM is 26184.6 MB. You may increase the value of parameter ‘perc_reserved_

mem_max’ to a value higher than 0.40 to force full pinnning.

Partial pinning of data of ‘transformer’ to reserved RAM

Found 200 tied weights for a total of 0.00 MB, last : blocks.49.attn.q_proj.output_scale ↔ blocks.49.attn.v_proj.output_scale

The model was partially pinned to reserved RAM: 100 large blocks spread across 18547.98 MB

Hooked to model ‘transformer’ (MiniMaxH3Model)

Async plan for model ‘transformer’ : base size of 58.22 MB will be preloaded with a 370.96 MB async circular shuttle Partial pinning of data of ‘text_encoder’ to reserved RAM

Unable to pin more tensors for this model as the maximum reservable memory has been reached (5331.80).

The model was partially pinned to reserved RAM: 24 large blocks spread across 5331.80 MB

Hooked to model ‘text_encoder’ (Qwen3VLTextModel)

Async plan for model ‘text_encoder’ : base size of 1483.75 MB will be preloaded with a 465.16 MB async circular shuttle Unable to pin data of ‘vision_encoder’ to reserved RAM as there is no reserved RAM left. Transfer speed from RAM to VRAM may be slower.

Hooked to model ‘vision_encoder’ (Qwen3VLVisionModel)

Unable to pin data of ‘vae’ to reserved RAM as there is no reserved RAM left. Transfer speed from RAM to VRAM may be slower.

Hooked to model ‘vae’ (ModuleDict) Unable to pin data of ‘video_encoder’ to reserved RAM as there is no reserved RAM left. Transfer speed from RAM to VRAM may be slower.

Hooked to model ‘video_encoder’ (ModuleDict)

Unable to pin data of ‘audio_vae’ to reserved RAM as there is no reserved RAM left. Transfer speed from RAM to VRAM may be slower.

Hooked to model ‘audio_vae’ (MiniMaxH3AudioVAE)

Traceback (most recent call last): File “C:\pinokio\api\wan.git\app\wgp.py”, line 7589, in generate_media

samples = wan_model.generate( ^^^^^^^^^^^^^^^^^^^

File “C:\pinokio\api\wan.git\app\models\minimax_h3\pipeline.py”, line 31, in wrapped

return method(self, *args, **kwargs) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^

File “C:\pinokio\api\wan.git\app\venv\Lib\site-packages\torch\utils_contextlib.py”, line 124, in decorate_context

return func(*args, **kwargs) ^^^^^^^^^^^^^^^^^^^^^

File “C:\pinokio\api\wan.git\app\models\minimax_h3\pipeline.py”, line 407, in generate

self._add_image_condition(image_start, 0, presentation, visual_latents, keyframes)

File “C:\pinokio\api\wan.git\app\models\minimax_h3\pipeline.py”, line 267, in _add_image_condition

latent = self._encode_video(video) ^^^^^^^^^^^^^^^^^^^^^^^^^

File “C:\pinokio\api\wan.git\app\models\minimax_h3\pipeline.py”, line 231, in _encode_video

return self.vae.encode_condition(video.unsqueeze(0).to(device=self.device, dtype=self.vae._model_dtype), keep_all_latents=keep_all_latents).cpu() ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^

File “C:\pinokio\api\wan.git\app\venv\Lib\site-packages\torch\utils_device.py”, line 109, in torch_function

return func(*args, **kwargs) ^^^^^^^^^^^^^^^^^^^^^

torch.AcceleratorError: CUDA error: out of memory

Search for cudaErrorMemoryAllocation' in https://docs.nvidia.com/cuda/cuda-runtime-api/group__CUDART__TYPES.html for more information. CUDA kernel errors might be asynchronously reported at some other API call, so the stacktrace below might be incorrect. For debugging consider passing CUDA_LAUNCH_BLOCKING=1 Compile with

TORCH_USE_CUDA_DSA` to enable device-side assertions.

Error Queue autosaved successfully to error_queue.zip

── more in #generative-ai 4 stories · sorted by recency
── more on @minimax h3 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/cuda-error-out-of-me…] indexed:0 read:2min 2026-08-11 ·