{"slug": "i-am-having-two-major-issues-with-minimax-h3-ref2va", "title": "I am having two major issues with MiniMax-H3 Ref2VA", "summary": "A user troubleshooting MiniMax-H3's Ref2VA reference-audio feature reported two unresolved problems: preserving a recognizable speaker while generating new timed dialogue without an extra voice, and preserving the person while replacing the original scene. In a separate test using public assets from ComfyUI issue #16155, extra words were recognized in 3/3 seeds with audible reference audio versus 0/3 with equally long silence or zeroed audio latents while retaining the reference block, pointing to the encoded reference values rather than the presence of a slot. The user's runs used the minimax_h3_ref2va_pruned_int8_convrot.safetensors checkpoint, and no matched official-BF16 test exists.", "body_md": "For now, I tried testing a similar reference-audio issue using publicly available materials:\n\nI would treat your goals as **two separate conditioning problems**, not as a reason to discard your WebUI or workflow: (a) preserve a recognizable speaker while generating *new, correctly timed dialogue* without an extra voice, and (b) preserve the *person* while replacing the original scene. I reproduced a related reference-audio anomaly using another person’s public example, but **not your exact overlapping-voices output**. The distinction matters throughout.\n\n**1. How closely should Ref2VA reproduce the speaker?**\n\nUsing an audio clip to guide the speaker’s timbre and delivery is a documented Ref2VA use case. The [official full-reference format](https://github.com/MiniMax-AI/MiniMax-H3/blob/main/.agents/skills/h3-prompt-writing/references/ref-en.txt) distinguishes `reference` (use characteristics without directly copying the signal) from `fully_copy` (reuse the source audio). It does **not** promise a particular numerical voice-similarity score or perfect identity for every newly scripted line. A recognizable voice is a reasonable goal; it is useful to judge **voice resemblance**, **correct words**, **intelligibility**, **speech timing**, and **unwanted additional voices** separately. Success on one is not proof of success on the others.\n\n**2. Is the apparent “two speakers at once” behavior known?**\n\nThere are neighboring reports: [ComfyUI #16155](https://github.com/Comfy-Org/ComfyUI/issues/16155) describes unexpected speech and timing errors with an audio reference; [MiniMax-H3 discussions #64](https://huggingface.co/MiniMaxAI/MiniMax-H3/discussions/64) and [#76](https://huggingface.co/MiniMaxAI/MiniMax-H3/discussions/76) include reports of reference material intruding into dialogue or clipped/gibberish vocalization; [MiniMax-H3 #17](https://github.com/MiniMax-AI/MiniMax-H3/issues/17) concerns voice binding between characters. Those are **related symptoms, not proof of the same bug**. Your “two voices” might mean genuinely simultaneous speakers, an extra utterance, speaker drift, or an echo/mixed track. Without your generated audio I cannot decide which.\n\nIn a separate test using #16155’s public assets, extra words were recognized in **3/3 seeds with audible reference audio**, versus **0/3** with equally long silence or zeroed audio latents **while retaining the reference block**. This points toward the *encoded reference values*, not merely the presence of a slot. It neither reproduces your overlapping voices nor identifies a root cause; the controls and limitations are below.\n\n**3. Could three LoRAs interfere with the audio?**\n\nYes, potentially: a “visual” adapter may alter shared audiovisual layers, depending on the modules it targets. Your three LoRAs appear to be **sequential patches before the final sample**, not three separate renders. Compare one fixed-seed run with all custom LoRAs off; if the problem clears, restore them individually. My related anomaly occurred **without LoRAs or Heretic**, so those are not prerequisites for *that* behavior, but they could still contribute in *your* pipeline.\n\n**4. Is strength `1.0` on each LoRA too much; should it be `0.5–0.7`?**\n\nThree `1.0` settings are worth testing, but **they do not add up to a meaningful global “3.0 strength.”** Effects depend on target layers and adapter interactions. Nor is `0.5–0.7` a universal safe range. Try **all off → one at a time → necessary combinations**, then sweep strengths only where a difference appears; this preserves your visual goal.\n\n**5. Could the pruned INT8 ConvRot checkpoint be responsible?**\n\nStill an open variable. My runs used `minimax_h3_ref2va_pruned_int8_convrot.safetensors` and produced both clean ASR-control transcripts and reference-dependent extras: INT8 does not **uniformly** break dialogue, but might affect this failure. **No matched official-BF16 test exists.** With 12 GB VRAM, check the graph and one-seed contrasts first. The [model card](https://huggingface.co/MiniMaxAI/MiniMax-H3) and [native ComfyUI guide](https://github.com/Comfy-Org/docs/blob/main/tutorials/video/minimax/minimax-h3-native.mdx) document the distinct deployment artifacts; successful loading alone does not establish output parity.\n\n**6. Must the audio be attached to both Ref2VA nodes (`400` and `587`)?**\n\n**Not necessarily.** `MiniMaxH3ReferenceToVideo` builds conditioning and an initial audio-video latent; it is not a sampler. Your plan to attach audio to **final node `587`** may be sound. Check the **JSON actually submitted to `/prompt` after Gradio rewrites it**: trace both final BasicGuider conditioning **and** sampler latent, and establish what `400` still supplies. Adding audio to `400` without that trace could introduce a new constraint. The [native implementation](https://github.com/Comfy-Org/ComfyUI/blob/a4b5a045e56fc334903db8457b728b64e006119c/comfy_extras/nodes_minimax_h3.py) and [node docs](https://github.com/Comfy-Org/embedded-docs/blob/main/comfyui_embedded_docs/docs/MiniMaxH3ReferenceToVideo/en.md) are useful references.\n\n**7. Why does the old background persist despite requesting a basement?**\n\nA full reference picture encodes both **the person and the old scene**. If a first-frame/guide anchor survives in the actual graph, it may impose still more scene continuity. Distinguish *identity reference* from *frame preservation*: compare the full picture with a **subject-focused crop or background-removed reference**, keeping the target basement/action/seed fixed. You have **already** softened preservation wording, added a single-subject constraint, and switched `max` to `match`; their outcome is not established. I did not reproduce this background issue.\n\n**8. `Is ref_image_size=\"max\"` …?**\n\nThe public post ends at this heading, so I will not invent the missing question. `ref_image_size` controls **image preprocessing**, not a linear reference-strength setting. The [official node docs](https://github.com/Comfy-Org/docs/blob/main/built-in-nodes/MiniMaxH3ReferenceToVideo.mdx) explain that `match` scales references toward the output pixel area, while `max` keeps greater detail (up to the 2048px short-edge reference convention), at potentially much higher token/compute cost. Neither is a dedicated background switch; the result after your `match` change is unknown.\n\n`587`, `400`, and the media loaders. This is mostly a static check and costs no render. Do not assume the original workflow screenshot equals the submitted graph.`<Audio N>` to `(S1)`, use `reference` rather than `fully_copy` for new dialogue, keep the sample clean, and specify the exact line. For timing, connect speech to an `match` setting; compare full-scene versus subject-focused image. Check whether anything actually anchors the first frame or reuses an old visual guide. Score The goal remains **new speech in the target voice** and **the same person in a new room**.\n\nUseful optional evidence: one failing **submitted** `/prompt` JSON (private paths removed), a short generated-audio excerpt, the intended line, seed, reference duration, and checkpoint/adapter names. These are not prerequisites for the suggestions below.\n\n**No-render audio check:** If the voice sounds doubled, compare **ComfyUI’s raw decoded audio** (if retained) with the **final Gradio/WebUI-exported or muxed audio** from the same run. Check left and right channels separately before downmixing. **Only after export?** Inspect muxing, duplicated tracks, and post-processing first. **Already in the raw audio?** Focus on generation, reference binding, or decoding. If the raw track was not saved, mark this comparison as unavailable instead of inferring an export bug; when both tracks exist, no new inference is needed.\n\n``` php\nGradio workflow template\n  -> file substitutions / prompt edits / adapter chain\n  -> actual payload sent to POST /prompt\n     -> LoadAudio -> 587.ref_audios.ref_audio_0\n     -> 587.positive -> final BasicGuider (?)\n     -> 587.LATENT   -> final sampler input (?)\n     -> 400 outputs -> which actual consumers, if any?\n     -> final H3 model / guider / sampler -> decode / output\n```\n\nThe question marks are **edges to check**, not accusations that they are wrong. `ref_audios.ref_audio_0` is a **zero-based API field**, while model text uses the **one-based** label `<Audio 1>`; that apparent mismatch is normal. Other enabled audio-bearing reference media can affect label ordering, so inspect the packed/reference list rather than treating numbering as a universal equivalence. To classify “double voice,” distinguish: **overlapping independent speakers**, **extra speech before/after the line**, **timbre drift**, and **duplicated/reverberant audio**. These suggest different causes; ASR alone cannot distinguish them.\n\nI used the **public image and voice WAV from someone else’s** [ComfyUI issue #16155](https://github.com/Comfy-Org/ComfyUI/issues/16155), where the requested line was `This park sucks.` around 2 seconds. This is not a reproduction of your particular reference, output, multi-node graph, or two-simultaneous-voices claim. My Colab setup used native ComfyUI, pruned INT8 H3 Ref2VA, a **standard (non-Heretic) Qwen**, **no LoRAs**, 512×288, 124 frames at 24 fps (~5.17 s), 20 sampling steps, and three deliberately reused seeds (`4242`, `5151`, `6262`).\n\nThe decisive panel kept the reference block in place while changing the encoded values:\n\n| Reference input | Extra words recognized beyond requested line | Interpretation | \n|---|---|---|\n| Original audible 7.05 s WAV | **3/3 seeds** , both Whisper models | Related anomaly reproduced in this setup | \n| Same-duration **time-reversed** WAV | **2/3** agreed;**1/3** ASR disagreement | Ordinary forward-intelligible English is not clearly required | \n| Same-duration **PCM silence, Audio-VAE encoded** | **0/3** , both ASR models | Reference block still exists; encoded silence is not a zero latent | \n| Original reference **latent zeroed after encoding** | **0/3** , both ASR models | Same reference block shape, but an artificial out-of-distribution control | \n\nCropping the reference still left extras in **2/3** seeds. Two Whisper versions agreed on most observations but are not human listening tests; neither words alone nor waveform activity establish concurrent speakers, literal copying, or restored voice identity.\n\n**Interpretation:** audio-derived conditioning values matter in this small panel, while the responsible feature and root cause remain unknown. The expanded controls follow.", "url": "https://wpnews.pro/news/i-am-having-two-major-issues-with-minimax-h3-ref2va", "canonical_source": "https://discuss.huggingface.co/t/i-am-having-two-major-issues-with-minimax-h3-ref2va/190009#post_2", "published_at": "2026-10-09 18:53:49+00:00", "updated_at": "2026-10-09 19:24:16.837273+00:00", "lang": "en", "topics": ["generative-ai", "ai-tools", "ai-products"], "entities": ["MiniMax-H3", "Ref2VA", "ComfyUI", "Hugging Face", "MiniMaxAI", "minimax_h3_ref2va_pruned_int8_convrot.safetensors", "Heretic", "LoRA"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/i-am-having-two-major-issues-with-minimax-h3-ref2va", "markdown": "https://wpnews.pro/news/i-am-having-two-major-issues-with-minimax-h3-ref2va.md", "text": "https://wpnews.pro/news/i-am-having-two-major-issues-with-minimax-h3-ref2va.txt", "jsonld": "https://wpnews.pro/news/i-am-having-two-major-issues-with-minimax-h3-ref2va.jsonld"}}