# I am having two major issues with MiniMax-H3 Ref2VA

> Source: <https://discuss.huggingface.co/t/i-am-having-two-major-issues-with-minimax-h3-ref2va/190009#post_2>
> Published: 2026-10-09 18:53:49+00:00

For now, I tried testing a similar reference-audio issue using publicly available materials:

I would treat your goals as **two separate conditioning problems**, not as a reason to discard your WebUI or workflow: (a) preserve a recognizable speaker while generating *new, correctly timed dialogue* without an extra voice, and (b) preserve the *person* while replacing the original scene. I reproduced a related reference-audio anomaly using another person’s public example, but **not your exact overlapping-voices output**. The distinction matters throughout.

**1. How closely should Ref2VA reproduce the speaker?**

Using an audio clip to guide the speaker’s timbre and delivery is a documented Ref2VA use case. The [official full-reference format](https://github.com/MiniMax-AI/MiniMax-H3/blob/main/.agents/skills/h3-prompt-writing/references/ref-en.txt) distinguishes `reference` (use characteristics without directly copying the signal) from `fully_copy` (reuse the source audio). It does **not** promise a particular numerical voice-similarity score or perfect identity for every newly scripted line. A recognizable voice is a reasonable goal; it is useful to judge **voice resemblance**, **correct words**, **intelligibility**, **speech timing**, and **unwanted additional voices** separately. Success on one is not proof of success on the others.

**2. Is the apparent “two speakers at once” behavior known?**

There are neighboring reports: [ComfyUI #16155](https://github.com/Comfy-Org/ComfyUI/issues/16155) describes unexpected speech and timing errors with an audio reference; [MiniMax-H3 discussions #64](https://huggingface.co/MiniMaxAI/MiniMax-H3/discussions/64) and [#76](https://huggingface.co/MiniMaxAI/MiniMax-H3/discussions/76) include reports of reference material intruding into dialogue or clipped/gibberish vocalization; [MiniMax-H3 #17](https://github.com/MiniMax-AI/MiniMax-H3/issues/17) concerns voice binding between characters. Those are **related symptoms, not proof of the same bug**. Your “two voices” might mean genuinely simultaneous speakers, an extra utterance, speaker drift, or an echo/mixed track. Without your generated audio I cannot decide which.

In a separate test using #16155’s public assets, extra words were recognized in **3/3 seeds with audible reference audio**, versus **0/3** with equally long silence or zeroed audio latents **while retaining the reference block**. This points toward the *encoded reference values*, not merely the presence of a slot. It neither reproduces your overlapping voices nor identifies a root cause; the controls and limitations are below.

**3. Could three LoRAs interfere with the audio?**

Yes, potentially: a “visual” adapter may alter shared audiovisual layers, depending on the modules it targets. Your three LoRAs appear to be **sequential patches before the final sample**, not three separate renders. Compare one fixed-seed run with all custom LoRAs off; if the problem clears, restore them individually. My related anomaly occurred **without LoRAs or Heretic**, so those are not prerequisites for *that* behavior, but they could still contribute in *your* pipeline.

**4. Is strength `1.0` on each LoRA too much; should it be `0.5–0.7`?**

Three `1.0` settings are worth testing, but **they do not add up to a meaningful global “3.0 strength.”** Effects depend on target layers and adapter interactions. Nor is `0.5–0.7` a universal safe range. Try **all off → one at a time → necessary combinations**, then sweep strengths only where a difference appears; this preserves your visual goal.

**5. Could the pruned INT8 ConvRot checkpoint be responsible?**

Still an open variable. My runs used `minimax_h3_ref2va_pruned_int8_convrot.safetensors` and produced both clean ASR-control transcripts and reference-dependent extras: INT8 does not **uniformly** break dialogue, but might affect this failure. **No matched official-BF16 test exists.** With 12 GB VRAM, check the graph and one-seed contrasts first. The [model card](https://huggingface.co/MiniMaxAI/MiniMax-H3) and [native ComfyUI guide](https://github.com/Comfy-Org/docs/blob/main/tutorials/video/minimax/minimax-h3-native.mdx) document the distinct deployment artifacts; successful loading alone does not establish output parity.

**6. Must the audio be attached to both Ref2VA nodes (`400` and `587`)?**

**Not necessarily.** `MiniMaxH3ReferenceToVideo` builds conditioning and an initial audio-video latent; it is not a sampler. Your plan to attach audio to **final node `587`** may be sound. Check the **JSON actually submitted to `/prompt` after Gradio rewrites it**: trace both final BasicGuider conditioning **and** sampler latent, and establish what `400` still supplies. Adding audio to `400` without that trace could introduce a new constraint. The [native implementation](https://github.com/Comfy-Org/ComfyUI/blob/a4b5a045e56fc334903db8457b728b64e006119c/comfy_extras/nodes_minimax_h3.py) and [node docs](https://github.com/Comfy-Org/embedded-docs/blob/main/comfyui_embedded_docs/docs/MiniMaxH3ReferenceToVideo/en.md) are useful references.

**7. Why does the old background persist despite requesting a basement?**

A full reference picture encodes both **the person and the old scene**. If a first-frame/guide anchor survives in the actual graph, it may impose still more scene continuity. Distinguish *identity reference* from *frame preservation*: compare the full picture with a **subject-focused crop or background-removed reference**, keeping the target basement/action/seed fixed. You have **already** softened preservation wording, added a single-subject constraint, and switched `max` to `match`; their outcome is not established. I did not reproduce this background issue.

**8. `Is ref_image_size="max"` …?**

The public post ends at this heading, so I will not invent the missing question. `ref_image_size` controls **image preprocessing**, not a linear reference-strength setting. The [official node docs](https://github.com/Comfy-Org/docs/blob/main/built-in-nodes/MiniMaxH3ReferenceToVideo.mdx) explain that `match` scales references toward the output pixel area, while `max` keeps greater detail (up to the 2048px short-edge reference convention), at potentially much higher token/compute cost. Neither is a dedicated background switch; the result after your `match` change is unknown.

`587`, `400`, and the media loaders. This is mostly a static check and costs no render. Do not assume the original workflow screenshot equals the submitted graph.`<Audio N>` to `(S1)`, use `reference` rather than `fully_copy` for new dialogue, keep the sample clean, and specify the exact line. For timing, connect speech to an `match` setting; compare full-scene versus subject-focused image. Check whether anything actually anchors the first frame or reuses an old visual guide. Score The goal remains **new speech in the target voice** and **the same person in a new room**.

Useful optional evidence: one failing **submitted** `/prompt` JSON (private paths removed), a short generated-audio excerpt, the intended line, seed, reference duration, and checkpoint/adapter names. These are not prerequisites for the suggestions below.

**No-render audio check:** If the voice sounds doubled, compare **ComfyUI’s raw decoded audio** (if retained) with the **final Gradio/WebUI-exported or muxed audio** from the same run. Check left and right channels separately before downmixing. **Only after export?** Inspect muxing, duplicated tracks, and post-processing first. **Already in the raw audio?** Focus on generation, reference binding, or decoding. If the raw track was not saved, mark this comparison as unavailable instead of inferring an export bug; when both tracks exist, no new inference is needed.

``` php
Gradio workflow template
  -> file substitutions / prompt edits / adapter chain
  -> actual payload sent to POST /prompt
     -> LoadAudio -> 587.ref_audios.ref_audio_0
     -> 587.positive -> final BasicGuider (?)
     -> 587.LATENT   -> final sampler input (?)
     -> 400 outputs -> which actual consumers, if any?
     -> final H3 model / guider / sampler -> decode / output
```

The question marks are **edges to check**, not accusations that they are wrong. `ref_audios.ref_audio_0` is a **zero-based API field**, while model text uses the **one-based** label `<Audio 1>`; that apparent mismatch is normal. Other enabled audio-bearing reference media can affect label ordering, so inspect the packed/reference list rather than treating numbering as a universal equivalence. To classify “double voice,” distinguish: **overlapping independent speakers**, **extra speech before/after the line**, **timbre drift**, and **duplicated/reverberant audio**. These suggest different causes; ASR alone cannot distinguish them.

I used the **public image and voice WAV from someone else’s** [ComfyUI issue #16155](https://github.com/Comfy-Org/ComfyUI/issues/16155), where the requested line was `This park sucks.` around 2 seconds. This is not a reproduction of your particular reference, output, multi-node graph, or two-simultaneous-voices claim. My Colab setup used native ComfyUI, pruned INT8 H3 Ref2VA, a **standard (non-Heretic) Qwen**, **no LoRAs**, 512×288, 124 frames at 24 fps (~5.17 s), 20 sampling steps, and three deliberately reused seeds (`4242`, `5151`, `6262`).

The decisive panel kept the reference block in place while changing the encoded values:

| Reference input | Extra words recognized beyond requested line | Interpretation | 
|---|---|---|
| Original audible 7.05 s WAV | **3/3 seeds** , both Whisper models | Related anomaly reproduced in this setup | 
| Same-duration **time-reversed** WAV | **2/3** agreed;**1/3** ASR disagreement | Ordinary forward-intelligible English is not clearly required | 
| Same-duration **PCM silence, Audio-VAE encoded** | **0/3** , both ASR models | Reference block still exists; encoded silence is not a zero latent | 
| Original reference **latent zeroed after encoding** | **0/3** , both ASR models | Same reference block shape, but an artificial out-of-distribution control | 

Cropping the reference still left extras in **2/3** seeds. Two Whisper versions agreed on most observations but are not human listening tests; neither words alone nor waveform activity establish concurrent speakers, literal copying, or restored voice identity.

**Interpretation:** audio-derived conditioning values matter in this small panel, while the responsible feature and root cause remain unknown. The expanded controls follow.
