I am having two major issues with MiniMax-H3 Ref2VA A user troubleshooting MiniMax-H3's Ref2VA reference-audio feature reported two unresolved problems: preserving a recognizable speaker while generating new timed dialogue without an extra voice, and preserving the person while replacing the original scene. In a separate test using public assets from ComfyUI issue #16155, extra words were recognized in 3/3 seeds with audible reference audio versus 0/3 with equally long silence or zeroed audio latents while retaining the reference block, pointing to the encoded reference values rather than the presence of a slot. The user's runs used the minimax_h3_ref2va_pruned_int8_convrot.safetensors checkpoint, and no matched official-BF16 test exists. For now, I tried testing a similar reference-audio issue using publicly available materials: I would treat your goals as two separate conditioning problems , not as a reason to discard your WebUI or workflow: a preserve a recognizable speaker while generating new, correctly timed dialogue without an extra voice, and b preserve the person while replacing the original scene. I reproduced a related reference-audio anomaly using another person’s public example, but not your exact overlapping-voices output . The distinction matters throughout. 1. How closely should Ref2VA reproduce the speaker? Using an audio clip to guide the speaker’s timbre and delivery is a documented Ref2VA use case. The official full-reference format https://github.com/MiniMax-AI/MiniMax-H3/blob/main/.agents/skills/h3-prompt-writing/references/ref-en.txt distinguishes reference use characteristics without directly copying the signal from fully copy reuse the source audio . It does not promise a particular numerical voice-similarity score or perfect identity for every newly scripted line. A recognizable voice is a reasonable goal; it is useful to judge voice resemblance , correct words , intelligibility , speech timing , and unwanted additional voices separately. Success on one is not proof of success on the others. 2. Is the apparent “two speakers at once” behavior known? There are neighboring reports: ComfyUI 16155 https://github.com/Comfy-Org/ComfyUI/issues/16155 describes unexpected speech and timing errors with an audio reference; MiniMax-H3 discussions 64 https://huggingface.co/MiniMaxAI/MiniMax-H3/discussions/64 and 76 https://huggingface.co/MiniMaxAI/MiniMax-H3/discussions/76 include reports of reference material intruding into dialogue or clipped/gibberish vocalization; MiniMax-H3 17 https://github.com/MiniMax-AI/MiniMax-H3/issues/17 concerns voice binding between characters. Those are related symptoms, not proof of the same bug . Your “two voices” might mean genuinely simultaneous speakers, an extra utterance, speaker drift, or an echo/mixed track. Without your generated audio I cannot decide which. In a separate test using 16155’s public assets, extra words were recognized in 3/3 seeds with audible reference audio , versus 0/3 with equally long silence or zeroed audio latents while retaining the reference block . This points toward the encoded reference values , not merely the presence of a slot. It neither reproduces your overlapping voices nor identifies a root cause; the controls and limitations are below. 3. Could three LoRAs interfere with the audio? Yes, potentially: a “visual” adapter may alter shared audiovisual layers, depending on the modules it targets. Your three LoRAs appear to be sequential patches before the final sample , not three separate renders. Compare one fixed-seed run with all custom LoRAs off; if the problem clears, restore them individually. My related anomaly occurred without LoRAs or Heretic , so those are not prerequisites for that behavior, but they could still contribute in your pipeline. 4. Is strength 1.0 on each LoRA too much; should it be 0.5–0.7 ? Three 1.0 settings are worth testing, but they do not add up to a meaningful global “3.0 strength.” Effects depend on target layers and adapter interactions. Nor is 0.5–0.7 a universal safe range. Try all off → one at a time → necessary combinations , then sweep strengths only where a difference appears; this preserves your visual goal. 5. Could the pruned INT8 ConvRot checkpoint be responsible? Still an open variable. My runs used minimax h3 ref2va pruned int8 convrot.safetensors and produced both clean ASR-control transcripts and reference-dependent extras: INT8 does not uniformly break dialogue, but might affect this failure. No matched official-BF16 test exists. With 12 GB VRAM, check the graph and one-seed contrasts first. The model card https://huggingface.co/MiniMaxAI/MiniMax-H3 and native ComfyUI guide https://github.com/Comfy-Org/docs/blob/main/tutorials/video/minimax/minimax-h3-native.mdx document the distinct deployment artifacts; successful loading alone does not establish output parity. 6. Must the audio be attached to both Ref2VA nodes 400 and 587 ? Not necessarily. MiniMaxH3ReferenceToVideo builds conditioning and an initial audio-video latent; it is not a sampler. Your plan to attach audio to final node 587 may be sound. Check the JSON actually submitted to /prompt after Gradio rewrites it : trace both final BasicGuider conditioning and sampler latent, and establish what 400 still supplies. Adding audio to 400 without that trace could introduce a new constraint. The native implementation https://github.com/Comfy-Org/ComfyUI/blob/a4b5a045e56fc334903db8457b728b64e006119c/comfy extras/nodes minimax h3.py and node docs https://github.com/Comfy-Org/embedded-docs/blob/main/comfyui embedded docs/docs/MiniMaxH3ReferenceToVideo/en.md are useful references. 7. Why does the old background persist despite requesting a basement? A full reference picture encodes both the person and the old scene . If a first-frame/guide anchor survives in the actual graph, it may impose still more scene continuity. Distinguish identity reference from frame preservation : compare the full picture with a subject-focused crop or background-removed reference , keeping the target basement/action/seed fixed. You have already softened preservation wording, added a single-subject constraint, and switched max to match ; their outcome is not established. I did not reproduce this background issue. 8. Is ref image size="max" …? The public post ends at this heading, so I will not invent the missing question. ref image size controls image preprocessing , not a linear reference-strength setting. The official node docs https://github.com/Comfy-Org/docs/blob/main/built-in-nodes/MiniMaxH3ReferenceToVideo.mdx explain that match scales references toward the output pixel area, while max keeps greater detail up to the 2048px short-edge reference convention , at potentially much higher token/compute cost. Neither is a dedicated background switch; the result after your match change is unknown. 587 , 400 , and the media loaders. This is mostly a static check and costs no render. Do not assume the original workflow screenshot equals the submitted graph.