For now, I tried testing a similar reference-audio issue using publicly available materials:
I would treat your goals as two separate conditioning problems, not as a reason to discard your WebUI or workflow: (a) preserve a recognizable speaker while generating new, correctly timed dialogue without an extra voice, and (b) preserve the person while replacing the original scene. I reproduced a related reference-audio anomaly using another person’s public example, but not your exact overlapping-voices output. The distinction matters throughout.
1. How closely should Ref2VA reproduce the speaker?
Using an audio clip to guide the speaker’s timbre and delivery is a documented Ref2VA use case. The official full-reference format distinguishes reference (use characteristics without directly copying the signal) from fully_copy (reuse the source audio). It does not promise a particular numerical voice-similarity score or perfect identity for every newly scripted line. A recognizable voice is a reasonable goal; it is useful to judge voice resemblance, correct words, intelligibility, speech timing, and unwanted additional voices separately. Success on one is not proof of success on the others.
2. Is the apparent “two speakers at once” behavior known?
There are neighboring reports: ComfyUI #16155 describes unexpected speech and timing errors with an audio reference; MiniMax-H3 discussions #64 and #76 include reports of reference material intruding into dialogue or clipped/gibberish vocalization; MiniMax-H3 #17 concerns voice binding between characters. Those are related symptoms, not proof of the same bug. Your “two voices” might mean genuinely simultaneous speakers, an extra utterance, speaker drift, or an echo/mixed track. Without your generated audio I cannot decide which.
In a separate test using #16155’s public assets, extra words were recognized in 3/3 seeds with audible reference audio, versus 0/3 with equally long silence or zeroed audio latents while retaining the reference block. This points toward the encoded reference values, not merely the presence of a slot. It neither reproduces your overlapping voices nor identifies a root cause; the controls and limitations are below.
3. Could three LoRAs interfere with the audio?
Yes, potentially: a “visual” adapter may alter shared audiovisual layers, depending on the modules it targets. Your three LoRAs appear to be sequential patches before the final sample, not three separate renders. Compare one fixed-seed run with all custom LoRAs off; if the problem clears, restore them individually. My related anomaly occurred without LoRAs or Heretic, so those are not prerequisites for that behavior, but they could still contribute in your pipeline.
4. Is strength 1.0 on each LoRA too much; should it be 0.5–0.7?
Three 1.0 settings are worth testing, but they do not add up to a meaningful global “3.0 strength.” Effects depend on target layers and adapter interactions. Nor is 0.5–0.7 a universal safe range. Try all off → one at a time → necessary combinations, then sweep strengths only where a difference appears; this preserves your visual goal.
5. Could the pruned INT8 ConvRot checkpoint be responsible?
Still an open variable. My runs used minimax_h3_ref2va_pruned_int8_convrot.safetensors and produced both clean ASR-control transcripts and reference-dependent extras: INT8 does not uniformly break dialogue, but might affect this failure. No matched official-BF16 test exists. With 12 GB VRAM, check the graph and one-seed contrasts first. The model card and native ComfyUI guide document the distinct deployment artifacts; successful alone does not establish output parity.
6. Must the audio be attached to both Ref2VA nodes (400 and 587)?
Not necessarily. MiniMaxH3ReferenceToVideo builds conditioning and an initial audio-video latent; it is not a sampler. Your plan to attach audio to final node 587 may be sound. Check the JSON actually submitted to /prompt after Gradio rewrites it: trace both final BasicGuider conditioning and sampler latent, and establish what 400 still supplies. Adding audio to 400 without that trace could introduce a new constraint. The native implementation and node docs are useful references.
7. Why does the old background persist despite requesting a basement?
A full reference picture encodes both the person and the old scene. If a first-frame/guide anchor survives in the actual graph, it may impose still more scene continuity. Distinguish identity reference from frame preservation: compare the full picture with a subject-focused crop or background-removed reference, keeping the target basement/action/seed fixed. You have already softened preservation wording, added a single-subject constraint, and switched max to match; their outcome is not established. I did not reproduce this background issue.
8. Is ref_image_size="max" …?
The public post ends at this heading, so I will not invent the missing question. ref_image_size controls image preprocessing, not a linear reference-strength setting. The official node docs explain that match scales references toward the output pixel area, while max keeps greater detail (up to the 2048px short-edge reference convention), at potentially much higher token/compute cost. Neither is a dedicated background switch; the result after your match change is unknown.
587, 400, and the media s. This is mostly a static check and costs no render. Do not assume the original workflow screenshot equals the submitted graph.<Audio N> to (S1), use reference rather than fully_copy for new dialogue, keep the sample clean, and specify the exact line. For timing, connect speech to an match setting; compare full-scene versus subject-focused image. Check whether anything actually anchors the first frame or reuses an old visual guide. Score The goal remains new speech in the target voice and the same person in a new room.
Useful optional evidence: one failing submitted /prompt JSON (private paths removed), a short generated-audio excerpt, the intended line, seed, reference duration, and checkpoint/adapter names. These are not prerequisites for the suggestions below.
No-render audio check: If the voice sounds doubled, compare ComfyUI’s raw decoded audio (if retained) with the final Gradio/WebUI-exported or muxed audio from the same run. Check left and right channels separately before downmixing. Only after export? Inspect muxing, duplicated tracks, and post-processing first. Already in the raw audio? Focus on generation, reference binding, or decoding. If the raw track was not saved, mark this comparison as unavailable instead of inferring an export bug; when both tracks exist, no new inference is needed.
Gradio workflow template
-> file substitutions / prompt edits / adapter chain
-> actual payload sent to POST /prompt
-> LoadAudio -> 587.ref_audios.ref_audio_0
-> 587.positive -> final BasicGuider (?)
-> 587.LATENT -> final sampler input (?)
-> 400 outputs -> which actual consumers, if any?
-> final H3 model / guider / sampler -> decode / output
The question marks are edges to check, not accusations that they are wrong. ref_audios.ref_audio_0 is a zero-based API field, while model text uses the one-based label <Audio 1>; that apparent mismatch is normal. Other enabled audio-bearing reference media can affect label ordering, so inspect the packed/reference list rather than treating numbering as a universal equivalence. To classify “double voice,” distinguish: overlapping independent speakers, extra speech before/after the line, timbre drift, and duplicated/reverberant audio. These suggest different causes; ASR alone cannot distinguish them.
I used the public image and voice WAV from someone else’s ComfyUI issue #16155, where the requested line was This park sucks. around 2 seconds. This is not a reproduction of your particular reference, output, multi-node graph, or two-simultaneous-voices claim. My Colab setup used native ComfyUI, pruned INT8 H3 Ref2VA, a standard (non-Heretic) Qwen, no LoRAs, 512×288, 124 frames at 24 fps (~5.17 s), 20 sampling steps, and three deliberately reused seeds (4242, 5151, 6262).
The decisive panel kept the reference block in place while changing the encoded values:
| Reference input | Extra words recognized beyond requested line | Interpretation |
|---|---|---|
| Original audible 7.05 s WAV | 3/3 seeds , both Whisper models | Related anomaly reproduced in this setup |
| Same-duration time-reversed WAV | 2/3 agreed;1/3 ASR disagreement | Ordinary forward-intelligible English is not clearly required |
| Same-duration PCM silence, Audio-VAE encoded | 0/3 , both ASR models | Reference block still exists; encoded silence is not a zero latent |
| Original reference latent zeroed after encoding | 0/3 , both ASR models | Same reference block shape, but an artificial out-of-distribution control |
Cropping the reference still left extras in 2/3 seeds. Two Whisper versions agreed on most observations but are not human listening tests; neither words alone nor waveform activity establish concurrent speakers, literal copying, or restored voice identity.
Interpretation: audio-derived conditioning values matter in this small panel, while the responsible feature and root cause remain unknown. The expanded controls follow.