MiniMax-H3 ref2va: aligned video guides (IC-LoRA style v2v) for ai-toolkit A developer contributed a patch to ai-toolkit's MiniMax-H3 diffusion model extension that fixes reference-to-video-alignment (ref2va) handling of control images and videos in IC-LoRA-style video-to-video workflows. The change refactors keyframe conversion into a helper and preserves the nested per-batch-item structure of control_tensor_list, which previously collapsed multiple references into a single block and caused an "image features and image tokens do not match" error in the Qwen3-VL processor. | | diff --git a/extensions built in/diffusion models/minimax h3/minimax h3.py b/extensions built in/diffusion models/minimax h3/minimax h3.py | | | index 4c21927..2509789 100644 | | | --- a/extensions built in/diffusion models/minimax h3/minimax h3.py | | | +++ b/extensions built in/diffusion models/minimax h3/minimax h3.py | | | @@ -537,44 +537,56 @@ class MinimaxH3Model BaseModel : | | | control tensors arrive in 0, 1 ; the Qwen3-VL processor wants PIL | | | keyframes per prompt = None len prompt | | | if control images is not None: | | | - if isinstance control images, torch.Tensor : | | | - images = control images i for i in range control images.shape 0 | | | - elif isinstance control images, list : | | | - images = | | | - c 0 if isinstance c, torch.Tensor and c.ndim == 4 else c | | | - for c in control images | | | - | | | - else: | | | - images = control images | | | - pil images = | | | - for img in images: | | | + | | | + def to keyframe img : | | | if isinstance img, torch.Tensor : | | | if img.ndim == 4: | | | img = img 0 | | | arr = img.float .clamp 0, 1 255 .round .to torch.uint8 | | | - pil images.append | | | - self. present image control | | | - Image.fromarray arr.permute 1, 2, 0 .cpu .numpy | | | - | | | + return self. present image control | | | + Image.fromarray arr.permute 1, 2, 0 .cpu .numpy | | | | | | - elif isinstance img, str : | | | + if isinstance img, str : | | | a control VIDEO path: 2 fps timestamped presentation over | | | the SAME frames the latent rows use dataset treatment | | | when caching training embeds, sample-length at sampling | | | ds cfg = getattr self, " ref video dataset config", None | | | - pil images.append | | | - load video ref for te | | | - self, img, ds cfg, max frames=self. sample ref max frames | | | - | | | + return load video ref for te | | | + self, img, ds cfg, max frames=self. sample ref max frames | | | | | | - else: | | | - pil images.append img | | | - if len pil images == 1: | | | - keyframes per prompt = pil images len prompt | | | - elif len pil images == len prompt : | | | - keyframes per prompt = img for img in pil images | | | + return img | | | + | | | + batch.control tensor list is item ref : several references per | | | + batch item. It must stay nested per item -- collapsing it emits a | | | + single