arXiv:2609.17987v1 Announce Type: new Abstract: The traditional Batak Ulos weaving industry faces growing challenges in producing diverse, innovative motifs due to limitations in conventional, manually driven design methods. This study proposes a multimodal generative framework integrating a fine-tuned Latent Diffusion Model (Stable Diffusion XL v1.0 via LoRA) with a Multimodal Large Language Model (LLaMA 1.5-7B) to enable controllable, culturally faithful Ulos motif generation. Four complementary conditioning mechanisms: text, image, representation, and semantic map (via ControlNet) jointly guide the generation process, each governing a distinct aspect from semantic intent to spatial layout. A five level ablation study across three scenarios (shape transformation, colour variation, and high-complexity input) shows that conditioning effectiveness is not proportional to the number of mechanisms combined: Text + Image + Semantic Map achieved the best FID (270) and CLIP Score (0.65 - 0.70) but the weakest SSIM (0.65), while Text + Image + Representation offered the best overall balance, with stable SSIM (0.84) and competitive FID (280). Combining all four mechanisms yielded the weakest FID (330), indicating conflicting optimization signals. Qualitative evaluation by nine weavers and thirty public participants confirmed statistically significant positive acceptance (Wilcoxon, p=0.007 and p<0.001, respectively). A web-based prototype supporting text-to-image and image-to-image generation was also developed, offering a practical digital design tool for cultural heritage preservation.
Multimodal Conditioning of Fine-Tuned Stable Diffusion XL for Controllable and Culturally Faithful Ulos Motif Generation
A study on arXiv (2609.17987v1) reports that a multimodal framework combining a LoRA fine-tuned Stable Diffusion XL v1.0 with the LLaMA 1.5-7B multimodal large language model can generate controllable, culturally faithful Batak Ulos motifs. In a five-level ablation across shape transformation, colour variation, and high-complexity input, Text + Image + Semantic Map conditioning achieved the best FID (270) and CLIP Score (0.65-0.70) but the weakest SSIM (0.65), Text + Image + Representation gave the best balance with SSIM 0.84 and FID 280, and combining all four mechanisms produced the weakest FID (330) due to conflicting optimization signals. Nine weavers and thirty public participants rated the output positively with statistical significance (Wilcoxon p=0.007 and p<0.001), and a web-based text-to-image and image-to-image prototype was built as a digital design tool for cultural heritage preservation.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.