cd /news/generative-ai/multimodal-conditioning-of-fine-tune… · home topics generative-ai article
[ARTICLE · art-132258] src=machinebrief.com ↗ pub= topic=generative-ai verified=true sentiment=↑ positive

Multimodal Conditioning of Fine-Tuned Stable Diffusion XL for Controllable and Culturally Faithful Ulos Motif Generation

A study on arXiv (2609.17987v1) reports that a multimodal framework combining a LoRA fine-tuned Stable Diffusion XL v1.0 with the LLaMA 1.5-7B multimodal large language model can generate controllable, culturally faithful Batak Ulos motifs. In a five-level ablation across shape transformation, colour variation, and high-complexity input, Text + Image + Semantic Map conditioning achieved the best FID (270) and CLIP Score (0.65-0.70) but the weakest SSIM (0.65), Text + Image + Representation gave the best balance with SSIM 0.84 and FID 280, and combining all four mechanisms produced the weakest FID (330) due to conflicting optimization signals. Nine weavers and thirty public participants rated the output positively with statistical significance (Wilcoxon p=0.007 and p<0.001), and a web-based text-to-image and image-to-image prototype was built as a digital design tool for cultural heritage preservation.

by read1 min views1 publishedSep 17, 2026

arXiv:2609.17987v1 Announce Type: new Abstract: The traditional Batak Ulos weaving industry faces growing challenges in producing diverse, innovative motifs due to limitations in conventional, manually driven design methods. This study proposes a multimodal generative framework integrating a fine-tuned Latent Diffusion Model (Stable Diffusion XL v1.0 via LoRA) with a Multimodal Large Language Model (LLaMA 1.5-7B) to enable controllable, culturally faithful Ulos motif generation. Four complementary conditioning mechanisms: text, image, representation, and semantic map (via ControlNet) jointly guide the generation process, each governing a distinct aspect from semantic intent to spatial layout. A five level ablation study across three scenarios (shape transformation, colour variation, and high-complexity input) shows that conditioning effectiveness is not proportional to the number of mechanisms combined: Text + Image + Semantic Map achieved the best FID (270) and CLIP Score (0.65 - 0.70) but the weakest SSIM (0.65), while Text + Image + Representation offered the best overall balance, with stable SSIM (0.84) and competitive FID (280). Combining all four mechanisms yielded the weakest FID (330), indicating conflicting optimization signals. Qualitative evaluation by nine weavers and thirty public participants confirmed statistically significant positive acceptance (Wilcoxon, p=0.007 and p<0.001, respectively). A web-based prototype supporting text-to-image and image-to-image generation was also developed, offering a practical digital design tool for cultural heritage preservation.

── more in #generative-ai 4 stories · sorted by recency
── more on @stable diffusion xl v1.0 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/multimodal-condition…] indexed:0 read:1min 2026-09-17 ·