Humasak Simanjuntak, Tamara Yunika Sianipar, Bronson T. M Siallagan, Difya Laurensya Ambarita, Arlinta Barus
Competent domain application with practical cultural value and a useful negative result, but minimal methodological novelty, tiny dataset, and poor absolute generation quality (FID 270-330) limit broad impact.
The traditional Batak Ulos weaving industry faces growing challenges in producing diverse, innovative motifs due to limitations in conventional, manually driven design methods. This study proposes a multimodal generative framework integrating a fine-tuned Latent Diffusion Model (Stable Diffusion XL v1.0 via LoRA) with a Multimodal Large Language Model (LLaMA 1.5-7B) to enable controllable, culturally faithful Ulos motif generation. Four complementary conditioning mechanisms: text, image, representation, and semantic map (via ControlNet) jointly guide the generation process, each governing a distinct aspect from semantic intent to spatial layout. A five level ablation study across three scenarios (shape transformation, colour variation, and high-complexity input) shows that conditioning effectiveness is not proportional to the number of mechanisms combined: Text + Image + Semantic Map achieved the best FID (270) and CLIP Score (0.65 - 0.70) but the weakest SSIM (0.65), while Text + Image + Representation offered the best overall balance, with stable SSIM (0.84) and competitive FID (280). Combining all four mechanisms yielded the weakest FID (330), indicating conflicting optimization signals. Qualitative evaluation by nine weavers and thirty public participants confirmed statistically significant positive acceptance (Wilcoxon, p=0.007 and p<0.001, respectively). A web-based prototype supporting text-to-image and image-to-image generation was also developed, offering a practical digital design tool for cultural heritage preservation.
Core Contribution. This paper presents an applied generative framework for producing Batak Ulos weaving motifs by combining four conditioning mechanisms into a fine-tuned Stable Diffusion XL pipeline: text prompts, image (latent) conditioning, "representation" conditioning derived from a fine-tuned LLaVA-1.5-7B captioner, and semantic-map conditioning via ControlNet. The central empirical finding is that conditioning effectiveness is *not* monotonic in the number of mechanisms combined — Text+Image+Representation gives the best fidelity/structure balance (SSIM ≈0.84, FID ≈280), Text+Image+Semantic Map gives best FID/CLIP but poor SSIM, and combining all four *degrades* FID (≈330), interpreted as conflicting optimization signals. The work also contributes a curated (small) Ulos dataset, a structured captioning template, a shape-based semantic segmentation pipeline (Otsu/CCA/Hough/contour) for map generation, and a deployed web prototype. The contribution is primarily a domain application and an engineering integration rather than a methodological novelty — every component (SDXL, LoRA, LLaVA, ControlNet, EncDiff-style cross-attention) is off-the-shelf.
Methodological Rigor. The experimental design is reasonably thorough for an applied paper: a five-level ablation, three task scenarios, parameter sweeps over diffusion steps/strength/guidance, three complementary automatic metrics, and a human evaluation with nonparametric statistical testing (Wilcoxon, Cronbach's alpha reliability). However, several rigor concerns are significant. First, the dataset is tiny — 125 curated images augmented to 431, with only 7 test images — which severely limits statistical reliability of FID (FID is unstable and biased at small sample sizes). Second, the absolute FID values (270–330) are extremely high by generative-modeling standards (strong models typically report FID <50), suggesting the generated distribution is far from the reference; the paper does not reckon with what such high FID means for the "culturally faithful" claim. Third, there is no quantitative comparison against the prior Ulos/batik methods it critiques (SinGAN, StyleGAN, GenBatik, the authors' own prior diffusion work). The human evaluation, while positive and significant, tests only against a neutral Likert midpoint — a weak bar — and involves small, potentially non-independent samples (9 weavers, 30 public). The claim that "more conditioning isn't better" rests on single-run metric differences without variance estimates or significance testing across configurations.
Potential Impact. The impact is niche. The primary beneficiaries are researchers and practitioners working on computational cultural-heritage preservation, particularly Indonesian textile digitization, and possibly the local weaving industry via the prototype tool. The general lesson — that stacking conditioning signals can produce conflicting gradients rather than additive gains — is a practically useful observation for practitioners building multi-conditioned diffusion pipelines, but it is neither surprising nor rigorously established here. The framework could serve as a template for other ethnic-motif domains (batik, songket, ikat), giving it some transferability, but it does not advance the underlying methods.
Timeliness & Relevance. Generative AI for cultural heritage is a genuinely active and relevant area, and the combination of MLLM-derived semantic conditioning with diffusion is current. The topic addresses a real practical need (motif diversification for a declining craft). However, the technical framing does not target a mainstream ML bottleneck.
Strengths. (1) Thorough, well-documented integration of multiple modern components with clear architectural detail (channel dimensions, LoRA ranks, ControlNet residual injection). (2) Combined quantitative + human evaluation, with domain experts (weavers) — a meaningful and often-neglected validation for heritage applications. (3) A deployed prototype adds translational value. (4) Clear, well-organized writing with explicit conditioning mechanism descriptions.
Limitations. (1) Minimal methodological novelty — an assembly of existing tools. (2) Very small dataset and test set undermine metric reliability. (3) Absolute generation quality (FID) is poor and unexamined. (4) No head-to-head quantitative baseline against prior art despite extensive related-work critique. (5) Reproducibility limited by proprietary/field-collected data and no code release. (6) The disentanglement claims (EncDiff adaptation) are asserted but not empirically demonstrated. (7) Interpretation of ablation results is post-hoc and not statistically supported.
Other observations. The paper is honest about limitations (dataset size, evaluator pool). The interesting negative result — degradation when combining all four mechanisms — is the most citable insight, but its evidential basis is thin. The dataset, if released, would have modest community value given its size. Arxiv metadata shows a future date (2026), consistent with a recent preprint.
Overall, this is a competent applied/engineering paper with genuine practical and cultural value in a narrow domain, but limited methodological originality, weak absolute generation quality, and evidence constrained by data scarcity. Expected scientific impact is modest.
Generated Sep 17, 2026
Competent domain application with practical cultural value and a useful negative result, but minimal methodological novelty, tiny dataset, and poor absolute generation quality (FID 270-330) limit broad impact.