arXiv:2608.28678v1 Announce Type: new Abstract: Scalable Vector Graphics (SVGs) power much of the modern visual ecosystem, yet state-of-the-art generative models focus almost entirely on rasterized images. We explore whether inference-time methods can unlock SVG generation capabilities in off-the-shelf vision-language models (VLMs). We systematically evaluate a constrained iterative refinement harness that combines visual feedback, structured editing, and constrained decoding to characterize the capabilities and limitations of current VLMs for SVG generation. Across multiple VLMs and generation settings, we find that constrained decoding improves compilation success rates, while iterative refinement reveals a deficit in visual reasoning and self-correction. Our results highlight both the promise and current limitations of using inference-time methods to adapt general-purpose VLMs for SVG generation.
Evaluating Constrained Iterative Refinement for Scalable Vector Graphics Generation with Off-the-Shelf VLMs
A new arXiv study (2608.28678v1) finds that constrained decoding improves compilation success rates for generating Scalable Vector Graphics (SVGs) with off-the-shelf vision-language models (VLMs), but iterative refinement exposes deficits in visual reasoning and self-correction. The research systematically evaluates a constrained iterative refinement harness combining visual feedback, structured editing, and constrained decoding across multiple VLMs and settings, highlighting both promise and limitations of inference-time methods for adapting general-purpose VLMs to SVG generation.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.