# VLM Fine-Tuning for End-to-End Combinatorial Optimization

> Source: <https://www.machinebrief.com/news/vlm-fine-tuning-for-end-to-end-combinatorial-optimization-j0j7>
> Published: 2026-09-30 04:00:00+00:00

arXiv:2609.37175v1 Announce Type: new 
Abstract: Large language models (LLMs) have provided a unified interface for end-to-end combinatorial optimization (CO), but textual serialization alone may obscure spatial and relational structures that are important for generating effective CO solutions. This paper presents a general-purpose vision-language solver that augments textual instance descriptions with input-derived visual representations. A single vision-language model (VLM) is applied across different CO tasks and trained using supervised fine-tuning followed by verifier-guided reinforcement learning. While the visual inputs contain no gold solutions or solution-derived information, our experiments show that the VLM generally improves solution quality over its text-only counterpart, with particularly clear gains on more complex CO problems such as CVRP and JSSP. The advantage of visual information is more pronounced at large problem scales.
