SolveEdit: Benchmarking Visual Problem Solving in Generative Models Researchers introduced SolveEdit, a benchmark for evaluating visual problem solving in generative models, according to the paper's headline. The benchmark targets real-world visual tasks such as arranging objects, repairing layouts, and tracing routes, which require understanding a scene, inferring what must change to achieve a goal, and realizing the change. The work positions visual problem solving as a distinct evaluation axis from abstract reasoning benchmarks. Machine intelligence is often evaluated through abstract reasoning problems, yet many real-world problems are visual, such as arranging objects, repairing layouts, or tracing routes. Solving these problems requires understanding a scene, inferring what must change to achieve a goal, and realizing th