arXiv:2607.14099v1 Announce Type: new Abstract: Deploying Vision-Language Models (VLMs) in real-world settings requires not only strong visual reasoning but also stability under sustained conversational pressure. We introduce Just Keep Prompting (JKP), a multi-turn evaluation framework that measures VLM epistemic stability when users repeatedly challenge, question, or contradict a model's answer. JKP probes models for up to 10 follow-up turns using three strategies: Adversarial Negation (repeated rejection), Pure Socratic Interrogation (repeated calls to reassess certainty), and Context-Aware Socratic Summarization (reflecting the model's prior rationale back before asking for reconsideration). We evaluate GPT-4o, Gemini 2.5 Pro, and Qwen3-VL-30B on a subset of the STAR benchmark across 720 multi-turn runs. Aggregate accuracy changes modestly from Turn 0 to Turn 10, but trajectory-level analysis reveals substantial instability: correct answers regress, wrong answers recover, and many runs exhibit repeated answer flipping. Repeated prompting has bounded upside and often acts as a destabilizer rather than a reasoning aid. The effect is strongly model-dependent: Qwen3-VL-30B achieves the highest final accuracy but becomes confidently wrong under direct contradiction; Gemini 2.5 Pro is comparatively stable but token-expensive; GPT-4o is the most brittle and oscillatory. These findings reveal that multi-turn VLM evaluation captures not just additional reasoning but pressure-response profiles: how models trade off visual grounding, calibration, and conversational compliance under repeated challenge.
Just Keep Prompting: Evaluating Repetitive Socratic Prompting in VLMs
A new evaluation framework called Just Keep Prompting (JKP) reveals that repeatedly challenging vision-language models (VLMs) with follow-up questions destabilizes their answers rather than improving reasoning, according to a study posted on arXiv. Testing GPT-4o, Gemini 2.5 Pro, and Qwen3-VL-30B on the STAR benchmark across 720 multi-turn runs, researchers found that while aggregate accuracy changes modestly from Turn 0 to Turn 10, trajectory-level analysis shows substantial instability: correct answers regress, wrong answers recover, and many runs exhibit repeated answer flipping. The effect is strongly model-dependent, with GPT-4o being the most brittle and oscillatory, Qwen3-VL-30B achieving the highest final accuracy but becoming confidently wrong under direct contradiction, and Gemini 2.5 Pro being comparatively stable but token-expensive.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.