{"slug": "larger-context-window-fewer-overcorrections-optimizing-prompts-and-batching-for", "title": "Larger Context Window, Fewer Overcorrections: Optimizing Prompts and Batching for Minimal-Edit Grammatical Error Correction", "summary": "A prompt-based approach to minimal-edit grammatical error correction reached an F0.5 score of 78.32 on the BEA-2019 test set, closing the gap to the fine-tuned single-model state of the art (Staruch et al., 2025) to 0.38 points, according to an arXiv paper (2609.10810v1). The method combines taxonomy-based instructions that bound correctable edits, batching multiple uncorrected sentences into a single input context to reduce overcorrection via an attention dilution effect, and LLM-assisted prompt optimization powered by Gemini 3.1-Pro. Code, prompts, and outputs are publicly available.", "body_md": "arXiv:2609.10810v1 Announce Type: new \nAbstract: Minimal-edit Grammatical Error Correction (GEC) is a challenging task for zero- and few-shot prompted Large Language Models (LLMs), which systematically overcorrect and degrade $F_{0.5}$ by rewriting well-formed spans. While fine-tuning provides an effective solution, it imposes substantial infrastructure demands. We introduce a prompt-based approach that closes the gap to fine-tuned models through three advances in GEC prompting methodology. First, we introduce taxonomy-based instructions to enforce minimal-edit constraints with a comprehensive list of grammatical error rules, equipping the LLM with a bounded, metric-aligned scope of correctable edits, which benefits the strongest models while remaining model-dependent overall. Second, we show that batching multiple uncorrected sentences into a single input context acts as a targeted regularizer against overcorrection, systematically reducing the edit rate across diverse LLM families; we hypothesize this arises from attention dilution effect induced by the bounded capacity of self-attention scores. Finally, LLM-assisted Prompt Optimization refines these instructions. Powered by Gemini 3.1-Pro, our prompt achieves $F_{0.5}=78.32$ on the BEA-2019 test set - establishing a new prompt-based SOTA while shrinking the gap to the fine-tuned single-model SOTA (Staruch et al., 2025) to a mere $0.38$ points. Code, prompts, and outputs are publicly available.", "url": "https://wpnews.pro/news/larger-context-window-fewer-overcorrections-optimizing-prompts-and-batching-for", "canonical_source": "https://arxiv.org/abs/2609.10810", "published_at": "2026-09-11 04:00:00+00:00", "updated_at": "2026-09-11 04:26:01.341421+00:00", "lang": "en", "topics": ["large-language-models", "natural-language-processing", "ai-research", "generative-ai"], "entities": ["BEA-2019", "Gemini 3.1-Pro", "Staruch et al., 2025", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/larger-context-window-fewer-overcorrections-optimizing-prompts-and-batching-for", "markdown": "https://wpnews.pro/news/larger-context-window-fewer-overcorrections-optimizing-prompts-and-batching-for.md", "text": "https://wpnews.pro/news/larger-context-window-fewer-overcorrections-optimizing-prompts-and-batching-for.txt", "jsonld": "https://wpnews.pro/news/larger-context-window-fewer-overcorrections-optimizing-prompts-and-batching-for.jsonld"}}