{"slug": "repair-before-reinforce-context-augmented-knowledge-graph-reasoning-for-multi", "title": "Repair Before Reinforce: Context-Augmented Knowledge Graph Reasoning for Multi-Hop Question Answering", "summary": "A context-augmented training framework for multi-hop question answering, validated on disease-specific knowledge graphs for Gastroparesis and Diabetes extracted with GraphMERT, consistently improved multi-hop performance over KG-only supervision when applied to the Qwen3-14B model, according to an arXiv paper (2609.12230v1). The framework attaches supporting triples from the same source text chunk to each primary KG triple to form a context graph, and an LLM-judged, history-aware adaptive repair pipeline brought the models to 100% accuracy on the cleaned retained one-hop validation sets. Reinforcement learning initialized from repaired supervised fine-tuning checkpoints produced larger and more stable gains on 3-hop, 4-hop, and 5-hop tasks across both diseases.", "body_md": "arXiv:2609.12230v1 Announce Type: new \nAbstract: Question-answering often requires reasoning across multiple connected facts rather than retrieving a single isolated relation. Knowledge graphs (KGs) provide a structured way to represent such facts, but training large language models (LLMs) only on isolated KG head-relation-tail triples may limit their ability to learn the surrounding context needed for multi-hop reasoning. In this work, we propose a context-augmented training framework for multi-hop question-answering. Although generally applicable, we validate the framework in the context of disease-specific KGs, extracted using a reliable KG extraction framework called GraphMERT, for Gastroparesis and Diabetes. For each primary KG triple, we attach supporting triples extracted from the same source text chunk to form a context graph (CG). This creates two supervision settings: KG-grounded supervision, which uses only the target KG triple or path, and CG-grounded supervision, which uses the target KG triple or path together with supporting context triples. We train the Qwen3-14B model using supervised fine-tuning (SFT) under both settings, producing KGModel and CGModel variants. To strengthen the lower-hop factual foundation of the models, we introduce an LLM-judged, history-aware adaptive repair pipeline that identifies unresolved one-hop failures, continually fine-tunes on targeted repair examples, and removes or quarantines problematic noisy triples. This repair stage enables the models to reach 100% accuracy on the cleaned retained one-hop validation sets. Finally, we employ reinforcement learning (RL) using lower-hop question-answer items and evaluate generalization on harder 3-hop, 4-hop, and 5-hop tasks. Across both diseases, context-augmented supervision consistently improves multi-hop performance over KG-only supervision. RL initialized from repaired SFT checkpoints yields larger and more stable gains.", "url": "https://wpnews.pro/news/repair-before-reinforce-context-augmented-knowledge-graph-reasoning-for-multi", "canonical_source": "https://arxiv.org/abs/2609.12230", "published_at": "2026-09-14 04:00:00+00:00", "updated_at": "2026-09-14 04:27:46.664383+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "natural-language-processing", "machine-learning"], "entities": ["GraphMERT", "Qwen3-14B", "Gastroparesis", "Diabetes", "KGModel", "CGModel"], "alternates": {"html": "https://wpnews.pro/news/repair-before-reinforce-context-augmented-knowledge-graph-reasoning-for-multi", "markdown": "https://wpnews.pro/news/repair-before-reinforce-context-augmented-knowledge-graph-reasoning-for-multi.md", "text": "https://wpnews.pro/news/repair-before-reinforce-context-augmented-knowledge-graph-reasoning-for-multi.txt", "jsonld": "https://wpnews.pro/news/repair-before-reinforce-context-augmented-knowledge-graph-reasoning-for-multi.jsonld"}}