{"slug": "benchmarking-fine-tuning-and-retrieval-strategies-for-a-multimodal-language-on", "title": "Benchmarking Fine-tuning and Retrieval Strategies for a Multimodal Language Model on the NRC Reactor Operator Licensing Examination", "summary": "A 31-billion-parameter open-weight multimodal model (Gemma 4 31B-IT) fine-tuned with supervised fine-tuning and retrieval-augmented generation using fixed-size chunking passed 8 of 14 U.S. Nuclear Regulatory Commission Reactor Operator licensing examinations, meeting the 80% human passing criterion, while no configuration without fine-tuning passed any exam. The study, published on arXiv, found that the preferred chunking strategy reverses depending on the model's training state and that retrieval-augmented fine-tuning underperforms standard supervised fine-tuning in matching search environments.", "body_md": "arXiv:2607.22067v1 Announce Type: new\nAbstract: The integration of large language models (LLMs) into the nuclear power industry requires outputs grounded in domain-specific knowledge. This study evaluates a 31-billion-parameter open-weight multimodal model (Gemma 4 31B-IT) on its capacity to apply nuclear knowledge by benchmarking eight model-retrieval configurations against the U.S. Nuclear Regulatory Commission (NRC) Reactor Operator licensing examination. We evaluate 14 Generic Fundamentals Examinations (GFE) from the 2015-2021 March sittings (seven pressurized and seven boiling water reactor exams) using the standard 80% human passing criterion. The base model is compared against configurations utilizing supervised fine-tuning (SFT) on Gemini-distilled chain-of-thought (CoT) rationales, retrieval-augmented generation (RAG) with BM25 sparse retrieval over the U.S. Department of Energy Fundamentals Handbook, and retrieval-augmented fine-tuning (RAFT). Within the retrieval pipeline, we compare fixed-size sliding-window chunking against structure-aware chunking. The SFT configuration with fixed-size chunking RAG met the criterion on 8 of the 14 examinations, outperforming all alternatives, whereas no configuration without fine-tuning passed any. Aggregate accuracy reached 79.7%, with a confidence interval spanning the threshold, and 80.2% on PWR items specifically. Furthermore, two regularities emerged: the preferred chunking strategy reverses depending on the model's training state, and RAFT underperforms compared to standard SFT in matching search environments. These results demonstrate which combination of fine-tuning and search approaches achieves operator-level capabilities.", "url": "https://wpnews.pro/news/benchmarking-fine-tuning-and-retrieval-strategies-for-a-multimodal-language-on", "canonical_source": "https://arxiv.org/abs/2607.22067", "published_at": "2026-07-27 04:00:00+00:00", "updated_at": "2026-07-27 04:25:44.142575+00:00", "lang": "en", "topics": ["large-language-models", "artificial-intelligence", "ai-research"], "entities": ["Gemma 4 31B-IT", "U.S. Nuclear Regulatory Commission", "U.S. Department of Energy", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/benchmarking-fine-tuning-and-retrieval-strategies-for-a-multimodal-language-on", "markdown": "https://wpnews.pro/news/benchmarking-fine-tuning-and-retrieval-strategies-for-a-multimodal-language-on.md", "text": "https://wpnews.pro/news/benchmarking-fine-tuning-and-retrieval-strategies-for-a-multimodal-language-on.txt", "jsonld": "https://wpnews.pro/news/benchmarking-fine-tuning-and-retrieval-strategies-for-a-multimodal-language-on.jsonld"}}