{"slug": "open-meditron-an-auditable-pipeline-for-clinical-llms", "title": "Open Meditron: An Auditable Pipeline for Clinical LLMs", "summary": "Researchers introduced Fully Open Meditron, the first fully open pipeline for building clinical large language models (LLMs), comprising a clinician-audited training corpus, reproducible framework, and evaluation protocol. The Apertus-70B-MeditronFO variant improved +6.6 points over its base (47.2% to 53.8%) on aggregate medical benchmarks, establishing a new fully open state-of-the-art. The pipeline enforces system-wide decontamination and end-to-end validation by a four-physician panel, demonstrating that fully open models can achieve top performance without sacrificing auditability.", "body_md": "# Computer Science > Artificial Intelligence\n\n[Submitted on 15 May 2026 (\n\n[v1](https://arxiv.org/abs/2605.16215v1)), last revised 29 May 2026 (this version, v2)]# Title:Fully Open Meditron: An Auditable Pipeline for Clinical LLMs\n\n[View PDF](/pdf/2605.16215)\n\n[HTML (experimental)](https://arxiv.org/html/2605.16215v2)\n\nAbstract:Clinical decision support systems (CDSS) require scrutable, auditable pipelines that enable rigorous, reproducible validation. Yet current LLM-based CDSS remain largely opaque. Most \"open\" models are open-weight only, releasing parameters while withholding the data provenance, curation procedures, and generation pipelines that determine model behavior. Fully Open (FO) models, which expose the complete training stack end-to-end, do not currently exist in medicine. We introduce Fully Open Meditron, the first fully open pipeline for building LLM-CDSS, comprising a clinician-audited training corpus, a reproducible data construction and training framework, and a use-aligned evaluation protocol. The corpus unifies eight public medical QA datasets into a normalized conversational format and expands coverage with three clinician-vetted synthetic extensions: exam-style QA, guideline-grounded QA derived from 46,469 clinical practice guidelines, and clinical vignettes. The pipeline enforces system-wide decontamination, gold-label resampling of teacher generations, and end-to-end validation by a four-physician panel. We evaluate using an LLM-as-a-judge protocol over expert-written clinical vignettes, calibrated against 204 human raters. We apply the recipe to five FO base models (Apertus-70B/8B-Instruct, OLMo-2-32B-SFT, EuroLLM-22B/9B-Instruct). All MeditronFO variants are preferred over their bases. Apertus-70B-MeditronFO improves +6.6 points over its base (47.2% to 53.8%) on aggregate medical benchmarks, establishing a new FO SoTA. Gemma-3-27B-MeditronFO is preferred over MedGemma in 58.6% of LLM-as-a-judge comparisons and outperforms it on HealthBench (58% vs 55.9%). These results show that fully open pipelines can achieve state-of-the-art domain-specific performance without sacrificing auditability or reproducibility.\n\n## Submission history\n\nFrom: Xavier Theimer-Lienhard [[view email](/show-email/92d2e130/2605.16215)]\n\n**Fri, 15 May 2026 17:29:08 UTC (603 KB)**\n\n[[v1]](/abs/2605.16215v1)**[v2]** Fri, 29 May 2026 15:56:10 UTC (603 KB)\n\n### References & Citations\n\nLoading...\n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer\n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers\n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps\n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations\n\n*(*[What are Smart Citations?](https://www.scite.ai/))# Code, Data and Media Associated with this Article\n\nalphaXiv\n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers\n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub\n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub\n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face\n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast\n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower\n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender\n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [ Learn more about arXivLabs](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/open-meditron-an-auditable-pipeline-for-clinical-llms", "canonical_source": "https://arxiv.org/abs/2605.16215", "published_at": "2026-07-21 10:19:02+00:00", "updated_at": "2026-07-21 10:53:35.596532+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "ai-ethics"], "entities": ["Fully Open Meditron", "Apertus-70B-MeditronFO", "Gemma-3-27B-MeditronFO", "MedGemma", "HealthBench", "Apertus-70B/8B-Instruct", "OLMo-2-32B-SFT", "EuroLLM-22B/9B-Instruct"], "alternates": {"html": "https://wpnews.pro/news/open-meditron-an-auditable-pipeline-for-clinical-llms", "markdown": "https://wpnews.pro/news/open-meditron-an-auditable-pipeline-for-clinical-llms.md", "text": "https://wpnews.pro/news/open-meditron-an-auditable-pipeline-for-clinical-llms.txt", "jsonld": "https://wpnews.pro/news/open-meditron-an-auditable-pipeline-for-clinical-llms.jsonld"}}