Open Meditron: An Auditable Pipeline for Clinical LLMs Researchers introduced Fully Open Meditron, the first fully open pipeline for building clinical large language models (LLMs), comprising a clinician-audited training corpus, reproducible framework, and evaluation protocol. The Apertus-70B-MeditronFO variant improved +6.6 points over its base (47.2% to 53.8%) on aggregate medical benchmarks, establishing a new fully open state-of-the-art. The pipeline enforces system-wide decontamination and end-to-end validation by a four-physician panel, demonstrating that fully open models can achieve top performance without sacrificing auditability. Computer Science Artificial Intelligence Submitted on 15 May 2026 v1 https://arxiv.org/abs/2605.16215v1 , last revised 29 May 2026 this version, v2 Title:Fully Open Meditron: An Auditable Pipeline for Clinical LLMs View PDF /pdf/2605.16215 HTML experimental https://arxiv.org/html/2605.16215v2 Abstract:Clinical decision support systems CDSS require scrutable, auditable pipelines that enable rigorous, reproducible validation. Yet current LLM-based CDSS remain largely opaque. Most "open" models are open-weight only, releasing parameters while withholding the data provenance, curation procedures, and generation pipelines that determine model behavior. Fully Open FO models, which expose the complete training stack end-to-end, do not currently exist in medicine. We introduce Fully Open Meditron, the first fully open pipeline for building LLM-CDSS, comprising a clinician-audited training corpus, a reproducible data construction and training framework, and a use-aligned evaluation protocol. The corpus unifies eight public medical QA datasets into a normalized conversational format and expands coverage with three clinician-vetted synthetic extensions: exam-style QA, guideline-grounded QA derived from 46,469 clinical practice guidelines, and clinical vignettes. The pipeline enforces system-wide decontamination, gold-label resampling of teacher generations, and end-to-end validation by a four-physician panel. We evaluate using an LLM-as-a-judge protocol over expert-written clinical vignettes, calibrated against 204 human raters. We apply the recipe to five FO base models Apertus-70B/8B-Instruct, OLMo-2-32B-SFT, EuroLLM-22B/9B-Instruct . All MeditronFO variants are preferred over their bases. Apertus-70B-MeditronFO improves +6.6 points over its base 47.2% to 53.8% on aggregate medical benchmarks, establishing a new FO SoTA. Gemma-3-27B-MeditronFO is preferred over MedGemma in 58.6% of LLM-as-a-judge comparisons and outperforms it on HealthBench 58% vs 55.9% . These results show that fully open pipelines can achieve state-of-the-art domain-specific performance without sacrificing auditability or reproducibility. Submission history From: Xavier Theimer-Lienhard view email /show-email/92d2e130/2605.16215 Fri, 15 May 2026 17:29:08 UTC 603 KB v1 /abs/2605.16215v1 v2 Fri, 29 May 2026 15:56:10 UTC 603 KB References & Citations Loading... Bibliographic and Citation Tools Bibliographic Explorer What is the Explorer? https://info.arxiv.org/labs/showcase.html arxiv-bibliographic-explorer Connected Papers What is Connected Papers? https://www.connectedpapers.com/about Litmaps What is Litmaps? https://www.litmaps.co/ scite Smart Citations What are Smart Citations? https://www.scite.ai/ Code, Data and Media Associated with this Article alphaXiv What is alphaXiv? https://alphaxiv.org/ CatalyzeX Code Finder for Papers What is CatalyzeX? https://www.catalyzex.com DagsHub What is DagsHub? https://dagshub.com/ Gotit.pub What is GotitPub? http://gotit.pub/faq Hugging Face What is Huggingface? https://huggingface.co/huggingface ScienceCast What is ScienceCast? https://sciencecast.org/welcome Demos Recommenders and Search Tools Influence Flower What are Influence Flowers? https://influencemap.cmlab.dev/ CORE Recommender What is CORE? https://core.ac.uk/services/recommender arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs https://info.arxiv.org/labs/index.html .