cd /news/large-language-models/treeprobe-a-tibetan-medicine-benchma… · home topics large-language-models article
[ARTICLE · art-85620] src=machinebrief.com ↗ pub= topic=large-language-models verified=true sentiment=· neutral

TreeProbe : A Tibetan Medicine Benchmark for Cultural Bias in LLMs

Researchers introduced TreeProbe, the first cultural-bias benchmark for Tibetan medicine, containing 4,719 expert-adjudicated items covering 467 diseases and 10 subtasks, and found that current large language models exhibit systematic external ontology drift in native Tibetan medical contexts. The benchmark, organized around the native Tree of Medicine framework, reveals that models diverge in whether they drift toward biomedical or traditional Chinese medicine reasoning, highlighting the need for linguistically inclusive and epistemically fair medical AI systems.

read1 min views1 publishedAug 4, 2026

arXiv:2608.00640v1 Announce Type: new Abstract: Large language models are increasingly viewed as a potential means of mitigating global health inequities, yet their outputs often reflect dominant high-resource medical traditions and provide limited coverage of traditional medical knowledge systems. Tibetan medicine, one of the world's four major traditional medical systems, has an independent and highly structured theoretical framework. When models lack grounded understanding of Tibetan medicine, they may fall back on dominant epistemic systems and distort the native knowledge structure during reasoning. However, quantitative tools for evaluating cultural bias in Tibetan medicine remain largely absent. To address this gap, we introduce TreeProbe, the first cultural-bias benchmark organized around the native Tree of Medicine framework in Tibetan medicine. It contains 4,719 expert-adjudicated items covering 467 diseases and 10 subtasks along the three roots. Experiments on representative LLMs show that current models remain limited in native Tibetan medical contexts and exhibit systematic external ontology drift. Further analysis reveals that models diverge in whether they drift toward biomedical or TCM reasoning, shaped by pretraining data composition and surface resemblance between TCM and Tibetan medicine. TreeProbe provides a diagnostic benchmark for developing medical AI systems that are both linguistically inclusive and epistemically fair. Code and data are available in an anonymous repository at https://anonymous.4open.science/r/TreeProbe/.

── more in #large-language-models 4 stories · sorted by recency
── more on @treeprobe 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/treeprobe-a-tibetan-…] indexed:0 read:1min 2026-08-04 ·