cd /news/natural-language-processing/quanling-cross-branch-validation-of-… · home › topics › natural-language-processing › article
[ARTICLE · art-147319] src=arxiv.org ↗ pub= topic=natural-language-processing verified=true sentiment=· neutral

QuanLing: Cross-Branch Validation of Language Distance Quantification on Western Romance

A new arXiv paper (2610.08851v1) extends the QuanLing framework for quantifying language distance to Western Romance languages, finding Portuguese and Spanish closest (LaBSE distance 0.0229) and French and Italian most distant (0.0338) across 150 four-language parallel sentences. The study, which combines LaBSE sentence embedding distances, tokenization fragmentation rates from four monolingual BERT tokenizers, and mBERT masked language model mutual intelligibility, reports that LaBSE and mBERT rankings agree on 4 of 6 language pairs. Western Romance shows a wider absolute distance span than North Germanic (0.011 vs. 0.008) with comparable relative ratios (1.48 vs. 1.67), and French shows higher MLM predictability (36.12% top-1 accuracy vs. 29.28% for Italian).

by read1 min views3 publishedOct 8, 2026

arXiv:2610.08851v1 Announce Type: new Abstract: Quantifying language distance among closely related languages remains a core challenge in quantitative linguistics. Our previous work [1] introduced QuanLing (Quantitative Linguistics via Pretrained Language Models), a quantitative framework combining language distance metrics (sentence embedding distance, tokenization fragmentation rate) with language property analysis (MLM prediction probability), validated on North Germanic (Danish, Norwegian Bokm{\aa}l, Swedish). This paper extends QuanLing to Western Romance--French, Portuguese, Spanish, Italian--testing cross-branch applicability with the same metric family and aggregation protocol as our North Germanic study, adapted for four languages (English anchor, quadruplet construction). Using 150 four-language parallel sentences, we compute LaBSE sentence embedding distances, tokenization fragmentation rates from four monolingual BERT tokenizers, and mBERT masked language model mutual intelligibility. Results show that Portuguese--Spanish are closest (LaBSE distance 0.0229), French--Italian most distant (0.0338); LaBSE and mBERT rankings agree on 4 of 6 pairs, confirming cross-model robustness. Western Romance shows a wider absolute distance span than North Germanic (0.011 vs. 0.008) but comparable relative ratios (1.48 vs. 1.67), consistent with longer divergence time. French exhibits notably higher MLM predictability (36.12% top-1 accuracy vs. 29.28% for Italian), reflecting its orthography--phonology decoupling. This cross-branch validation provides further evidence for QuanLing's generalizability beyond a single language branch.

── more in #natural-language-processing 4 stories · sorted by recency
── more on @quanling 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/quanling-cross-branc…] indexed:0 read:1min 2026-10-08 · —