cd /news/natural-language-processing/byte-pair-encoding-captures-aspects-… · home › topics › natural-language-processing › article
[ARTICLE · art-147245] src=aclanthology.org ↗ pub= topic=natural-language-processing verified=true sentiment=· neutral

Byte-pair Encoding Captures Aspects of Word Meaning that Distributional Semantics Overlooks

Byte-pair encoding subword tokens capture aspects of word meaning that distributional semantics overlooks, according to a study by David A. Haslett, Antoni B. Chan, and Janet H. Hsiao published in Transactions of the Association for Computational Linguistics, Volume 14, pages 2425–2448. Across 18 languages, shared tokens predicted greater semantic priming effects in humans — for example, people recognize bromine faster after reading fluorine — beyond what distributional similarity explained, and across seven languages shared tokens boosted semantic priming beyond what shared morphemes explained. The authors also found that the trend toward larger token vocabularies obscures category markers, and that language models underestimate the semantic relatedness of words that share tokens even when they have access to subword tokens.

read1 min views1 publishedOct 7, 2026
Byte-pair Encoding Captures Aspects of Word Meaning that Distributional Semantics Overlooks
Image: Aclanthology (auto-discovered)
Abstract

Large language models learn word meanings through distributional semantics (i.e., patterns in text), and distribution does not always capture information about taxonomy and function, which are essential to human representations of word meaning. However, sound and spelling often conveys such information, and language models decompose many words into subword tokens, which may identify category markers. For example, GPT-5 segments fluorine and bromine into fluor + ine and brom + ine, which share the category marker ine. We provide evidence that, across 18 languages, shared tokens predict greater semantic priming effects in humans (e.g., people recognize bromine faster after reading fluorine), beyond what distributional similarity explains. Furthermore, across seven languages, shared tokens boost semantic priming beyond what shared morphemes explain. This suggests that subword tokens could help language models arrive at human-like representations of word meanings. However, we also provide evidence that the trend towards larger vocabularies of tokens obscures category markers, and even when language models have access to subword tokens, they underestimate the semantic relatedness of words that share tokens.

- Anthology ID:
- 2026.tacl-1.113
- Volume:
- [Transactions of the Association for Computational Linguistics, Volume 14](https://aclanthology.org/volumes/2026.tacl-1/)
- Month:
- Year:
  • 2026
  • Address:
  • Cambridge, MA
- Venue:
- [TACL](https://aclanthology.org/venues/tacl/)
- SIG:
- Publisher:
  • MIT Press
- Note:
- Pages:
  • 2425–2448
- Language:
- URL:
- [https://aclanthology.org/2026.tacl-1.113/](https://aclanthology.org/2026.tacl-1.113/)
- DOI:
- [10.1162/tacl.a.810](https://doi.org/10.1162/tacl.a.810)
- Cite (ACL):
- Cite (Informal):
- [Byte-pair Encoding Captures Aspects of Word Meaning that Distributional Semantics Overlooks](https://aclanthology.org/2026.tacl-1.113/) (Haslett et al., TACL 2026)
- PDF:
- [https://aclanthology.org/2026.tacl-1.113.pdf](https://aclanthology.org/2026.tacl-1.113.pdf)
── more in #natural-language-processing 4 stories · sorted by recency
── more on @david a. haslett 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/byte-pair-encoding-c…] indexed:0 read:1min 2026-10-07 · —