cd /news/natural-language-processing/percept-a-corpus-for-pos-tagging-and… · home topics natural-language-processing article
[ARTICLE · art-92999] src=arxiv.org ↗ pub= topic=natural-language-processing verified=true sentiment=· neutral

PERCEPT: A Corpus for POS Tagging and Analysis of Persian-English Code-Mixing

Researchers introduced PERCEPT, the first publicly available large-scale Persian-English code-mixed corpus annotated with Universal Dependencies POS tags, comprising 6,800 posts from X, Instagram, and Digikala. The corpus, developed with an LLM-assisted annotation framework, reveals that nouns are the predominant category for code-mixed words and that the triggering effect is more pronounced in Digikala. The dataset is available at https://github.com/kalhorghazal/PERCEPT.

read1 min views1 publishedAug 12, 2026

arXiv:2608.10109v1 Announce Type: new Abstract: Social media has become a major venue for multilingual communication, where users frequently mix multiple languages within a single utterance. Although code-mixed corpora have been developed for several language pairs, Persian-English code-mixing remains relatively underexplored. Existing Persian resources lack Universal Dependencies (UD) part-of-speech (POS) annotations for code-mixed words, limiting both linguistic analyses and the development of syntax-aware NLP models. To address this gap, we introduce PERCEPT, the first publicly available large-scale Persian-English code-mixed corpus annotated with Universal Dependencies POS tags for code-mixed words. The dataset comprises 6,800 posts collected from X, Instagram, and Digikala. We further present an LLM-assisted annotation framework that automatically assigns POS tags and document-level topics. Human evaluation demonstrates high agreement between the automatically generated annotations and gold annotations, confirming the reliability of the annotations. Using PERCEPT, we conduct the first comprehensive linguistic analysis of Persian-English code-mixing across multiple social media platforms. Our analyses reveal that nouns are the predominant category for code-mixed words, while the distributions of other POS categories vary across platforms. We further find that the positional distribution of code-mixed words is remarkably consistent across platforms, whereas the triggering effect is substantially more pronounced in Digikala. PERCEPT is publicly available at https://github.com/kalhorghazal/PERCEPT.

── more in #natural-language-processing 4 stories · sorted by recency
── more on @percept 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/percept-a-corpus-for…] indexed:0 read:1min 2026-08-12 ·