cd /news/natural-language-processing/analyzing-traditional-and-neural-app… · home topics natural-language-processing article
[ARTICLE · art-126500] src=arxiv.org ↗ pub= topic=natural-language-processing verified=true sentiment=· neutral

Analyzing Traditional and Neural Approaches to Multilingual Readability Assessment

A new arXiv paper (2609.10792v1) tests whether transformer-based models internalize the same linguistic features as traditional feature-based classifiers in Automatic Readability Assessment across Arabic, English, French, Hindi, and Russian using the ReadMe++ dataset. Using Shapley Additive Explanations (SHAP) to identify features driving traditional classifiers and TCAV concept sets to probe multilingual XLM-R and language-specific encoders, the researchers found transformers recover surface-length, syntactic, and lexical-diversity signals and reflect the ordinal CEFR structure of traditional models. Alignment varied by model family, language, and layer, with language-specific encoders tracking traditional models more clearly than XLM-R, and high linear separability did not always imply directional influence.

by read1 min views1 publishedSep 11, 2026

arXiv:2609.10792v1 Announce Type: new Abstract: Transformer-based models excel at Automatic Readability Assessment (ARA), yet feature-based models remain in active use because their predictions tie back to linguistic properties. This matters because readability labels are subjective and rater-dependent, so high accuracy on noisy ground truth may reflect surface patterns rather than the linguistic structure that defines difficulty. We test whether transformers internalize the same features as traditional models across Arabic, English, French, Hindi, and Russian using the ReadMe++ dataset. Shapley Additive Explanations (SHAP) identify the features driving traditional classifiers, which we then use as TCAV concept sets to probe multilingual XLM-R and language-specific encoders. Transformers recover surface-length, syntactic, and lexical-diversity signals, and reflect the ordinal CEFR structure of the traditional models. Alignment varies by model family, language, and layer, with language-specific encoders tracking traditional models more clearly than XLM-R. High linear separability does not always imply directional influence, limiting linear probing for count-based readability features.

── more in #natural-language-processing 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/analyzing-traditiona…] indexed:0 read:1min 2026-09-11 ·