{"slug": "analyzing-traditional-and-neural-approaches-to-multilingual-readability", "title": "Analyzing Traditional and Neural Approaches to Multilingual Readability Assessment", "summary": "A new arXiv paper (2609.10792v1) tests whether transformer-based models internalize the same linguistic features as traditional feature-based classifiers in Automatic Readability Assessment across Arabic, English, French, Hindi, and Russian using the ReadMe++ dataset. Using Shapley Additive Explanations (SHAP) to identify features driving traditional classifiers and TCAV concept sets to probe multilingual XLM-R and language-specific encoders, the researchers found transformers recover surface-length, syntactic, and lexical-diversity signals and reflect the ordinal CEFR structure of traditional models. Alignment varied by model family, language, and layer, with language-specific encoders tracking traditional models more clearly than XLM-R, and high linear separability did not always imply directional influence.", "body_md": "arXiv:2609.10792v1 Announce Type: new \nAbstract: Transformer-based models excel at Automatic Readability Assessment (ARA), yet feature-based models remain in active use because their predictions tie back to linguistic properties. This matters because readability labels are subjective and rater-dependent, so high accuracy on noisy ground truth may reflect surface patterns rather than the linguistic structure that defines difficulty. We test whether transformers internalize the same features as traditional models across Arabic, English, French, Hindi, and Russian using the ReadMe++ dataset. Shapley Additive Explanations (SHAP) identify the features driving traditional classifiers, which we then use as TCAV concept sets to probe multilingual XLM-R and language-specific encoders. Transformers recover surface-length, syntactic, and lexical-diversity signals, and reflect the ordinal CEFR structure of the traditional models. Alignment varies by model family, language, and layer, with language-specific encoders tracking traditional models more clearly than XLM-R. High linear separability does not always imply directional influence, limiting linear probing for count-based readability features.", "url": "https://wpnews.pro/news/analyzing-traditional-and-neural-approaches-to-multilingual-readability", "canonical_source": "https://arxiv.org/abs/2609.10792", "published_at": "2026-09-11 04:00:00+00:00", "updated_at": "2026-09-11 04:25:58.965835+00:00", "lang": "en", "topics": ["natural-language-processing", "machine-learning", "ai-research", "large-language-models"], "entities": ["arXiv", "ReadMe++", "XLM-R", "SHAP", "TCAV", "CEFR"], "alternates": {"html": "https://wpnews.pro/news/analyzing-traditional-and-neural-approaches-to-multilingual-readability", "markdown": "https://wpnews.pro/news/analyzing-traditional-and-neural-approaches-to-multilingual-readability.md", "text": "https://wpnews.pro/news/analyzing-traditional-and-neural-approaches-to-multilingual-readability.txt", "jsonld": "https://wpnews.pro/news/analyzing-traditional-and-neural-approaches-to-multilingual-readability.jsonld"}}