{"slug": "sk-bench-a-native-first-benchmark-for-evaluating-large-language-models-in-slovak", "title": "sk-bench: A Native-First Benchmark for Evaluating Large Language Models in Slovak", "summary": "A new arXiv paper (2610.09152v1) introduces sk-bench, a native-first Slovak benchmark comprising 30 datasets and 33 scored task variants across ten skill categories, and reports that the best open-weights model trails proprietary APIs by 12.6 points across 55 evaluated models. The authors find that continued Slovak pretraining lowers Qwen3-14B's overall score by 13.9 points, with a small instruction set restoring three quarters of that loss, while test-time reasoning improves scores by 8.5 to 12.5 points for models of 9B and above. The data and code are released at https://github.com/slovak-nlp/sk-bench.", "body_md": "arXiv:2610.09152v1 Announce Type: new \nAbstract: Multilingual LLM benchmarks omit Slovak, a morphologically rich West Slavic language of five million speakers, or cover it only by machine translation. We present sk-bench, a native-first Slovak benchmark with 30 datasets (33 scored task variants) across ten skill categories. Eleven resources are introduced or first packaged for generative-LLM evaluation, including IFEval-SK with Slovak-adapted instruction checkers and native Chiby/SKJ1 resources for Slovak grammar and morphology. We evaluate 55 open- and closed-weights models under one harness. The best open model trails proprietary APIs by 12.6 points. Model rankings are similar for native and translated closed-form data ($\\rho\\geq0.98$), though translation separates the strongest models less well. By contrast, human-authored and LLM-generated QA questions rank models differently ($\\rho=0.72$). For Qwen3-14B, continued Slovak pretraining lowers the overall score by 13.9 points. A small instruction set restores three quarters of that loss. Test-time reasoning improves scores by 8.5 to 12.5 points for models of 9B and above. Together, these findings suggest four design lessons for other under-resourced languages: use native data where translation fails, plan instruction repair after language adaptation, enable test-time reasoning before scaling up, and avoid overinvesting in target-language prompts. We release the data and code at https://github.com/slovak-nlp/sk-bench", "url": "https://wpnews.pro/news/sk-bench-a-native-first-benchmark-for-evaluating-large-language-models-in-slovak", "canonical_source": "https://arxiv.org/abs/2610.09152", "published_at": "2026-10-08 04:00:00+00:00", "updated_at": "2026-10-08 04:18:32.622016+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "natural-language-processing", "machine-learning", "artificial-intelligence"], "entities": ["sk-bench", "Qwen3-14B", "arXiv", "IFEval-SK", "Chiby", "SKJ1", "slovak-nlp"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/sk-bench-a-native-first-benchmark-for-evaluating-large-language-models-in-slovak", "markdown": "https://wpnews.pro/news/sk-bench-a-native-first-benchmark-for-evaluating-large-language-models-in-slovak.md", "text": "https://wpnews.pro/news/sk-bench-a-native-first-benchmark-for-evaluating-large-language-models-in-slovak.txt", "jsonld": "https://wpnews.pro/news/sk-bench-a-native-first-benchmark-for-evaluating-large-language-models-in-slovak.jsonld"}}