cd /news/large-language-models/sk-bench-a-native-first-benchmark-fo… · home › topics › large-language-models › article
[ARTICLE · art-147325] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

sk-bench: A Native-First Benchmark for Evaluating Large Language Models in Slovak

A new arXiv paper (2610.09152v1) introduces sk-bench, a native-first Slovak benchmark comprising 30 datasets and 33 scored task variants across ten skill categories, and reports that the best open-weights model trails proprietary APIs by 12.6 points across 55 evaluated models. The authors find that continued Slovak pretraining lowers Qwen3-14B's overall score by 13.9 points, with a small instruction set restoring three quarters of that loss, while test-time reasoning improves scores by 8.5 to 12.5 points for models of 9B and above. The data and code are released at https://github.com/slovak-nlp/sk-bench.

by read1 min views2 publishedOct 8, 2026

arXiv:2610.09152v1 Announce Type: new Abstract: Multilingual LLM benchmarks omit Slovak, a morphologically rich West Slavic language of five million speakers, or cover it only by machine translation. We present sk-bench, a native-first Slovak benchmark with 30 datasets (33 scored task variants) across ten skill categories. Eleven resources are introduced or first packaged for generative-LLM evaluation, including IFEval-SK with Slovak-adapted instruction checkers and native Chiby/SKJ1 resources for Slovak grammar and morphology. We evaluate 55 open- and closed-weights models under one harness. The best open model trails proprietary APIs by 12.6 points. Model rankings are similar for native and translated closed-form data ($\rho\geq0.98$), though translation separates the strongest models less well. By contrast, human-authored and LLM-generated QA questions rank models differently ($\rho=0.72$). For Qwen3-14B, continued Slovak pretraining lowers the overall score by 13.9 points. A small instruction set restores three quarters of that loss. Test-time reasoning improves scores by 8.5 to 12.5 points for models of 9B and above. Together, these findings suggest four design lessons for other under-resourced languages: use native data where translation fails, plan instruction repair after language adaptation, enable test-time reasoning before scaling up, and avoid overinvesting in target-language prompts. We release the data and code at https://github.com/slovak-nlp/sk-bench

── more in #large-language-models 4 stories · sorted by recency
── more on @sk-bench 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/sk-bench-a-native-fi…] indexed:0 read:1min 2026-10-08 · —