arXiv:2610.09152v1 Announce Type: new Abstract: Multilingual LLM benchmarks omit Slovak, a morphologically rich West Slavic language of five million speakers, or cover it only by machine translation. We present sk-bench, a native-first Slovak benchmark with 30 datasets (33 scored task variants) across ten skill categories. Eleven resources are introduced or first packaged for generative-LLM evaluation, including IFEval-SK with Slovak-adapted instruction checkers and native Chiby/SKJ1 resources for Slovak grammar and morphology. We evaluate 55 open- and closed-weights models under one harness. The best open model trails proprietary APIs by 12.6 points. Model rankings are similar for native and translated closed-form data ($\rho\geq0.98$), though translation separates the strongest models less well. By contrast, human-authored and LLM-generated QA questions rank models differently ($\rho=0.72$). For Qwen3-14B, continued Slovak pretraining lowers the overall score by 13.9 points. A small instruction set restores three quarters of that loss. Test-time reasoning improves scores by 8.5 to 12.5 points for models of 9B and above. Together, these findings suggest four design lessons for other under-resourced languages: use native data where translation fails, plan instruction repair after language adaptation, enable test-time reasoning before scaling up, and avoid overinvesting in target-language prompts. We release the data and code at https://github.com/slovak-nlp/sk-bench
sk-bench: A Native-First Benchmark for Evaluating Large Language Models in Slovak
A new arXiv paper (2610.09152v1) introduces sk-bench, a native-first Slovak benchmark comprising 30 datasets and 33 scored task variants across ten skill categories, and reports that the best open-weights model trails proprietary APIs by 12.6 points across 55 evaluated models. The authors find that continued Slovak pretraining lowers Qwen3-14B's overall score by 13.9 points, with a small instruction set restoring three quarters of that loss, while test-time reasoning improves scores by 8.5 to 12.5 points for models of 9B and above. The data and code are released at https://github.com/slovak-nlp/sk-bench.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.