cd /news/large-language-models/orqa-an-occupation-realistic-questio… · home topics large-language-models article
[ARTICLE · art-128730] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

ORQA: An Occupation-Realistic Question and Answer Framework for LLM Professional Knowledge

A new benchmark called ORQA, built by connecting O*NET occupations to trusted occupation-specific websites, found that the best large language models score only about 58-62% on occupation-level professional knowledge questions, according to the arXiv paper arXiv:2609.12366v1. The question set covers 116 occupations across all 21 SOC major groups, with 480 questions sourced from 187 websites, and was used to test 15 frontier and open-weight models. Claude Opus 4.6, GPT-5.4 and Claude Sonnet 4.6 performed best at roughly 58-62%, while smaller open-weight models reached about 33-41%, with healthcare occupations highest at 78% and individual occupations such as Sheet Metal Workers and Fish and Game Wardens near zero.

by read1 min views1 publishedSep 14, 2026

arXiv:2609.12366v1 Announce Type: new Abstract: We present ORQA, a method for testing occupation-level knowledge in large language models. Prior methods either map abstract LLM skills to occupations via task definitions or utilize expert knowledge which is difficult to obtain at scale and expensive. ORQA complements both of these methods by connecting O*NET occupations to trusted occupation-specific websites (such as regulatory agencies, licensing bodies, professional organizations, and government publications) and converting these into source-traceable question-answer pairs. A combination of an automated pipeline and human review produces a set of high quality questions about occupations. The question set created via our method covers 116 occupations from all 21 major groups in the SOC, with 480 questions sourced from 187 different websites. Each question is designed to probe a real-world skill question that is relevant to the occupation in question. We test 15 state-of-the-art frontier and open-weight models via this method. Claude Opus 4.6, GPT-5.4 and Claude Sonnet 4.6 all perform the best at approximately 58-62% while smaller open-weight models achieve approximately 33-41% performance. Performance varies significantly across occupations. Healthcare-related occupations achieve the highest performance (78%) while Office and Administrative Support achieve approximately 40%. Performance on individual occupations (e.g. Sheet Metal Workers and Fish and Game Wardens) is essentially zero. We also find that open-ended questions and weighting by wage bill do not significantly affect the ranking of models on this benchmark. We believe that leveraging existing trusted occupation-specific information to test LLM knowledge in professional domains may be a scalable and useful method for evaluating occupation-level AI performance in the future. Results and data are available at orqabench.org.

── more in #large-language-models 4 stories · sorted by recency
── more on @orqa 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/orqa-an-occupation-r…] indexed:0 read:1min 2026-09-14 ·