{"slug": "open-sota-registry", "title": "Open SOTA Registry", "summary": "CodeSOTA, an open data terminal for reinforcement learning environments and state-of-the-art models, has launched a registry tracking 9,102 results, 163 models, 371 datasets, and 9 capability areas, with a representative result per capability area and an API schema. The registry is cited by researchers and analysts, including the University of Surrey and AAAI 2026, and features a research graph with 7 questions, 14 experiments, 18 claims, and 6 findings. The platform aims to help AI labs identify which benchmarks and environments still separate frontier models, with examples like ARC-AGI-1 spanning 31.4 points, ARC-Challenge 0.8, and GSM8K 0 in a matched snapshot.", "body_md": "[Featured guideWhich AI model fits your GPU?Benchmark-first picks for every card, RTX 3060 to MI300X.3060→Qwen3-8B4090→35B-A3B Q4H100→picked 70BMI300X→Kimi K2.6Read →](/hardware/best-model-by-gpu)\n\n# The data terminal for\n\n*RL environments & SOTA-with-code.*\n\nOne aggregated, dated, source-tiered registry of the evals, RL environments, models, and papers that move the frontier — the place AI labs check to see which environments actually separate models, and the team that *builds* the verifiable-reward environments that do.\n\n## Three ways in. *All backed by the same evidence.*\n\nWhether you are choosing what to cite, deciding what to train on, or tracking the frontier, you land on the same dated, source-tiered registry underneath.\n\n[Pick an eval01Browse benchmarks →](/benchmarks)\n\n### Browse benchmarks by what they prove.\n\nEvery eval with status, saturation, and lift evidence — so you cite numbers that still separate models, not dead ones.\n\n[Improve a capability02Open the recommender →](/rl-environments#capability)\n\n### Get the environments that lift it.\n\nPick a capability gap; get RL environments and datasets ranked by discriminative power — the ones most likely to move a strong model.\n\n[See what’s new03Open the feed →](/recent_papers)\n\n### The frontier feed.\n\nLatest evals, environments, models, and papers — chronological, dated, and linked back to the registry rows they touch.\n\n## One representative result *per capability area*.\n\nA compact snapshot of the nine capability areas. The homepage only prints a model when the registry row is verified and has an inspectable source URL; otherwise the row stays pending instead of promoting a stale or weak claim.\n\n- Results\n- 9,102\n- Models tracked\n- 163\n- Datasets indexed\n- 371\n- Capability areas\n- 9\n\n[API schema →](/api-landing/sota)\n\n| Capability | Evidence | Trusted model | Metric | Score | Source | Snapshot |\n|---|---|---|---|---|---|---|\n| Language & Knowledge |\n|\n\n[OCRBench v2](/benchmark/ocrbench-v2)[ovis2-5-9b](/model/ovis2-5-9b)[WildASR](/benchmark/wildasr)[VQA-v2](/benchmark/vqa-v2)[SWE-bench Verified](/benchmark/swe-bench-verified)[Claude Opus 4.7](/model/claude-opus-4-7)[GAIA](/benchmark/gaia)[MTEB](/benchmark/mteb)[Atari 2600](/benchmark/atari-2600)[MVTec-AD](/benchmark/mvtec-ad)`curl https://www.codesota.com/api/sota/swe-bench`\n\n[Docs](/api-landing/sota)\n\n[Cited & referenced byResearchers and analysts](/cited-by)\n\ncite the registry. Univ. of Surrey · AAAI 2026Tomasz Tunguz · Theory VenturesUseAIAPIAlternativeToHacker Newsr/MachineLearningSee all citations →\n\ncite the registry.\n\n## Results become useful when they answer *a question.*\n\nFollow public questions through hypotheses and experiments to dated evidence, durable claims, negative results, and reproduced findings.\n\nQUESTION → HYPOTHESIS → EXPERIMENT → EVIDENCE → CLAIM → FINDING\n\n- Questions\n- 7\n- Experiments\n- 14\n- Claims\n- 18\n- Findings\n- 6\n\n[Explore the research graph →](/research)\n\n[RQ-PG-0001partially answered](/research/questions/minimize-language-model-bpb-under-16mb-parameter-golf)\n\n### How do we minimize language-model BPB under a 16 MB artifact budget?\n\nThe pinned leaderboard moves from the 1.2244 baseline to 1.0565 BPB. The selected #1394–#1530 milestones improve from 1.08563 to 1.07336, but they are a frontier progression rather than one literal code lineage.\n\n[RQ-0006partially answered](/research/questions/reasoning-benchmarks-that-still-separate-frontier-models)\n\n### Which reasoning benchmarks still separate frontier models?\n\nIn this matched snapshot ARC-AGI-1 spans 31.4 points, ARC-Challenge spans 0.8, and GSM8K spans 0. Protocol-matched reruns are still required before calling this an intrinsic benchmark property.\n\n[RQ-0005open](/research/questions/benchmark-role-changes-with-model-scale)\n\n### When should a benchmark move from leading to primary to control?\n\nPythia reference results suggest overlapping dynamic ranges, but the promotion thresholds need repeated measurements and confidence intervals.\n\n## Practical routes. *Benchmarks as evidence.*\n\nThe top-level map is a navigation layer, not a perfect ontology. Capabilities, modalities, and vertical domains stay linked through task pages, benchmark sets, datasets, models, papers, and evidence rows.\n\n[RouteLanguage & KnowledgeMMLU-Pro · GPQA · MTEB](/nlp)\n\nReasoning, exams, retrieval, and knowledge-heavy language tasks.\n\n[Route63.40Vision & DocumentsCOCO · OCRBench · OmniDocBench](/ocr)\n\nImages, detection, OCR, layout, tables, and document parsing.\n\n[RouteAudio & SpeechWildASR · VoiceBench · ESC-50](/audio)\n\nASR, audio tagging, voice assistants, speech quality, and TTS.\n\n[RouteMultimodal MediaVQA-v2 · TextVQA · MMMU](/browse/multimodal)\n\nVQA, charts, video, image-text reasoning, and media understanding.\n\n[Route87.6%Code & Software EngineeringHumanEval · LiveCodeBench · SWE-bench](/code-generation)\n\nCode generation, repair, repository tasks, and verified software work.\n\n[Route87.6%Agents & Tool UseGAIA · WebArena · OSWorld](/agentic)\n\nLong-horizon tool use, browser work, OS tasks, and workflow execution.\n\n[RouteStructured Data & ForecastingMTEB · tabular · graph suites](/tasks)\n\nEmbeddings, retrieval, reranking, tabular prediction, graphs, and forecasting.\n\n[RouteRobotics, Control & RLAtari · Habitat · LIBERO](/robotics)\n\nSimulation, control, games, embodied agents, and manipulation.\n\n[RouteScience, Medicine & IndustryCheXpert · MVTec-AD · MedQA](/medical)\n\nScientific QA, medical imaging, industrial inspection, and applied AI.\n\n## A leaderboard row is not a fact *until it can be inspected.*\n\nCodeSOTA is useful only if the evidence is visible. The homepage now surfaces the provenance contract before editorial lineages and release notes.\n\n### Dated scores\n\nRows carry access dates and snapshot context so old frontier claims do not masquerade as current facts.\n\n### Metric direction\n\nEvery benchmark declares whether higher or lower is better before a winner is selected.\n\n### Source tiers\n\nPaper, vendor, reproduced, and registry-maintained rows are labeled separately.\n\n### Provenance trail\n\nBenchmark pages connect model, paper, dataset, code, and source URL where available.\n\n## The neutral measure for the environments labs train on.\n\nA market of RL-environment startups is selling to frontier labs — and the labs keep asking the same question: does this environment actually separate models, or is it saturated? CodeSOTA is the independent party that answers it. If you build environments, we certify yours discriminates. If you train models, we tell you which ones are worth the run.\n\n## The registry is *callable.*\n\nAgents and notebooks should not scrape leaderboards. They should call a stable, source-aware endpoint and cache the snapshot they used.\n\n[API docs](/api-landing/sota)\n\n**codesota/sota-api**\n\n```\ncurl https://www.codesota.com/api/sota/swe-bench\ncurl https://www.codesota.com/api/sota?area=vision-documents\n{\n  \"task\": \"swe-bench\",\n  \"metric\": \"resolve rate\",\n  \"direction\": \"higher\",\n  \"leader\": {\n    \"model\": \"registry top pick\",\n    \"score\": \"dated value\",\n    \"source\": \"paper | vendor | reproduced\",\n    \"snapshot_id\": \"2026-04-27\"\n  }\n}\n```\n\n## Tell CodeSOTA what to watch.\n\n*Use your words, not our taxonomy.*\n\nYou do not need to know whether the thing is called SWE-bench Verified, agentic coding, open-weight models, or Polish LLMs. Write the alert you actually want. We map it to benchmark rules, show the interpretation back, and fall back to a weekly digest when the request is too broad.\n\nBroad request detected. Defaulting to a weekly digest.\n\n- Weekly digest of top CodeSOTA changes\n\n## Trained something\n\n*that beats the table?*\n\nSubmit a checkpoint, paper result, or correction with structured benchmark provenance. We validate the score, cross-check the source, and add the row to the registry with its date and evidence trail.", "url": "https://wpnews.pro/news/open-sota-registry", "canonical_source": "https://www.codesota.com", "published_at": "2026-08-25 21:12:12+00:00", "updated_at": "2026-08-25 21:46:25.347887+00:00", "lang": "en", "topics": ["ai-research", "ai-tools", "ai-infrastructure"], "entities": ["CodeSOTA", "University of Surrey", "AAAI 2026", "Tomasz Tunguz", "Theory Ventures", "ARC-AGI-1", "ARC-Challenge", "GSM8K"], "alternates": {"html": "https://wpnews.pro/news/open-sota-registry", "markdown": "https://wpnews.pro/news/open-sota-registry.md", "text": "https://wpnews.pro/news/open-sota-registry.txt", "jsonld": "https://wpnews.pro/news/open-sota-registry.jsonld"}}