cd /news/artificial-intelligence/asi-bench-at-the-dawn-of-artificial-… · home topics artificial-intelligence article
[ARTICLE · art-102430] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

ASI-Bench: At the Dawn of Artificial Superintelligence

Researchers introduced ASI-Bench, the first benchmark to jointly evaluate AI systems' innovative exploration and autonomous scientific execution across general research domains, built by over 40 experts at a cost of 31,000+ human hours and containing 60 project-level research tasks across 11 scientific domains. Across 18 state-of-the-art agent-model configurations, average scores dropped from 50.91 with full methodological guidance to 29.10 with only the method specified and 26.62 when agents determined methods themselves, showing current systems remain heavily dependent on human guidance. The benchmark is open for contributions at https://asibench.apexin.ai/submit.

read1 min views1 publishedAug 19, 2026

arXiv:2608.17271v1 Announce Type: new Abstract: Artificial superintelligence (ASI) requires AI to move beyond mastering existing knowledge toward exploring the unknown, creating new knowledge, and turning new ideas into verifiable results. However, the capabilities of today's AI systems are still largely built on learning, compressing, and applying existing human knowledge. Accordingly, existing benchmarks primarily test whether AI can produce correct answers based on learned knowledge, or whether it can complete tasks under extensive human guidance. We therefore introduce ASI-Bench, the first benchmark to jointly evaluate AI systems' capabilities of innovative exploration and autonomous scientific execution across general research domains, and the first to progressively withdraw human methodological guidance within the same research project to test how far AI can proceed on its own. Built by over 40 experts with the cost of 31,000+ human hours, ASI-Bench contains 60 project-level research tasks across 11 scientific domains and progressively reduces methodological guidance to test whether AI can independently select methods, conduct research, and produce verifiable results. All tasks undergo expert review, AI-assisted auditing, sandbox execution, and scorer validation. Across 18 state-of-the-art agent--model configurations, the average score drops from 50.91 with full methodological guidance to 29.10 with only the method specified and 26.62 when agents must determine the method themselves. This sharp decline shows that current systems remain heavily dependent on human guidance and are still far from autonomously conducting end-to-end, project-level scientific research. ASI-Bench is open to the world. We invite researchers and builders everywhere to contribute new tasks, challenge the limits of today's AI, and help accelerate humanity's collective path toward artificial superintelligence at https://asibench.apexin.ai/submit.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @asi-bench 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/asi-bench-at-the-daw…] indexed:0 read:1min 2026-08-19 ·