cd /news/artificial-intelligence/google-deepmind-and-harvard-propose-… · home topics artificial-intelligence article
[ARTICLE · art-127120] src=cryptobriefing.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Google DeepMind and Harvard propose vision-first path to AGI

A white paper titled "Visual General Intelligence: A White Paper," published on arXiv as 2608.25924 by more than 21 researchers from Google DeepMind, Harvard, and other institutions, argues that images, video, and geometric data — not text alone — should drive the path to artificial general intelligence. Contributors include Robert Geirhos of Google DeepMind and Yilun Du of Harvard, and the work grew out of discussions at the CVPR 2026 Visual General Intelligence Workshop. The paper presents no model, benchmark results, or product roadmap, instead proposing principles, benchmarks, and learning paradigms such as generative video models and self-supervised learning, and it builds on DeepMind's 2024 "Levels of AGI" and June 2026 "From AGI to ASI" papers.

by read2 min views4 publishedSep 11, 2026
Google DeepMind and Harvard propose vision-first path to AGI
Image: Cryptobriefing (auto-discovered)

Google 2015 logo (Wikimedia Commons, public domain)

A new white paper from over 21 AI researchers argues that images and video, not just text, could be the key to building artificial general intelligence

For years, the race to build artificial general intelligence has been dominated by one modality: language. GPT, Claude, Gemini. All of them learned to reason by digesting oceans of text. A new white paper from researchers at Google DeepMind, Harvard, and other leading institutions argues that approach might be incomplete, and that the path to AGI could run through what machines see rather than what they read. The paper, titled “Visual General Intelligence: A White Paper” and published on arXiv as 2608.25924, lays out a research agenda for what the authors call visual general intelligence, or VGI. The core thesis: AI systems should learn directly from images, videos, and geometric data to understand, predict, and act in the physical world.

What the paper actually says #

This is not a product announcement or a benchmark-beating model reveal. It is a position paper, a collective argument from more than 21 researchers about where the field should invest its attention next.

Among the contributors are Robert Geirhos from Google DeepMind and Yilun Du from Harvard. The work grew out of discussions at the CVPR 2026 Visual General Intelligence Workshop, one of the premier gatherings in the computer vision community.

Rather than presenting a single model or architecture, the paper discusses principles, benchmarks, and learning paradigms for building intelligence through visual experience. The researchers argue that generative video models and self-supervised learning, where systems train themselves by predicting what comes next in visual sequences, could provide a foundation for AGI that language alone cannot.

The paper also explores strategies for integrating multiple modalities. Vision and language aren’t positioned as competitors in this framework but as complementary channels, each capturing different slices of intelligence.

Building on DeepMind’s AGI framework #

The white paper doesn’t exist in isolation. It builds on a lineage of publications from Google DeepMind that have attempted to map the road to AGI and beyond.

In 2024, DeepMind published “Levels of AGI,” a paper that proposed a taxonomy for measuring progress toward general intelligence. More recently, in June 2026, DeepMind followed up with “From AGI to ASI,” which explored what might come after general intelligence: artificial superintelligence. The VGI white paper slots neatly into this intellectual arc, asking whether the current text-heavy paradigm is sufficient to reach the milestones those earlier papers described.

For now, no specific performance claims, model timelines, or product roadmaps accompany the paper. It is deliberately open-ended, designed to spark exploration rather than declare victory. Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our

Editorial Policy.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @google deepmind 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/google-deepmind-and-…] indexed:0 read:2min 2026-09-11 ·