{"slug": "google-deepmind-and-harvard-propose-vision-first-path-to-agi", "title": "Google DeepMind and Harvard propose vision-first path to AGI", "summary": "A white paper titled \"Visual General Intelligence: A White Paper,\" published on arXiv as 2608.25924 by more than 21 researchers from Google DeepMind, Harvard, and other institutions, argues that images, video, and geometric data — not text alone — should drive the path to artificial general intelligence. Contributors include Robert Geirhos of Google DeepMind and Yilun Du of Harvard, and the work grew out of discussions at the CVPR 2026 Visual General Intelligence Workshop. The paper presents no model, benchmark results, or product roadmap, instead proposing principles, benchmarks, and learning paradigms such as generative video models and self-supervised learning, and it builds on DeepMind's 2024 \"Levels of AGI\" and June 2026 \"From AGI to ASI\" papers.", "body_md": "Google 2015 logo (Wikimedia Commons, public domain)\n\n# Google DeepMind and Harvard propose vision-first path to AGI\n\nA new white paper from over 21 AI researchers argues that images and video, not just text, could be the key to building artificial general intelligence\n\nFor years, the race to build artificial general intelligence has been dominated by one modality: language. GPT, Claude, Gemini. All of them learned to reason by digesting oceans of text. A new white paper from researchers at [Google](https://cryptobriefing.com/markets/alphabet/) DeepMind, Harvard, and other leading institutions argues that approach might be incomplete, and that the path to AGI could run through what machines see rather than what they read.\n\nThe paper, titled “Visual General Intelligence: A White Paper” and published on arXiv as 2608.25924, lays out a research agenda for what the authors call visual general intelligence, or VGI. The core thesis: AI systems should learn directly from images, videos, and geometric data to understand, predict, and act in the physical world.\n\n## What the paper actually says\n\nThis is not a product announcement or a benchmark-beating model reveal. It is a position paper, a collective argument from more than 21 researchers about where the field should invest its attention next.\n\nAmong the contributors are Robert Geirhos from Google DeepMind and Yilun Du from Harvard. The work grew out of discussions at the CVPR 2026 Visual General Intelligence Workshop, one of the premier gatherings in the computer vision community.\n\nRather than presenting a single model or architecture, the paper discusses principles, benchmarks, and learning paradigms for building intelligence through visual experience. The researchers argue that generative video models and self-supervised learning, where systems train themselves by predicting what comes next in visual sequences, could provide a foundation for AGI that language alone cannot.\n\nThe paper also explores strategies for integrating multiple modalities. Vision and language aren’t positioned as competitors in this framework but as complementary channels, each capturing different slices of intelligence.\n\n## Building on DeepMind’s AGI framework\n\nThe white paper doesn’t exist in isolation. It builds on a lineage of publications from Google DeepMind that have attempted to map the road to AGI and beyond.\n\nIn 2024, DeepMind published “Levels of AGI,” a paper that proposed a taxonomy for measuring progress toward general intelligence. More recently, in June 2026, DeepMind followed up with “From AGI to ASI,” which explored what might come after general intelligence: artificial superintelligence. The VGI white paper slots neatly into this intellectual arc, asking whether the current text-heavy paradigm is sufficient to reach the milestones those earlier papers described.\n\nFor now, no specific performance claims, model timelines, or product roadmaps accompany the paper. It is deliberately open-ended, designed to spark exploration rather than declare victory.\n\n**Disclosure:** This article was edited by Editorial Team. For more information on how we create and review content, see our\n\n[Editorial Policy](https://cryptobriefing.com/editorial-policy/).", "url": "https://wpnews.pro/news/google-deepmind-and-harvard-propose-vision-first-path-to-agi", "canonical_source": "https://cryptobriefing.com/deepmind-harvard-visual-general-intelligence-agi/", "published_at": "2026-09-11 18:08:13+00:00", "updated_at": "2026-09-11 18:14:31.768986+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-research", "computer-vision", "generative-ai", "ai-safety"], "entities": ["Google DeepMind", "Harvard", "Robert Geirhos", "Yilun Du", "Visual General Intelligence: A White Paper", "arXiv", "CVPR 2026 Visual General Intelligence Workshop", "Levels of AGI"], "alternates": {"html": "https://wpnews.pro/news/google-deepmind-and-harvard-propose-vision-first-path-to-agi", "markdown": "https://wpnews.pro/news/google-deepmind-and-harvard-propose-vision-first-path-to-agi.md", "text": "https://wpnews.pro/news/google-deepmind-and-harvard-propose-vision-first-path-to-agi.txt", "jsonld": "https://wpnews.pro/news/google-deepmind-and-harvard-propose-vision-first-path-to-agi.jsonld"}}