Land or Water? Andrej Karpathy highlighted an LLM evaluation that asks a model "Land or Water?" for a latitude and longitude coordinate 16,200 times and plots the answers as an image, reporting that the models know the map from compressing the internet. A commenter noted a model answering "water" every time scores roughly 70% because most of the planet is ocean, making Sonnet 4.5's 60% worse than never guessing land, while 90%+ is where a model "really starts knowing the map." Another commenter said a model tested on the eval had a grasp of the globe on par with some of the best models from a year earlier, and that the architecture and low cost made it feasible to extract a labeled map of continents and countries. Andrej Karpathy on X: "Cool eval. Simply ask an LLM “Land or Water?” and give it a latitude and longitude coordinate as text. Ask 16,200 times, plot as image. The models know. From compressing the internet." Cool eval. Simply ask an LLM “Land or Water?” and give it a latitude and longitude coordinate as text. Ask 16,200 times, plot as image. The models know. From compressing the internet. Cool eval. Simply ask an LLM “Land or Water?” and give it a latitude and longitude coordinate as text. Ask 16,200 times, plot as image. The models know. From compressing the internet. fun detail: a model that just answers "water" every time scores roughly 70% here, since most of the planet is ocean. so Sonnet 4.5 at 60% is actually worse than never guessing land. 90%+ is where it really starts knowing the map remember this old post of mine? Tried it with Jev---it seems to have a grasp of the globe on par with some of the best models a year ago due to the architecture & low cost, it was also feasible to extract a labeled map of continents and countries. lots of interesting details