Can Geography Train AI? When Spatial Relationships Become the Labels
Most GeoAI models still depend on prepared training data. A building detector needs building footprints. A road model needs road centerlines. A land-cover model needs pixel classes. A vision-language model needs captions that explain what appears in an image.
This creates a bottleneck in Earth observation. Satellites, aircraft, and drones keep collecting imagery, but human annotation is slow. Tracing rooftops, drawing roads, marking flood boundaries, or separating crop types often needs domain knowledge. A non-specialist can usually identify a car in a street photo. Labelling a wetland, a seasonal water body, or a mixed crop field from above is harder.
Large datasets show how much effort this requires. Functional Map of the World contains more than one million satellite images from over 200 countries, with bounding-box labels across 63 categories. SpaceNet was built around high-resolution labelled imagery for tasks such as building footprint extraction and road network mapping. DynamicEarthNet combines daily Planet imagery with monthly pixel-level land-cover labels, showing how difficult dense temporal annotation can become.
These datasets are important, but they also show the constraint. High-quality labels are expensive, unevenly distributed, and often tied to specific places or tasks.
This raises the central question of this article: what if some of the information needed to train a GeoAI model already exists in geography itself?
Geography Already Contains Hidden Labels
A satellite image may not have a human label, but it is rarely just a grid of unknown pixels. Its coordinates already place it on the Earth. From that location, we can connect the image to roads, rivers, building footprints, land use, elevation, weather records, and older images of the same place.
These layers are not labels in the usual sense. A digital elevation model does not directly say “urban area.” An OpenStreetMap road line does not describe the full image. A past satellite scene does not replace a human annotation. Still, each one gives the model a clue about the place it is looking at.
Traditional supervision is simple: an image is paired with a human label, and the model learns from that label.
Geographic supervision works differently. The image is paired with spatial context and relationships. A model may learn that a grey roof lies beside a road, that a field sits near an irrigation canal, or that a water boundary follows low-lying terrain. These are weaker than expert labels, but they carry structure that raw pixels alone do not provide.
Early Earth observation and OpenStreetMap experiments showed that map layers can be used together with satellite imagery for semantic labelling. That idea has become more relevant as GeoAI moves toward models that learn from imagery, maps, location, and time together.
Geographical reference imagery is not always the complete ground truth. Maps can be old, incomplete, or misaligned. But they can still act as a source of supervision when labels are limited.

Geographic context layers turning an unlabelled satellite image into a training signal. Source: AI-generated
Five Ways Geography Could Teach a Model
Geography can provide training signals in several ways. These signals are weaker than hand-drawn labels, but they can still help a model learn from the structure around an image.
The first is location. Coordinates are not just numbers. They place an image inside a climate zone, terrain type, settlement pattern, and regional context. SatCLIP learns location embeddings by aligning satellite imagery with geographic coordinates, while GeoCLIP uses location-image alignment for worldwide geolocalization. In both cases, location becomes part of the learning signal.
The second is surrounding context. An unlabelled grey rectangle in an image may be hard to interpret by itself. But if it sits beside a highway, rail yard, port, or irrigation canal, that context gives the model useful clues about what the object might be.
The third is spatial relationship. Geography is not only about what exists, but how things connect. Roads form networks. Buildings sit beside streets. Rivers have upstream and downstream relationships. Sat2Graph uses graph-tensor encoding to extract road graphs from satellite imagery, showing why connectivity matters beyond pixel-level segmentation.
The fourth is time. Satellites revisit the same location repeatedly, creating natural pairs of images from different dates. Seasonal Contrast uses this structure for self-supervised learning, helping models learn what remains stable despite seasonal changes.
The fifth is physical context. Elevation, slope, drainage, coastlines, and hydrology can constrain what is likely or unlikely in an image. A flood prediction on a steep slope, for example, should be treated differently from one in a low-lying floodplain. Physics-guided flood modelling combines remote-sensing imagery with DEM-derived terrain features and hydrodynamic constraints.
Together, these signals suggest a different way to think about supervision. Geography is not only the thing GeoAI tries to map. It can also help teach the model how the world is structured.
This Is Already Starting to Happen
Using geography as a training signal is no longer only a concept. Several recent systems are testing how maps, geographic priors, and spatial graphs can influence how remote-sensing models learn.
These systems are different, but they share one idea: geographic data is beginning to move into the training process itself. It is not only something a model reads after prediction. It can also help form the representation the model learns.
Why This Could Change GeoAI
Geographic supervision could reduce some dependence on manual annotation. Open maps, elevation data, satellite archives, and repeated observations already describe parts of the world at scale. They cannot replace expert labels, but they can give models useful training signals before humans begin drawing polygons. SSL4EO-S12 shows how unlabeled Sentinel-1 and Sentinel-2 archives can support self-supervised pretraining across sensors, seasons, and locations.
It could also make learned representations richer. A model trained only on pixels may learn that a roof has a certain colour or texture. A model trained with geographic context can also learn that the roof sits beside a road, falls inside a residential area, and remains stable across several dates.
Regional transfer may improve, but not automatically. A road in Germany, Kenya, and India may look different, but its network role and relation to nearby buildings may carry useful structure.
This also brings GIS and computer vision closer together. Rasters, vector maps, terrain data, and time-series observations can become part of training, not only post-processing. UN-Habitat’s GeoAI toolkit reflects this broader use of satellite imagery, geospatial data, and planning information together.
GeoAI has mostly learned patterns inside geographic data. It may increasingly learn from the structure of geography itself.
But Geography Alone Is Not Always Ground Truth
Geographic data can be useful, but it can also be incomplete or outdated.
OpenStreetMap is useful because it provides roads, buildings, land use, and other mapped features across many places. But coverage is uneven. One global study estimated the OSM road network to be about 83% complete, while another found strong spatial inequalities in OSM building completeness. . If a training pipeline treats “no mapped building” as “no building exists,” the model may learn false negatives in places where mapping is incomplete.

Spatial distribution of OSM building completeness in 13,189 urban centers. Source: Herfort et al., 2023
Time adds another problem. A road, building, or land-use boundary may remain in a map after the landscape has changed. When an old vector layer is paired with newer imagery, the model receives a conflicting signal.
Coordinates can also become shortcuts. A model may learn that a crop is common in one region instead of learning its visual and spatial characteristics. That can make results look good in nearby test areas but weaker in new regions.
So geographic supervision should be treated as evidence, not truth. The stronger approach combines human labels, imagery, maps, location, time, spatial relationships, and physical context.
GeoAI has usually been trained to learn about geography. The next step may be letting geography itself become part of how the model learns.
Did you like this post? Follow us on our social media channels!
Read more and subscribe to our monthly newsletter!





















