#Contributing Writers #GeoAI #GeoDev #Ideas #Insights

Can Geography Train AI? When Spatial Relationships Become the Labels

Most GeoAI models still depend on prepared training data. A building detector needs building footprints. A road model needs road centerlines. A land-cover model needs pixel classes. A vision-language model needs captions that explain what appears in an image.

This creates a bottleneck in Earth observation. Satellites, aircraft, and drones keep collecting imagery, but human annotation is slow. Tracing rooftops, drawing roads, marking flood boundaries, or separating crop types often needs domain knowledge. A non-specialist can usually identify a car in a street photo. Labelling a wetland, a seasonal water body, or a mixed crop field from above is harder.

Large datasets show how much effort this requires. Functional Map of the World contains more than one million satellite images from over 200 countries, with bounding-box labels across 63 categories. SpaceNet was built around high-resolution labelled imagery for tasks such as building footprint extraction and road network mapping. DynamicEarthNet combines daily Planet imagery with monthly pixel-level land-cover labels, showing how difficult dense temporal annotation can become.

 


One example per category in Functional Map of the World
One example per category in Functional Map of the World

Sample output of baseline algorithm applied to SpaceNet test data
Sample output of baseline algorithm applied to SpaceNet test data

Visualization of the DynamicEarthNet dataset
Visualization of the DynamicEarthNet dataset

These datasets are important, but they also show the constraint. High-quality labels are expensive, unevenly distributed, and often tied to specific places or tasks.

This raises the central question of this article: what if some of the information needed to train a GeoAI model already exists in geography itself?

Geography Already Contains Hidden Labels

A satellite image may not have a human label, but it is rarely just a grid of unknown pixels. Its coordinates already place it on the Earth. From that location, we can connect the image to roads, rivers, building footprints, land use, elevation, weather records, and older images of the same place.

These layers are not labels in the usual sense. A digital elevation model does not directly say “urban area.” An OpenStreetMap road line does not describe the full image. A past satellite scene does not replace a human annotation. Still, each one gives the model a clue about the place it is looking at.

Traditional supervision is simple: an image is paired with a human label, and the model learns from that label.

Geographic supervision works differently. The image is paired with spatial context and relationships. A model may learn that a grey roof lies beside a road, that a field sits near an irrigation canal, or that a water boundary follows low-lying terrain. These are weaker than expert labels, but they carry structure that raw pixels alone do not provide.

Early Earth observation and OpenStreetMap experiments showed that map layers can be used together with satellite imagery for semantic labelling. That idea has become more relevant as GeoAI moves toward models that learn from imagery, maps, location, and time together.

Geographical reference imagery is not always the complete ground truth. Maps can be old, incomplete, or misaligned. But they can still act as a source of supervision when labels are limited.

Geographic context layers turning an unlabelled satellite image into a training signal. Source: AI-generated

Five Ways Geography Could Teach a Model

Geography can provide training signals in several ways. These signals are weaker than hand-drawn labels, but they can still help a model learn from the structure around an image.

The first is location. Coordinates are not just numbers. They place an image inside a climate zone, terrain type, settlement pattern, and regional context. SatCLIP learns location embeddings by aligning satellite imagery with geographic coordinates, while GeoCLIP uses location-image alignment for worldwide geolocalization. In both cases, location becomes part of the learning signal.

The second is surrounding context. An unlabelled grey rectangle in an image may be hard to interpret by itself. But if it sits beside a highway, rail yard, port, or irrigation canal, that context gives the model useful clues about what the object might be.

The third is spatial relationship. Geography is not only about what exists, but how things connect. Roads form networks. Buildings sit beside streets. Rivers have upstream and downstream relationships. Sat2Graph uses graph-tensor encoding to extract road graphs from satellite imagery, showing why connectivity matters beyond pixel-level segmentation.

The fourth is time. Satellites revisit the same location repeatedly, creating natural pairs of images from different dates. Seasonal Contrast uses this structure for self-supervised learning, helping models learn what remains stable despite seasonal changes.

The fifth is physical context. Elevation, slope, drainage, coastlines, and hydrology can constrain what is likely or unlikely in an image. A flood prediction on a steep slope, for example, should be treated differently from one in a low-lying floodplain. Physics-guided flood modelling combines remote-sensing imagery with DEM-derived terrain features and hydrodynamic constraints.

Together, these signals suggest a different way to think about supervision. Geography is not only the thing GeoAI tries to map. It can also help teach the model how the world is structured.

This Is Already Starting to Happen

Using geography as a training signal is no longer only a concept. Several recent systems are testing how maps, geographic priors, and spatial graphs can influence how remote-sensing models learn.

 


OSM-CLIP paper illustration

OSM-CLIP

Uses OpenStreetMap annotations as spatial supervision for remote-sensing image-text learning. It links roads, buildings, land-use areas, and other mapped features to satellite patches.

Reported result: 10.81 percentage-point average zero-shot gain over RemoteCLIP across 13 benchmarks.


Read source paper →


OSMDA paper illustration

OSMDA

Pairs aerial images with rendered OpenStreetMap tiles. A vision-language model reads map-like graphics and text to generate OSM-enriched captions for overhead imagery.

Key idea: reduce dependence on manually written captions and external teacher models.


Read source paper →


GeoPriorCLIP paper illustration

GeoPriorCLIP

Adds geographic priors to remote-sensing vision-language learning, using map-derived geometry, topology, and semantic attributes to guide feature alignment.

Key idea: spatial context becomes part of image-text representation learning.


Read source paper →


GeoLink paper illustration

GeoLink

Integrates OpenStreetMap vector data into remote-sensing foundation-model pretraining by connecting raster imagery with OSM entity graphs and spatial relationships.

Key idea: connect satellite pixels with map objects and their relationships.


Read source paper →

These systems are different, but they share one idea: geographic data is beginning to move into the training process itself. It is not only something a model reads after prediction. It can also help form the representation the model learns.

Why This Could Change GeoAI

Geographic supervision could reduce some dependence on manual annotation. Open maps, elevation data, satellite archives, and repeated observations already describe parts of the world at scale. They cannot replace expert labels, but they can give models useful training signals before humans begin drawing polygons. SSL4EO-S12 shows how unlabeled Sentinel-1 and Sentinel-2 archives can support self-supervised pretraining across sensors, seasons, and locations.

It could also make learned representations richer. A model trained only on pixels may learn that a roof has a certain colour or texture. A model trained with geographic context can also learn that the roof sits beside a road, falls inside a residential area, and remains stable across several dates.

Regional transfer may improve, but not automatically. A road in Germany, Kenya, and India may look different, but its network role and relation to nearby buildings may carry useful structure.

This also brings GIS and computer vision closer together. Rasters, vector maps, terrain data, and time-series observations can become part of training, not only post-processing. UN-Habitat’s GeoAI toolkit reflects this broader use of satellite imagery, geospatial data, and planning information together.

GeoAI has mostly learned patterns inside geographic data. It may increasingly learn from the structure of geography itself.

But Geography Alone Is Not Always Ground Truth

Geographic data can be useful, but it can also be incomplete or outdated.

OpenStreetMap is useful because it provides roads, buildings, land use, and other mapped features across many places. But coverage is uneven. One global study estimated the OSM road network to be about 83% complete, while another found strong spatial inequalities in OSM building completeness. . If a training pipeline treats “no mapped building” as “no building exists,” the model may learn false negatives in places where mapping is incomplete.

Spatial distribution of OSM building completeness in 13,189 urban centers. Source: Herfort et al., 2023

Time adds another problem. A road, building, or land-use boundary may remain in a map after the landscape has changed. When an old vector layer is paired with newer imagery, the model receives a conflicting signal.

Coordinates can also become shortcuts. A model may learn that a crop is common in one region instead of learning its visual and spatial characteristics. That can make results look good in nearby test areas but weaker in new regions.

So geographic supervision should be treated as evidence, not truth. The stronger approach combines human labels, imagery, maps, location, time, spatial relationships, and physical context.

GeoAI has usually been trained to learn about geography. The next step may be letting geography itself become part of how the model learns.


Did you like this post? Follow us on our social media channels!

Read more and subscribe to our monthly newsletter!

Say thanks for this article (0)
Our community is supported by:
Become a sponsor
#Contributing Writers
#Fun #GeoDev #Ideas #Insights
What Facebook Friendships Reveal About Europe’s Hidden Borders
Sebastian Walczak 07.5.2026
AWESOME 2
#Contributing Writers #Ideas #Insights #Science
Geo Embeddings Explained: Understanding Locations as Vectors
Aravindh Subramanian 12.21.2025
AWESOME 2
#GeoDev #Ideas #Insights
What Are Coordinate Systems and Why Do They Matter in Mapping
Sebastian Walczak 12.6.2025
AWESOME 0
Next article
#News

Tech for Earth: Methane, Penguins, and a Volcano Nobody Was Watching

Google and NASA released an AI model that found 23,000 methane plumes nobody had spotted, including at 24 of the 25 worst-emitting landfills on Earth. A glacier fell off Langtang Lirung and killed hundreds of people 100 km downstream. Europe put two Earth observation satellites up on one rocket for the first time. And a team at Durham worked out that the bright moving smudges in their winter radar images were emperor penguins.

Finding Methane Nobody Was Looking At

Google Research and NASA JPL published MAPL-EMIT in PNAS this month, and released the model, the code and the resulting plume database. The short version: it reads the full radiance spectrum from EMIT, the hyperspectral instrument on the International Space Station, and picks out methane plumes at 60 metres per pixel.

The technical move is combining spectral and spatial information in one vision transformer rather than running a matched filter per pixel. That lets the model use plume shape to trace a source, and to separate overlapping plumes in dense industrial areas where conventional methods smear them together. It was trained on 3.6 million physics-based synthetic plumes injected into real EMIT radiance data.

The results are worth stating precisely because they get rounded off in coverage. Against hand-annotated NASA L2B plume complexes across 1,084 EMIT granules, the model recovered 84 percent. It also flagged roughly 1.5 times as many plausible plumes as human analysts, which is where the “50 percent more” figure comes from, and about 23,000 additional plumes globally. Those two numbers measure different things, and the second is the more interesting one: these are candidate detections, not confirmed leaks, and the paper backs them with airborne comparisons and controlled release experiments rather than asserting them outright.

The landfill finding is the part with immediate policy weight. The model located plumes at 24 of the world’s 25 largest-emitting landfills. Europe’s Methane Regulation mandated satellite monitoring and a super-emitter alert system, but covers oil, gas and coal, not waste.

23,000
Extra methane plumes MAPL-EMIT found beyond the existing product
84%
Share of hand-annotated NASA plume complexes the model recovered
5 km
WeatherNext 3 surface temperature grid, down from 25 km
100 km
Distance the Nepal flood travelled downstream from the collapse
8 years
Sentinel-1 winter imagery behind the penguin colony tracking
Explore the data
Global methane plume map
Every plume MAPL-EMIT found, at 60 metres per pixel, released under a Creative Commons licence. The database sits in Earth Engine, the model is on Kaggle and the code is on GitHub.

Read the release →

Google Research and NASA JPL, published in PNAS

Weather Forecasting Stops Waiting

Google DeepMind and Google Research released WeatherNext 3 on 3 September. The headline numbers are resolution and cadence: a 5 km grid for surface temperature and moisture against 25 km before, refreshed hourly against every six hours.

The structural change underneath is more interesting than the accuracy claim. Most AI weather models train on and initialise from ECMWF analysis, which takes around five hours to assemble and updates every six. WeatherNext 3 layers live geostationary satellite mosaics on top, which cuts the lag between the last observation and the forecast from roughly seven hours to three or four. For nowcasting and severe weather that gap is the whole game.

It also outputs wind speed at 100 metres, near turbine hub height, plus cloud cover and solar radiation, which is a clear move towards grid operators and renewables forecasting rather than consumer weather. The model is running in Search, Gemini, Maps, the Maps Platform Weather API and Cloud.

One caveat worth carrying: the 50 percent precipitation improvement is Google’s own figure against its own predecessor. Independent live comparison exists through Brightband’s open leaderboard, but a peer-reviewed architecture paper had not appeared at the time of the announcement.

Side-by-side WeatherNext 2 and WeatherNext 3 temperature forecasts over the UK

Two-metre temperature forecasts over the UK. WeatherNext 2 on the left at 25 km (0.25°), WeatherNext 3 on the right at a native 5 km (0.05°). The finer grid resolves local topography instead of smoothing it into blocks. Credit: Google DeepMind

Nepal, and What the Satellites Saw

On the morning of 26 August a section of glacier on the north face of Langtang Lirung collapsed. The impact released energy equivalent to a magnitude 5.2 earthquake. The resulting flow of water, ice and rock travelled nearly 100 km down the Lende Khola and Trishuli valleys, destroyed the Gyirong Port crossing on the China border, and struck dozens of settlements. Hundreds were killed and thousands reported missing.

ESA published before and after imagery on 31 August, and the acquisition detail matters. The clearest pairing is a Landsat 9 scene from 26 August, roughly two hours after the collapse, against Sentinel-2 from 24 August. Both were processed through shortwave infrared, which separates water and ice from cloud in a valley system that is cloud-covered most of the time. Copernicus Emergency Management Service was activated for flood extent and damage assessment.

The attribution question is still open and should be treated that way. Analysis of satellite imagery found the glacier north of Langtang Lirung moving around 10 mm per month between January and August, with the rate accelerating in the weeks before failure. Researchers have pointed to warm conditions and permafrost instability as plausible contributors, since meltwater working into crevasses weakens ice-rock bonds. No rain was recorded at Rasuwa district headquarters that morning. That is a coherent picture, not a proven mechanism, and the published work so far is careful about saying so.

Before, 24 AugustSentinel-2 image of the Langtang valley before the flood
Sentinel-2 shortwave infrared, two days before the collapse. The Lende Khola and Trishuli valleys are narrow and clear.
After, 26 AugustLandsat 9 image of the same valley after the glacier collapse
Landsat 9, about two hours after the glacier failed. Shortwave infrared separates water and ice from the cloud that usually covers this region.
Images: contains modified Copernicus Sentinel data (2026) and Landsat 9 data, processed by ESA

Sentinel-2 natural colour before and after the Nepal flash flood

The same event in natural colour: Sentinel-2 on 27 August, the day after, against 12 August before the flood. Credit: contains modified Copernicus Sentinel data (2026), processed by ESA

Europe Launches Two at Once

Flight VV30 lifted off from Kourou at 03:21 CEST on 15 September carrying both FLEX and Copernicus Sentinel-3C. It was Vega-C’s first dual launch, using the Vespa adapter to stack the two: Sentinel-3C on top, deployed first, with FLEX encapsulated below and injected about an hour later.

Sentinel-3C continues the operational workhorse line, carrying OLCI, SLSTR, the SAR altimeter and a microwave radiometer for ocean colour, surface temperature and topography, feeding near real-time ocean and weather forecasting.

FLEX is the more unusual of the two. The Fluorescence Explorer measures chlorophyll fluorescence from terrestrial vegetation, the faint light plants re-emit during photosynthesis. That signal is a direct indicator of photosynthetic activity rather than a proxy inferred from greenness, which is what vegetation indices give you. It flies at 814 km in a 27-day repeat, designed to fly in tandem with Sentinel-3 so the fluorescence retrieval can be corrected using coincident optical and thermal data.
Vega-C lifts off carrying FLEX and Sentinel-3C

Flight VV30 leaving Kourou at 03:21 CEST on 15 September, Vega-C’s first dual launch. Credit: ESA

Diagram of FLEX flying in tandem with Sentinel-3

FLEX flies in tandem with Sentinel-3 so the faint fluorescence signal can be corrected using coincident atmospheric and land-surface data. Credit: ESA

Penguins in the Dark

The nicest piece of remote sensing this month came out of Durham University. Emperor penguins breed through the Antarctic winter, which is precisely when optical satellites are useless because there is no sunlight for months. Nearly 70 colonies sit on fast ice around the continent, many never visited by anyone.

Grant Macdonald’s team noticed bright clusters of pixels drifting around on Sentinel-1 SAR imagery against the smooth dark backdrop of stable fast ice. Icebergs and rough ice are also bright, but they do not move when the ice is fixed. Comparing against optical imagery from September, when light returns, confirmed the moving smudges were the colonies themselves. Birds about 1.2 metres tall, standing in dense huddles, scatter enough to show up against flat ice.

They then tracked three colonies at Atka Bay, Coulman Island and Cape Washington across eight breeding seasons from 2017 to 2024, often at sub-weekly resolution, combining winter SAR with summer optical. The behavioural findings are the payoff: colonies move more in winter than expected, sub-groups sometimes move away from the fast ice edge together despite the ice appearing stable, and one colony repeatedly shifted from sea ice onto glacier ice in early spring. The work is in Communications Earth and Environment, and all of it used free, routinely collected data that had been sitting in the archive for years.
Study sites and emperor penguin colonies visible in Sentinel-1 SAR

Study sites, and the colonies as they appear in Sentinel-1 SAR at Atka Bay, Coulman Island and Cape Washington: bright clusters against dark, smooth fast ice. Credit: Macdonald et al., Communications Earth and Environment, CC BY 4.0

Penguin guano and colonies in near-simultaneous optical and SAR imagery

The validation step: optical and SAR of the same spot within a day. Yellow arrows mark the bright SAR returns where guano is visible optically. Credit: Macdonald et al., Communications Earth and Environment, CC BY 4.0

NISAR Watches a Volcano for Nine Months

NISAR imaged Krasheninnikov on Kamchatka on 25 December 2025, while still finishing post-launch checks. Days later the volcano began erupting for the first time since roughly 1550, apparently woken by the magnitude 8.8 earthquake offshore in July 2025.

NASA assembled 17 frames through mid-August into a time-lapse showing lava filling an inner caldera, overflowing into the wider crater, then spreading into a fan. Each pixel covers about 10 by 10 metres, and lava reads brighter than the surrounding snow and rock because it scatters microwaves differently.

The point the science team makes is about consistency rather than any single image. Twice every 12 days, in high-resolution mode, from two look directions, over a volcano that nobody was monitoring closely because it had been quiet for five centuries. As Cornell’s Matthew Pritchard put it, there are volcanoes worldwide that have never had eyes on them like this.

Seventeen NISAR frames from December 2025 to August 2026. Credit: NASA’s Scientific Visualization Studio

The Month in Pictures

A few more frames from the month that did not fit above: the launch that put two Earth observation missions up at once, the valley in Nepal before it was destroyed, and eight winters of penguins.

Click any image to open it full size at the source.

Vega-C climbing out of Kourou

Vega-C climbing out of Kourou
Vega-C takes FLEX and Sentinel-3C into orbit on flight VV30.
ESA

Two satellites, one fairing

Two satellites, one fairing
Sentinel-3C on top of the Vespa adapter, FLEX encapsulated below it.
ESA

What FLEX actually measures

What FLEX actually measures
The faint fluorescence plants re-emit while photosynthesising.
ESA

ESOC, Darmstadt

ESOC, Darmstadt
Where the launch and early orbit phase is run for both satellites.
ESA

The Trishuli valley before

The Trishuli valley before
Sentinel-2 over the region northwest of Kathmandu ahead of the flood.
ESA / Copernicus

Atka Bay through one winter

Atka Bay through one winter
The colony tracked across winter 2023 in Sentinel-1 IW HH imagery, inside the dashed boundary.
Macdonald et al., CC BY 4.0

A year of colony movement

A year of colony movement
Atka Bay tracks, dot colour is day of year. Multiple dots per date mean the colony has split.
Macdonald et al., CC BY 4.0

Tools and Data

A few releases worth knowing about. Mapbox introduced location infrastructure aimed at AI agents rather than human map users, which is a telling shift in who the customer is. Vexcel launched UltraCam Condor 5.0, claiming the largest aerial camera footprint available. EarthDefine released a 3D building footprints API. And on the open side, Overture shipped its September data release.

From the LinkedIn feed, two open-source items stood out: GeoLibre, a free GIS that runs across platforms, and a walkthrough of turning 2D building footprints into 2.5D cityscapes. There was also a genuinely useful explainer on coordinate systems versus projections, which remains the single most common source of confusion for people starting out, and a list of 48 remote sensing definitions worth bookmarking.

Try it yourself
Overture September release
The latest monthly drop of open base map data: addresses, buildings, places, transportation and divisions, free to download and use.

Read the release notes →

Overture Maps Foundation

Worth a Look

NASA opened the Earth Modeling Nexus and detailed the Hamaq mission. USGS on Landsat products supporting water management worldwide. ESA’s Biomass mission imaged mangrove degradation in the Niger Delta, and Sentinel-3 caught Anak Krakatau erupting. NOAA’s summer drought summary in 12 maps is a good example of communicating a slow hazard. Esri opened the 2026 StoryMaps competition and published a guide to visualising time in Map Viewer. GIJN put out a course on geospatial data investigations for reporters. On the lighter side, PetaPixel on the first autumn colour of 2026 seen from orbit, Google Maps Mania on this year’s foliage map, and La Brujula Verde on the Tabula Peutingeriana, the only surviving road map of the Roman Empire.


Compiled by the Geoawesome team. Got something we missed? Reach out, we read everything.

Did you like this post? Read more and subscribe to our monthly newsletter!

Read on
Search