about
Open High-Resolution Satellite Imagery: The WorldStrat Dataset (arxiv.org)
13 points by Jimmc414 on Nov 20, 2022 | hide | past | pdf | 1 comment on HN

In plain words: A free collection pairs nearly 10,000 square kilometers of sharp 1.5-meter satellite photos with matching blurrier 10-meter shots, covering every land type plus places like refugee settlements and illegal mines. Trained on it, simple programs learn to sharpen the free blurry images toward the detail of costly sharp ones.

Abstract · Open High-Resolution Satellite Imagery: The WorldStrat Dataset -- With Application to Super-Resolution

Analyzing the planet at scale with satellite imagery and machine learning is a dream that has been constantly hindered by the cost of difficult-to-access highly-representative high-resolution imagery. To remediate this, we introduce here the WorldStrat dataset. The largest and most varied such publicly available dataset, at Airbus SPOT 6/7 satellites' high resolution of up to 1.5 m/pixel, empowered by European Space Agency's Phi-Lab as part of the ESA-funded QueryPlanet project, we curate nearly 10,000 sqkm of unique locations to ensure stratified representation of all types of land-use across the world: from agriculture to ice caps, from forests to multiple urbanization densities. We also enrich those with locations typically under-represented in ML datasets: sites of humanitarian interest, illegal mining sites, and settlements of persons at risk. We temporally-match each high-resolution image with multiple low-resolution images from the freely accessible lower-resolution Sentinel-2 satellites at 10 m/pixel. We accompany this dataset with an open-source Python package to: rebuild or extend the WorldStrat dataset, train and infer baseline algorithms, and learn with abundant tutorials, all compatible with the popular EO-learn toolbox. We hereby hope to foster broad-spectrum applications of ML to satellite imagery, and possibly develop from free public low-resolution Sentinel2 imagery the same power of analysis allowed by costly private high-resolution imagery. We illustrate this specific point by training and releasing several highly compute-efficient baselines on the task of Multi-Frame Super-Resolution. High-resolution Airbus imagery is CC BY-NC, while the labels and Sentinel2 imagery are CC BY, and the source code and pre-trained models under BSD. The dataset is available at https://zenodo.org/record/6810791 and the software package at https://github.com/worldstrat/worldstrat .

Julien Cornebise, Ivan Oršolić, Freddie Kalaitzis
arXiv:2207.06418 · eess.IV, cs.CV, cs.LG, stat.AP · submitted Jul 13, 2022 · updated May 31, 2025
abstract · pdf · html · Published in 36th Conference on Neural Information Processing Systems (NeurIPS 2022) Track on Datasets and Benchmarks

add comment on HN