about
Learning to Interpret Satellite Images in Global Scale Using Wikipedia (arxiv.org)
2 points by sel1 on Aug 13, 2019 | hide | past | pdf | discuss on HN

In plain words: Satellite photos are paired with Wikipedia articles about the same places, and a model learns by guessing facts about the article from the image. This beat the usual start from a model trained on labeled photos, raising its F1 score by up to 4.5%.

Abstract

Despite recent progress in computer vision, finegrained interpretation of satellite images remains challenging because of a lack of labeled training data. To overcome this limitation, we construct a novel dataset called WikiSatNet by pairing georeferenced Wikipedia articles with satellite imagery of their corresponding locations. We then propose two strategies to learn representations of satellite images by predicting properties of the corresponding articles from the images. Leveraging this new multi-modal dataset, we can drastically reduce the quantity of human-annotated labels and time required for downstream tasks. On the recently released fMoW dataset, our pre-training strategies can boost the performance of a model pre-trained on ImageNet by up to 4:5% in F1 score.

Burak Uzkent, Evan Sheehan, Chenlin Meng, Zhongyi Tang, Marshall Burke, David Lobell, Stefano Ermon
arXiv:1905.02506 · cs.CV, cs.LG · submitted May 7, 2019 · updated Aug 11, 2019
abstract · pdf · html · Accepted to IJCAI 2019

add comment on HN