In plain words: Bongard puzzles ask you to find the rule shared by example images, then classify a new one. Rules often need several examples together, but usual systems judge one by one; using the whole group lifted accuracy on natural images to 76.4%, beating the previous best.
Abstract · Support-Set Context Matters for Bongard Problems
Current machine learning methods struggle to solve Bongard problems, which are a type of IQ test that requires deriving an abstract "concept" from a set of positive and negative "support" images, and then classifying whether or not a new query image depicts the key concept. On Bongard-HOI, a benchmark for natural-image Bongard problems, most existing methods have reached at best 69% accuracy (where chance is 50%). Low accuracy is often attributed to neural nets' lack of ability to find human-like symbolic rules. In this work, we point out that many existing methods are forfeiting accuracy due to a much simpler problem: they do not adapt image features given information contained in the support set as a whole, and rely instead on information extracted from individual supports. This is a critical issue, because the "key concept" in a typical Bongard problem can often only be distinguished using multiple positives and multiple negatives. We explore simple methods to incorporate this context and show substantial gains over prior works, leading to new state-of-the-art accuracy on Bongard-LOGO (75.3%) and Bongard-HOI (76.4%) compared to methods with equivalent vision backbone architectures and strong performance on the original Bongard problem set (60.8%).
Nikhil Raghuraman, Adam W. Harley, Leonidas Guibas
arXiv:2309.03468 · cs.CV, cs.AI, cs.LG · submitted Sep 7, 2023 · updated Dec 1, 2024
abstract · pdf · html · TMLR October 2024. Code: https://github.com/nraghuraman/bongard-context
Here's a brief summary: Bongard problems are a type of visual reasoning task where the human/AI sees two sets of images: a “positive” set where each image depicts a certain concept and a “negative” set where images don’t depict the concept. The goal is to determine the concept given these sets. Unfortunately, while humans can often solve these problems in the ~90% accuracy range, AI performance is often in the ~60% range.
In our work, we devise simple methods that incorporate "cross-image context" to attain higher performance. "Cross-image context" means that we consider the similarities and differences between multiple positive and negative images. We find that performance of simple existing methods (e.g., k-nearest neighbors on deep features) can be dramatically improved by incorporating cross-image context, sometimes up to 10%!
In all, our work attains state-of-the-art performance on the two main Bongard datasets.
Let us know what you think!