about
Explain Yourself Leveraging Language Models for Commonsense Reasoning (arxiv.org)
51 points by sel1 on Jun 9, 2019 | hide | past | pdf | 2 comments on HN

In plain words: People wrote short explanations for commonsense questions, and a language model learned to write its own, which is then added to the question when the system trains and answers. This beat the previous best score on a hard commonsense quiz by 10%.

Abstract · Explain Yourself! Leveraging Language Models for Commonsense Reasoning

Deep learning models perform poorly on tasks that require commonsense reasoning, which often necessitates some form of world-knowledge or reasoning over information not immediately present in the input. We collect human explanations for commonsense reasoning in the form of natural language sequences and highlighted annotations in a new dataset called Common Sense Explanations (CoS-E). We use CoS-E to train language models to automatically generate explanations that can be used during training and inference in a novel Commonsense Auto-Generated Explanation (CAGE) framework. CAGE improves the state-of-the-art by 10% on the challenging CommonsenseQA task. We further study commonsense reasoning in DNNs using both human and auto-generated explanations including transfer to out-of-domain tasks. Empirical results indicate that we can effectively leverage language models for commonsense reasoning.

Nazneen Fatema Rajani, Bryan McCann, Caiming Xiong, Richard Socher
arXiv:1906.02361 · cs.CL · submitted Jun 6, 2019
abstract · pdf · html · Accepted at ACL, 11 pages total

add comment on HN

The model is trained end-to-end to answer common-sense-reasoning questions after generating explanations for its answers, using sample human explanations as part of the training data.

This results in improved performance on the question-answering task.

This is fascinating, although in hindsight, not entirely surprising: inducing a machine to learn to model human explanations helps the machine perform better in testing.

A natural question follows:

Can we find ways to induce much larger models to learn to generate human explanations about a growing number of subjects of increasing complexity?

Reminds me of the concept of "Social Stories": https://en.wikipedia.org/wiki/Social_Stories

> Social Stories are a concept devised by Carol Gray in 1991 to improve the social skills of people with autism spectrum disorders (ASD). The objective is to share information, which is often through a description of the events occurring around the subject and also why.

I'm wondering if this kind of "common sense" interactions could be leveraged to train models?

Here's a concrete example, "Being Angry and Safe": https://youtu.be/R8c_Br8I_Tc?t=28