In plain words: Instead of learning from fixed sets of labeled examples, a chatbot learns from its partner's replies during conversation, which act as natural feedback. A system that guesses the teacher's reply before answering learned to answer questions correctly with no reward signal at all.
Abstract · Dialog-based Language Learning
A long-term goal of machine learning research is to build an intelligent dialog agent. Most research in natural language understanding has focused on learning from fixed training sets of labeled data, with supervision either at the word level (tagging, parsing tasks) or sentence level (question answering, machine translation). This kind of supervision is not realistic of how humans learn, where language is both learned by, and used for, communication. In this work, we study dialog-based language learning, where supervision is given naturally and implicitly in the response of the dialog partner during the conversation. We study this setup in two domains: the bAbI dataset of (Weston et al., 2015) and large-scale question answering from (Dodge et al., 2015). We evaluate a set of baseline learning strategies on these tasks, and show that a novel model incorporating predictive lookahead is a promising approach for learning from a teacher's response. In particular, a surprising result is that it can learn to answer questions correctly without any reward-based supervision at all.
Jason Weston
arXiv:1604.06045 · cs.CL · submitted Apr 20, 2016 · updated Oct 24, 2016
abstract · pdf · html