about
Neural Symbolic Machines: Learning Semantic Parsers with Weak Supervision (arxiv.org)
94 points by fitzwatermellow on Nov 2, 2016 | hide | past | pdf | 8 comments on HN

In plain words: A system writes a program for each question and runs it on a large fact database, learning from question-answer pairs by trying programs and keeping ones that give the right answer. It beat the best previous approach on a question set without hand-built rules.

Abstract · Neural Symbolic Machines: Learning Semantic Parsers on Freebase with Weak Supervision

Harnessing the statistical power of neural networks to perform language understanding and symbolic reasoning is difficult, when it requires executing efficient discrete operations against a large knowledge-base. In this work, we introduce a Neural Symbolic Machine, which contains (a) a neural "programmer", i.e., a sequence-to-sequence model that maps language utterances to programs and utilizes a key-variable memory to handle compositionality (b) a symbolic "computer", i.e., a Lisp interpreter that performs program execution, and helps find good programs by pruning the search space. We apply REINFORCE to directly optimize the task reward of this structured prediction problem. To train with weak supervision and improve the stability of REINFORCE, we augment it with an iterative maximum-likelihood training process. NSM outperforms the state-of-the-art on the WebQuestionsSP dataset when trained from question-answer pairs only, without requiring any feature engineering or domain-specific knowledge.

Chen Liang, Jonathan Berant, Quoc Le, Kenneth D. Forbus, Ni Lao
arXiv:1611.00020 · cs.CL, cs.AI, cs.LG · submitted Oct 31, 2016 · updated Apr 23, 2017
abstract · pdf · html · ACL 2017 camera ready version

add comment on HN

I know that the field of deep learning / machine learning generally moves too fast for researchers to target conferences or journals, but several of the citations in the PDF of this article are missing (LaTeX inserted [?]).

(GRU, in case anyone is wondering, stands for "gated recurrent unit" and is a building block of standard LSTMs)

EDIT: now that I've finished the paper, I've realized that the citations are straight-up missing. That's no good, but I'm sure the authors just messed up the arxiv upload. If OP knows them, they should let them know... failing to include any citations at all is a quick way to decrease the credibility of an article.

A GRU is more like a simplification of the ideas in LSTM, rather than a building block. At a high level, it uses the hidden state as the memory of the cell (rather than a separate cell state) and it uses a single "update" gate, merging the forget and input gates. Overall it performs similarly to LSTM while being more computationally efficient (fewer matrices).
FWIW, my guess is that a lot of the novel stuff that has been released in the last week is because of the impending ICLR deadline (Friday). The review process for that conference allows the papers to be updated until the reviewers' decisions are made. So getting the text 'finalized' isn't an essential step right now.
I've let them know already. I think it is just a LaTeX compilation issue.
If you look at the source you see that the \cite{} statements do reference meaningful anchors but that the bibtext file seems to be missing from the LaTeX archive:

https://arxiv.org/format/1611.00020v1

A GRU is a simplified version of the LSTM, not a building block.
Hi, I am Chen Liang, the first author of the paper. Thanks for pointing out the Latex problem and sorry for the inconvenience.

We are trying to fix the Latex problem and submit a replacement to ArXiv soon. In the meantime, we hosted the PDF version of the paper on another link:

https://www.researchgate.net/profile/Chen_Liang14/publicatio...

Thanks and look forward to your feedbacks and suggestions :)

Interesting. It wasn't until the end of the paper and re-read the top that I noticed "Liang" isn't actually Percy Liang (with whom Berant has collaborated a lot in the past)