about
Information Extraction with Neural Networks and Free Noisy Supervision (arxiv.org)
60 points by breck on Dec 10, 2017 | hide | past | pdf | 3 comments on HN

In plain words: A character-level neural network is added to a rule-based text parser and trained for free by checking its extracted data against existing databases, instead of hand-labeled examples. It sharply raised precision over Bloomberg's heavily tuned production system for financial text.

Abstract · Information Extraction with Character-level Neural Networks and Free Noisy Supervision

We present an architecture for information extraction from text that augments an existing parser with a character-level neural network. The network is trained using a measure of consistency of extracted data with existing databases as a form of noisy supervision. Our architecture combines the ability of constraint-based information extraction systems to easily incorporate domain knowledge and constraints with the ability of deep neural networks to leverage large amounts of data to learn complex features. Boosting the existing parser's precision, the system led to large improvements over a mature and highly tuned constraint-based production information extraction system used at Bloomberg for financial language text.

Philipp Meerkamp, Zhengyi Zhou
arXiv:1612.04118 · cs.CL, cs.IR, cs.LG · submitted Dec 13, 2016 · updated Jan 24, 2017
abstract · pdf · html

add comment on HN

How does the pipeline compare to state-of-the-art information extraction suites (e.g. Apache OpenNLP) on standard datasets?
Where's the code?
Well. That was a short paper.