about
Reflex: Flexible Framework for Relation Extraction in Multiple Domains (arxiv.org)
3 points by sel1 on Jun 23, 2019 | hide | past | pdf | discuss on HN

In plain words: A shared, extendable toolkit tests how text systems pull relationships between named things out of sentences, running the same setup on general, biomedical, and clinical text. Testing many choices showed data cleanup steps mattered most, so papers that skip those details make comparisons unfair.

Abstract · REflex: Flexible Framework for Relation Extraction in Multiple Domains

Systematic comparison of methods for relation extraction (RE) is difficult because many experiments in the field are not described precisely enough to be completely reproducible and many papers fail to report ablation studies that would highlight the relative contributions of their various combined techniques. In this work, we build a unifying framework for RE, applying this on three highly used datasets (from the general, biomedical and clinical domains) with the ability to be extendable to new datasets. By performing a systematic exploration of modeling, pre-processing and training methodologies, we find that choices of pre-processing are a large contributor performance and that omission of such information can further hinder fair comparison. Other insights from our exploration allow us to provide recommendations for future research in this area.

Geeticka Chauhan, Matthew B. A. McDermott, Peter Szolovits
arXiv:1906.08318 · cs.CL · submitted Jun 19, 2019 · updated Jul 20, 2019
abstract · pdf · html · accepted by BioNLP 2019 at the Association of Computation Linguistics 2019

add comment on HN