about
Towards De-Identification of Legal Texts (arxiv.org)
2 points by sel1 on Oct 11, 2019 | hide | past | pdf | 1 comment on HN

In plain words: They tested standard text tools on lawsuit documents to see how well they spot and hide people's names while keeping the story readable. The tools missed at least one name in 84% of documents, so they need heavy tuning for legal writing.

Abstract · Towards De-identification of Legal Texts

In many countries, personal information that can be published or shared between organizations is regulated and, therefore, documents must undergo a process of de-identification to eliminate or obfuscate confidential data. Our work focuses on the de-identification of legal texts, where the goal is to hide the names of the actors involved in a lawsuit without losing the sense of the story. We present a first evaluation on our corpus of NLP tools in tasks such as segmentation, tokenization and recognition of named entities, and we analyze several evaluation measures for our de-identification task. Results are meager: 84% of the documents have at least one name not covered by NER tools, something that might lead to the re-identification of involved names. We conclude that tools must be strongly adapted for processing texts of this particular domain.

Diego Garat, Dina Wonsever
arXiv:1910.03739 · cs.CL · submitted Oct 9, 2019
abstract · pdf · html

add comment on HN

I would love to see this work continue; 84% is not bad for a starting point. Fascinating stuff!