about
Pkuseg: A Toolkit for Multi-Domain Chinese Word Segmentation (arxiv.org)
1 point by sel1 on Jun 29, 2019 | hide | past | pdf | discuss on HN

In plain words: A free toolkit splits Chinese text into words with a different model for each area, like web or medicine, not one for all. For new areas with little labeled text, it makes examples from unlabeled text via translation, and it worked well across several areas.

Abstract · PKUSEG: A Toolkit for Multi-Domain Chinese Word Segmentation

Chinese word segmentation (CWS) is a fundamental step of Chinese natural language processing. In this paper, we build a new toolkit, named PKUSEG, for multi-domain word segmentation. Unlike existing single-model toolkits, PKUSEG targets multi-domain word segmentation and provides separate models for different domains, such as web, medicine, and tourism. Besides, due to the lack of labeled data in many domains, we propose a domain adaptation paradigm to introduce cross-domain semantic knowledge via a translation system. Through this method, we generate synthetic data using a large amount of unlabeled data in the target domain and then obtain a word segmentation model for the target domain. We also further refine the performance of the default model with the help of synthetic data. Experiments show that PKUSEG achieves high performance on multiple domains. The new toolkit also supports POS tagging and model training to adapt to various application scenarios. The toolkit is now freely and publicly available for the usage of research and industry.

Ruixuan Luo, Jingjing Xu, Yi Zhang, Zhiyuan Zhang, Xuancheng Ren, Xu Sun
arXiv:1906.11455 · cs.CL · submitted Jun 27, 2019 · updated May 28, 2022
abstract · pdf · html

add comment on HN