conference-paper Open access

Neural Word Segmentation with Rich Pretraining

Research footprint

At a glance

Citations
126
References
49
Comments
0
Paper overview

Abstract

Neural word segmentation research has benefited from large-scale raw texts by leveraging them for pretraining character and word embeddings. On the other hand, statistical segmentation research has exploited richer sources of external information, such as punctuation, automatic segmentation and POS. We investigate the effectiveness of a range of external training sources for neural word segmentation by building a modular segmentation model, pretraining the most important submodule using rich external sources. Results show that such pretraining significantly improves the model, leading to accuracies competitive to the best methods on six benchmarks. * Equal contribution.

Record transparency

Publication details

DOI
10.18653/v1/p17-1078
OpenAlex
W2964093505
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.