A Corpus-Based Sampling to Build Training Data Set for Extracting Japanese Sentence Pattern
At a glance
- Citations
- 1
- References
- 16
- Comments
- 0
Abstract
Training data set plays an important role in Natural Language Processing (NLP) or Machine Learning (ML) Tasks. In the application of NLP in Japanese language education, construction of a high-quality training data set becomes the pre-requisite of automatic extraction of grammar knowledge where there are limited training data sets are available. In this work, a corpus-based method for building training data set was proposed aiming to reach a satisfactory performance in automatic extraction of Japanese sentence patterns in Japanese grammar. Furthermore, a machine learning algorithm based on Conditional Random Field (CRF) was applied to train a model using the manually annotated training data sets in experiments. A comparative evaluation was conducted in terms of our proposed method and a baseline method based on paper-based sampling. Experimental results indicated that our proposed method based on corpus-based sampling to build training data set achieved much higher accuracy than paper-based sampling.
Publication details
- DOI
- 10.1109/iceit54416.2022.9690759
- OpenAlex
- W4210622232
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.