conference-paper

A Corpus-Based Sampling to Build Training Data Set for Extracting Japanese Sentence Pattern

Research footprint

At a glance

Citations
1
References
16
Comments
0
Paper overview

Abstract

Training data set plays an important role in Natural Language Processing (NLP) or Machine Learning (ML) Tasks. In the application of NLP in Japanese language education, construction of a high-quality training data set becomes the pre-requisite of automatic extraction of grammar knowledge where there are limited training data sets are available. In this work, a corpus-based method for building training data set was proposed aiming to reach a satisfactory performance in automatic extraction of Japanese sentence patterns in Japanese grammar. Furthermore, a machine learning algorithm based on Conditional Random Field (CRF) was applied to train a model using the manually annotated training data sets in experiments. A comparative evaluation was conducted in terms of our proposed method and a baseline method based on paper-based sampling. Experimental results indicated that our proposed method based on corpus-based sampling to build training data set achieved much higher accuracy than paper-based sampling.

Record transparency

Publication details

DOI
10.1109/iceit54416.2022.9690759
OpenAlex
W4210622232
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.