conference-paper Open access

Long Unit Word Tokenization and Bunsetsu Segmentation of Historical Japanese

Research footprint

At a glance

Citations
1
References
0
Comments
0
Paper overview

Abstract

In Japanese, "bunsetsu" is the natural minimal phrase of a sentence; it serves as a natural boundary of a sentence for native speakers rather than words, and thus grammatical analysis in Japanese linguistics commonly operates on the basis of bunsetsu units.By contrast, because Japanese does not have delimiters between words, there are two major categories of word definitions: Short Unit Words (SUWs) and Long Unit Words (LUWs).SUW dictionaries are available, whereas LUW dictionaries are not.Hence, this study focuses on providing deep learning-based (or LLM-based) bunsetsu and LUWs parser for the Heian period (AD 794-1185) and evaluating its performances.We model the parser as a transformerbased joint sequential labels model that combines the bunsetsu BI tag, LUW BI tag, and LUW Part-of-Speech (POS) tag for each SUW token.We trained our models on the corpora of each period including contemporary and historical Japanese.The results ranged from 0.976 to 0.996 in the f1 value for both bunsetsu and LUW reconstruction indicating that our models achieved comparable performance with models for a contemporary Japanese corpus.Through statistical analysis and a diachronic case study, it was found that the estimation of bunsetsu could be influenced by the grammaticalization of morphemes.

Record transparency

Publication details

DOI
10.18653/v1/2024.ml4al-1.6
OpenAlex
W4402670241
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.