conference-paper Open access

What If Sentence-hood is Hard to Define: A Case Study in Chinese Reading Comprehension

Research footprint

At a glance

Citations
2
References
53
Comments
0
Paper overview

Abstract

Machine reading comprehension (MRC) is a challenging NLP task for it requires to carefully deal with all linguistic granularities from word, sentence to passage. For extractive MRC, the answer span has been shown mostly determined by key evidence linguistic units, in which it is a sentence in most cases. However, we recently discovered that sentences may not be clearly defined in many languages to different extents, so that this causes so-called location unit ambiguity problem and as a result makes it difficult for the model to determine which sentence exactly contains the answer span when sentence itself has not been clearly defined at all. Taking Chinese language as a case study, we explain and analyze such a linguistic phenomenon and correspondingly propose a reader with Explicit Span-Sentence Predication to alleviate such a problem. Our proposed reader eventually helps achieve new a state-of-the-art on Chinese MRC benchmark and shows great potential in dealing with other languages.

Record transparency

Publication details

DOI
10.18653/v1/2021.findings-emnlp.202
OpenAlex
W3212498084
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.