Kentaro Inui
11 papers in the PaperMetrix corpus
Papers by this author
-
A Joint Neural Model for Fine-Grained Named Entity Classification of Wikipedia Articles
2017 · IEICE Transactions on Information and Systems
This paper addresses the task of assigning labels of fine-grained named entity (NE) types to Wikipedia articles. Information of NE types are useful when extracting knowledge of NEs from natural language text. It is common …
-
Distantly Supervised Biomedical Knowledge Acquisition via Knowledge Graph Based Attention
2019
The increased demand for structured scientific knowledge has attracted considerable attention in extracting scientific relation from the ever growing scientific publications. Distant supervision is widely applied approach to automatically generate large amounts of labelled data …
-
An Empirical Study of Incorporating Pseudo Data into Grammatical Error Correction
2019 · arXiv (Cornell University)
The incorporation of pseudo data in the training of grammatical error correction models has been one of the main factors in improving the performance of such models. However, consensus is lacking on experimental configurations, namely, …
-
Language Models as an Alternative Evaluator of Word Order Hypotheses: A Case Study in Japanese
2020
We examine a methodology using neural language models (LMs) for analyzing the word order of language. This LM-based method has the potential to overcome the difficulties existing methods face, such as the propagation of preprocessor …
-
A System for Worldwide COVID-19 Information Aggregation
2020 · arXiv (Cornell University)
The global pandemic of COVID-19 has made the public pay close attention to related news, covering various domains, such as sanitation, treatment, and effects on education. Meanwhile, the COVID-19 condition is very different among the …
-
A Self-Refinement Strategy for Noise Reduction in Grammatical Error Correction
2020 · arXiv (Cornell University)
Existing approaches for grammatical error correction (GEC) largely rely on supervised learning with manually created GEC datasets. However, there has been little focus on verifying and ensuring the quality of the datasets, and on how …
-
Corruption Is Not All Bad: Incorporating Discourse Structure into Pre-training via Corruption for Essay Scoring
2020 · arXiv (Cornell University)
Existing approaches for automated essay scoring and document representation learning typically rely on discourse parsers to incorporate discourse structure into text representation. However, the performance of parsers is not always adequate, especially when they are …
-
Test-time Augmentation for Factual Probing
2023
Factual probing is a method that uses prompts to test if a language model “knows” certain world knowledge facts. A problem in factual probing is that small changes to the prompt can lead to large …
-
Repetition Neurons: How Do Language Models Produce Repetitions?
2024 · arXiv (Cornell University)
This paper introduces repetition neurons, regarded as skill neurons responsible for the repetition problem in text generation tasks. These neurons are progressively activated more strongly as repetition continues, indicating that they perceive repetition as a …
-
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge
2025 · arXiv (Cornell University)
While fine-tuning is the standard for injecting factual knowledge into large language models (LLMs), the mechanisms enabling reliable fact recall via unseen queries remain poorly understood. Common two-stage training strategies, which sequentially train on fact …
-
Neural Architectures for Fine-grained Entity Type Classification
2017
Sonse Shimaoka, Pontus Stenetorp, Kentaro Inui, Sebastian Riedel. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017.