Jiawei Han
15 papers in the PaperMetrix corpus
Papers by this author
-
Cross-type Biomedical Named Entity Recognition with Deep Multi-Task Learning
2018 · arXiv (Cornell University)
Motivation: State-of-the-art biomedical named entity recognition (BioNER) systems often require handcrafted features specific to each entity type, such as genes, chemicals and diseases. Although recent studies explored using neural network models for BioNER to free …
-
AspEm: Embedding Learning by Aspects in Heterogeneous Information Networks
2018 · arXiv (Cornell University)
Heterogeneous information networks (HINs) are ubiquitous in real-world applications. Due to the heterogeneity in HINs, the typed edges may not fully align with each other. In order to capture the semantic subtlety, we propose the …
-
Entropy-Based Subword Mining with an Application to Word Embeddings
2018
Recent literature has shown a wide variety of benefits to mapping traditional onehot representations of words and phrases to lower-dimensional real-valued vectors known as word embeddings. Traditionally, most word embedding algorithms treat each word as …
-
Learning Named Entity Tagger using Domain-Specific Dictionary
2018 · arXiv (Cornell University)
Recent advances in deep neural models allow us to build reliable named entity recognition (NER) systems without handcrafting features. However, such methods require large amounts of manually-labeled training data. There have been efforts on replacing …
-
Generating Representative Headlines for News Stories
2020 · arXiv (Cornell University)
Millions of news articles are published online every day, which can be overwhelming for readers to follow. Grouping articles that are reporting the same event into news stories is a common way of assisting readers …
-
Open-Domain Question Answering with Pre-Constructed Question Spaces
2020 · arXiv (Cornell University)
Open-domain question answering aims at solving the task of locating the answers to user-generated questions in massive collections of documents. There are two families of solutions available: retriever-readers, and knowledge-graph-based approaches. A retriever-reader usually first …
-
The Future is not One-dimensional: Complex Event Schema Induction by Graph Modeling for Event Prediction
2021 · arXiv (Cornell University)
Event schemas encode knowledge of stereotypical structures of events and their connections. As events unfold, schemas are crucial to act as a scaffolding. Previous work on event schema induction focuses either on atomic events or …
-
Structure-Augmented Reasoning Generation
2025 · arXiv (Cornell University)
Recent advances in Large Language Models (LLMs) have significantly improved complex reasoning capabilities. Retrieval-Augmented Generation (RAG) has further extended these capabilities by grounding generation in dynamically retrieved evidence, enabling access to information beyond the model's …
-
CoType
2017
Extracting entities and relations for types of interest from text is important for understanding massive text corpora. Traditionally, systems of entity relation extraction have relied on human-annotated corpora for training and adopted an incremental pipeline. …
-
Empower Sequence Labeling with Task-Aware Neural Language Model
2018 · Proceedings of the AAAI Conference on Artificial Intelligence
Linguistic sequence labeling is a general approach encompassing a variety of problems, such as part-of-speech tagging and named entity recognition. Recent advances in neural networks (NNs) make it possible to build reliable models without handcrafted …
-
Weakly-Supervised Neural Text Classification
2018
Deep neural networks are gaining increasing popularity for the classic text classification task, due to their strong expressive power and less requirement for feature engineering. Despite such attractiveness, neural text classification models suffer from the …
-
CoType: Joint Extraction of Typed Entities and Relations with Knowledge Bases
2016 · arXiv (Cornell University)
Extracting entities and relations for types of interest from text is important for understanding massive text corpora. Traditionally, systems of entity relation extraction have relied on human-annotated corpora for training and adopted an incremental pipeline. …
-
Empower Sequence Labeling with Task-Aware Neural Language Model
2017 · arXiv (Cornell University)
Linguistic sequence labeling is a general modeling approach that encompasses a variety of problems, such as part-of-speech tagging and named entity recognition. Recent advances in neural networks (NNs) make it possible to build reliable models …
-
Spherical Text Embedding
2019 · arXiv (Cornell University)
Unsupervised text embedding has shown great power in a wide range of NLP tasks. While text embeddings are typically learned in the Euclidean space, directional similarity is often more effective in tasks such as word …
-
Document-Level Event Argument Extraction by Conditional Generation
2021
Event extraction has long been treated as a sentence-level task in the IE community. We argue that this setting does not match human information seeking behavior and leads to incomplete and uninformative extraction results. We …