Jiwei Li
24 papers in the PaperMetrix corpus
Papers by this author
-
Glyce: Glyph-vectors for Chinese Character Representations
2019 · arXiv (Cornell University)
It is intuitive that NLP tasks for logographic languages like Chinese should benefit from the use of the glyph information in those languages. However, due to the lack of rich pictographic evidence in glyphs and …
-
Visualizing and Understanding Neural Models in NLP
2016
While neural networks have been successfully applied to many NLP tasks the resulting vectorbased models are very difficult to interpret. For example it's not clear how they achieve compositionality, building sentence meaning from the meanings …
-
A General Framework for Defending Against Backdoor Attacks via Influence Graph
2021 · arXiv (Cornell University)
In this work, we propose a new and general framework to defend against backdoor attacks, inspired by the fact that attack triggers usually follow a \textsc{specific} type of attacking pattern, and therefore, poisoned training examples …
-
Cascaded CNN-resBiLSTM-CTC: An End-to-End Acoustic Model For Speech\n Recognition
2018 · arXiv (Cornell University)
Automatic speech recognition (ASR) tasks are resolved by end-to-end deep\nlearning models, which benefits us by less preparation of raw data, and easier\ntransformation between languages. We propose a novel end-to-end deep learning\nmodel architecture namely cascaded CNN-resBiLSTM-CTC. …
-
GPT-NER: Named Entity Recognition via Large Language Models
2023 · arXiv (Cornell University)
Despite the fact that large-scale Language Models (LLM) have achieved SOTA performances on a variety of NLP tasks, its performance on NER is still significantly below supervised baselines. This is due to the gap between …
-
Pushing the Limits of ChatGPT on NLP Tasks
2023 · arXiv (Cornell University)
Despite the success of ChatGPT, its performances on most NLP tasks are still well below the supervised baselines. In this work, we looked into the causes, and discovered that its subpar performance was caused by …
-
Ranking-Enhanced Unsupervised Sentence Representation Learning
2023
Yeon Seonwoo, Guoyin Wang, Changmin Seo, Sajal Choudhary, Jiwei Li, Xiang Li, Puyang Xu, Sunghyun Park, Alice Oh. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
-
Are Human-generated Demonstrations Necessary for In-context Learning?
2023 · arXiv (Cornell University)
Despite the promising few-shot ability of large language models (LLMs), the standard paradigm of In-context Learning (ICL) suffers the disadvantages of susceptibility to selected demonstrations and the intricacy to generate these demonstrations. In this paper, …
-
DTI with Minimal Data: Image Translation Based Distortion Correction and FA Map Generation for Clinical Efficiency
2025 · Proceedings on CD-ROM - International Society for Magnetic Resonance in Medicine. Scientific Meeting and Exhibition/Proceedings of the International Society for Magnetic Resonance in Medicine, Scientific Meeting and Exhibition
Motivation: Current DTI processing requires around 30 directions per shell to ensure data quality, which is difficult to obtain in certain vulnerable patient groups. Goal(s): To reduce number of DWI directions needed in clinical study. …
-
Visualizing and Understanding Neural Models in NLP
2015 · arXiv (Cornell University)
While neural networks have been successfully applied to many NLP tasks the resulting vector-based models are very difficult to interpret. For example it's not clear how they achieve {\em compositionality}, building sentence meaning from the …
-
A Diversity-Promoting Objective Function for Neural Conversation Models
2015 · arXiv (Cornell University)
Sequence-to-sequence neural network models for generation of conversational responses tend to generate safe, commonplace responses (e.g., "I don't know") regardless of the input. We suggest that the traditional objective function, i.e., the likelihood of output …
-
A Hierarchical Neural Autoencoder for Paragraphs and Documents
2015 · arXiv (Cornell University)
Jiwei Li, Thang Luong, Dan Jurafsky. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
-
A Persona-Based Neural Conversation Model
2016 · arXiv (Cornell University)
We present persona-based models for handling the issue of speaker consistency in neural response generation. A speaker model encodes personas in distributed embeddings that capture individual characteristics such as background information and speaking style. A …
-
A Simple, Fast Diverse Decoding Algorithm for Neural Generation
2016 · arXiv (Cornell University)
In this paper, we propose a simple, fast decoding algorithm that fosters diversity in neural generation. The algorithm modifies the standard beam search algorithm by adding an inter-sibling ranking penalty, favoring choosing hypotheses from diverse …
-
Adversarial Learning for Neural Dialogue Generation
2017 · arXiv (Cornell University)
In this paper, drawing intuition from the Turing test, we propose using adversarial training for open-domain dialogue generation: the system is trained to produce sequences that are indistinguishable from human-generated dialogue utterances. We cast the …
-
Data Noising as Smoothing in Neural Network Language Models
2017 · arXiv (Cornell University)
Data noising is an effective technique for regularizing neural network models. While noising is widely adopted in application domains such as vision and speech, commonly used noising primitives have not been developed for discrete sequence-level …
-
Is Word Segmentation Necessary for Deep Learning of Chinese Representations?
2019
Segmenting a chunk of text into words is usually the first step of processing Chinese text, but its necessity has rarely been explored.
-
Entity-Relation Extraction as Multi-Turn Question Answering
2019
In this paper, we propose a new paradigm for the task of entity-relation extraction. We cast the task as a multi-turn question answering problem, i.e., the extraction of entities and relations is transformed to the …
-
When Are Tree Structures Necessary for Deep Learning of Representations?
2015 · arXiv (Cornell University)
Recursive neural models, which use syntactic parse trees to recursively generate representations bottom-up, are a popular architecture. But there have not been rigorous evaluations showing for exactly which tasks this syntax-based method is appropriate. In …
-
A Diversity-Promoting Objective Function for Neural Conversation Models
2016
Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, Bill Dolan. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
-
Neural Net Models of Open-domain Discourse Coherence
2017
Discourse coherence is strongly associated with text quality, making it important to natural language generation and understanding. Yet existing models of coherence focus on measuring individual aspects of coherence (lexical overlap, rhetorical structure, entity centering) …
-
Dice Loss for Data-imbalanced NLP Tasks
2020
Many NLP tasks such as tagging and machine reading comprehension (MRC) are faced with the severe data imbalance issue: negative examples significantly outnumber positive ones, and the huge number of easy-negative examples overwhelms training. The …
-
CorefQA: Coreference Resolution as Query-based Span Prediction
2020
In this paper, we present CorefQA, an accurate and extensible approach for the coreference resolution task. We formulate the problem as a span prediction task, like in question answering: A query is generated for each …
-
A Unified MRC Framework for Named Entity Recognition
2020
The task of named entity recognition (NER) is normally divided into nested NER and flat NER depending on whether named entities are nested or not. Models are usually separately developed for the two tasks, since …