Researcher profile

Furu Wei

38 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Read + Verify: Machine Reading Comprehension with Unanswerable Questions

    2018 · arXiv (Cornell University)

    Machine reading comprehension with unanswerable questions aims to abstain from answering when no answer can be inferred. In addition to extract answers, previous works usually predict an additional "no-answer" probability to detect unanswerable cases. However, …

  2. Unified Language Model Pre-training for Natural Language Understanding and Generation

    2019 · arXiv (Cornell University)

    This paper presents a new Unified pre-trained Language Model (UniLM) that can be fine-tuned for both natural language understanding and generation tasks. The model is pre-trained using three types of language modeling tasks: unidirectional, bidirectional, …

  3. Automatic Grammatical Error Correction for Sequence-to-sequence Text Generation: An Empirical Study

    2019

    Sequence-to-sequence (seq2seq) models have achieved tremendous success in text generation tasks. However, there is no guarantee that they can always generate sentences without grammatical errors. In this paper, we present a preliminary empirical study on …

  4. At Which Level Should We Extract? An Empirical Study on Extractive Document Summarization

    2020 · arXiv (Cornell University)

    Extractive methods have been proven effective in automatic document summarization. Previous works perform this task by identifying informative contents at sentence level. However, it is unclear whether performing extraction at sentence level is the best …

  5. Investigating Learning Dynamics of BERT Fine-Tuning

    2020

    Yaru Hao, Li Dong, Furu Wei, Ke Xu. Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing. 2020.

  6. Improving Pretrained Cross-Lingual Language Models via Self-Labeled Word Alignment

    2021 · arXiv (Cornell University)

    The cross-lingual language models are typically pretrained with masked language modeling on multilingual text or parallel sentences. In this paper, we introduce denoising word alignment as a new cross-lingual pre-training task. Specifically, the model first …

  7. Controllable Natural Language Generation with Contrastive Prefixes

    2022 · Findings of the Association for Computational Linguistics: ACL 2022

    To guide the generation of large pretrained language models (LM), previous work has focused on directly fine-tuning the language model or utilizing an attribute discriminator. In this work, we propose a novel lightweight framework for …

  8. MoEC: Mixture of Expert Clusters

    2022 · arXiv (Cornell University)

    Sparsely Mixture of Experts (MoE) has received great interest due to its promising scaling capability with affordable computational overhead. MoE converts dense layers into sparse experts, and utilizes a gated routing network to make experts …

  9. Text Morphing

    2018 · arXiv (Cornell University)

    In this paper, we introduce a novel natural language generation task, termed as text morphing, which targets at generating the intermediate sentences that are fluency and smooth with the two input sentences. We propose the …

  10. Revamping Multilingual Agreement Bidirectionally via Switched Back-translation for Multilingual Neural Machine Translation

    2022 · arXiv (Cornell University)

    Despite the fact that multilingual agreement (MA) has shown its importance for multilingual neural machine translation (MNMT), current methodologies in the field have two shortages: (i) require parallel data between multiple language pairs, which is …

  11. Joint Pre-Training with Speech and Bilingual Text for Direct Speech to Speech Translation

    2022 · arXiv (Cornell University)

    Direct speech-to-speech translation (S2ST) is an attractive research topic with many advantages compared to cascaded S2ST. However, direct S2ST suffers from the data scarcity problem because the corpora from speech of the source language to …

  12. HanoiT: Enhancing Context-aware Translation via Selective Context

    2023 · arXiv (Cornell University)

    Context-aware neural machine translation aims to use the document-level context to improve translation quality. However, not all words in the context are helpful. The irrelevant or trivial words may bring some noise and distract the …

  13. K-Level Reasoning: Establishing Higher Order Beliefs in Large Language Models for Strategic Reasoning

    2024 · arXiv (Cornell University)

    Strategic reasoning is a complex yet essential capability for intelligent agents. It requires Large Language Model (LLM) agents to adapt their strategies dynamically in multi-agent environments. Unlike static reasoning tasks, success in these contexts depends …

  14. Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language Models

    2024 · arXiv (Cornell University)

    Large language models (LLMs) demonstrate remarkable multilingual capabilities without being pre-trained on specially curated multilingual parallel corpora. It remains a challenging problem to explain the underlying mechanisms by which LLMs process multilingual texts. In this …

  15. HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria Decomposition

    2024

    Yuxuan Liu, Tianchi Yang, Shaohan Huang, Zihan Zhang, Haizhen Huang, Furu Wei, Weiwei Deng, Feng Sun, Qi Zhang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.

  16. Language Models as Inductive Reasoners

    2024

    Zonglin Yang, Li Dong, Xinya Du, Hao Cheng, Erik Cambria, Xiaodong Liu, Jianfeng Gao, Furu Wei. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). …

  17. Question Answering over Freebase with Multi-Column Convolutional Neural Networks

    2015

    Li Dong, Furu Wei, Ming Zhou, Ke Xu. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.

  18. Selective Encoding for Abstractive Sentence Summarization

    2017

    We propose a selective encoding model to extend the sequence-to-sequence framework for abstractive sentence summarization. It consists of a sentence encoder, a selective gate network, and an attention equipped decoder. The sentence encoder and decoder …

  19. S-Net: From Answer Extraction to Answer Generation for Machine Reading Comprehension

    2017 · arXiv (Cornell University)

    In this paper, we present a novel approach to machine reading comprehension for the MS-MARCO dataset. Unlike the SQuAD dataset that aims to answer a question with exact text spans in a passage, the MS-MARCO …

  20. Gated Self-Matching Networks for Reading Comprehension and Question Answering

    2017

    In this paper, we present the gated selfmatching networks for reading comprehension style question answering, which aims to answer questions from a given passage. We first match the question and passage with gated attention-based recurrent …

  21. Faithful to the Original: Fact Aware Neural Abstractive Summarization

    2018 · Proceedings of the AAAI Conference on Artificial Intelligence

    Unlike extractive summarization, abstractive summarization has to fuse different parts of the source text, which inclines to create fake facts. Our preliminary study reveals nearly 30% of the outputs from a state-of-the-art neural summarization system …

  22. Hierarchical Attention Flow for Multiple-Choice Reading Comprehension

    2018 · Proceedings of the AAAI Conference on Artificial Intelligence

    In this paper, we focus on multiple-choice reading comprehension which aims to answer a question given a passage and multiple candidate options. We present the hierarchical attention flow to adequately leverage candidate options to model …

  23. Fluency Boost Learning and Inference for Neural Grammatical Error Correction

    2018

    Most of the neural sequence-to-sequence (seq2seq) models for grammatical error correction (GEC) have two limitations: (1) a seq2seq model may not be well generalized with only limited error-corrected data; (2) a seq2seq model may fail …

  24. Retrieve, Rerank and Rewrite: Soft Template Based Neural Summarization

    2018

    Most previous seq2seq summarization systems purely depend on the source text to generate summaries, which tends to work unstably. Inspired by the traditional template-based summarization approaches, this paper proposes to use existing summaries as soft …

  25. Reaching Human-level Performance in Automatic Grammatical Error Correction: An Empirical Study

    2018 · arXiv (Cornell University)

    Neural sequence-to-sequence (seq2seq) approaches have proven to be successful in grammatical error correction (GEC). Based on the seq2seq framework, we propose a novel fluency boost learning and inference mechanism. Fluency boosting learning generates diverse error-corrected …

  26. Neural Latent Extractive Document Summarization

    2018

    Extractive summarization models require sentence-level labels, which are usually created heuristically (e.g., with rule-based methods) given that most summarization datasets only have document-summary pairs. Since these labels might be suboptimal, we propose a latent variable …

  27. Neural Document Summarization by Jointly Learning to Score and Select Sentences

    2018

    Sentence scoring and sentence selection are two main steps in extractive document summarization systems. However, previous works treat them as two separated subtasks. In this paper, we present a novel end-to-end neural network framework for …

  28. Formality Style Transfer with Hybrid Textual Annotations

    2019 · arXiv (Cornell University)

    Formality style transformation is the task of modifying the formality of a given sentence without changing its content. Its challenge is the lack of large-scale sentence-aligned parallel data. In this paper, we propose an omnivorous …

  29. HIBERT: Document Level Pre-training of Hierarchical Bidirectional Transformers for Document Summarization

    2019

    Neural extractive summarization models usually employ a hierarchical encoder for document encoding and they are trained using sentence-level labels, which are created heuristically using rule-based methods. Training the hierarchical encoder with these inaccurate labels is …

  30. A Dependency-Based Neural Network for Relation Classification

    2015

    Yang Liu, Furu Wei, Sujian Li, Heng Ji, Ming Zhou, Houfeng Wang. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume …

  31. Attention-Guided Answer Distillation for Machine Reading Comprehension

    2018

    Despite that current reading comprehension systems have achieved significant advancements, their promising performances are often obtained at the cost of making an ensemble of numerous models. Besides, existing approaches are also vulnerable to adversarial attacks. …

  32. Response Generation by Context-Aware Prototype Editing

    2019 · Proceedings of the AAAI Conference on Artificial Intelligence

    Open domain response generation has achieved remarkable progress in recent years, but sometimes yields short and uninformative responses. We propose a new paradigm, prototypethen-edit for response generation, that first retrieves a prototype response from a …

  33. Read + Verify: Machine Reading Comprehension with Unanswerable Questions

    2019 · Proceedings of the AAAI Conference on Artificial Intelligence

    Machine reading comprehension with unanswerable questions aims to abstain from answering when no answer can be inferred. In addition to extract answers, previous works usually predict an additional “no-answer” probability to detect unanswerable cases. However, …

  34. Faithful to the Original: Fact Aware Neural Abstractive Summarization

    2017 · arXiv (Cornell University)

    Unlike extractive summarization, abstractive summarization has to fuse different parts of the source text, which inclines to create fake facts. Our preliminary study reveals nearly 30% of the outputs from a state-of-the-art neural summarization system …

  35. Visualizing and Understanding the Effectiveness of BERT

    2019

    Yaru Hao, Li Dong, Furu Wei, Ke Xu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.

  36. Text Embeddings by Weakly-Supervised Contrastive Pre-training

    2022 · arXiv (Cornell University)

    This paper presents E5, a family of state-of-the-art text embeddings that transfer well to a wide range of tasks. The model is trained in a contrastive manner with weak supervision signals from our curated large-scale …

  37. PromptBERT: Improving BERT Sentence Embeddings with Prompts

    2022

    Ting Jiang, Jian Jiao, Shaohan Huang, Zihan Zhang, Deqing Wang, Fuzhen Zhuang, Furu Wei, Haizhen Huang, Denvy Deng, Qi Zhang. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.

  38. Ranking with Recursive Neural Networks and Its Application to Multi-Document Summarization

    2015 · Proceedings of the AAAI Conference on Artificial Intelligence

    We develop a Ranking framework upon Recursive Neural Networks (R2N2) to rank sentences for multi-document summarization. It formulates the sentence ranking task as a hierarchical regression process, which simultaneously measures the salience of a sentence …