Researcher profile

Hua Wu

22 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. SgSum: Transforming Multi-document Summarization into Sub-graph Selection

    2021 · arXiv (Cornell University)

    Most of existing extractive multi-document summarization (MDS) methods score each sentence individually and extract salient sentences one by one to compose a summary, which have two main drawbacks: (1) neglecting both the intra and cross-document …

  2. Semi-Supervised Learning for Neural Machine Translation

    2016 · arXiv (Cornell University)

    While end-to-end neural machine translation (NMT) has made remarkable progress recently, NMT systems only rely on parallel corpora for parameter estimation. Since parallel corpora are usually limited in quantity, quality, and coverage, especially for low-resource …

  3. Multi-Task Learning for Multiple Language Translation

    2015

    Daxiang Dong, Hua Wu, Wei He, Dianhai Yu, Haifeng Wang. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long …

  4. Learning to Respond with Deep Neural Networks for Retrieval-Based Human-Computer Conversation System

    2016

    To establish an automatic conversation system between humans and computers is regarded as one of the most hardcore problems in computer science, which involves interdisciplinary techniques in information retrieval, natural language processing, artificial intelligence, etc. …

  5. Improved Neural Machine Translation with SMT Features

    2016 · Proceedings of the AAAI Conference on Artificial Intelligence

    Neural machine translation (NMT) conducts end-to-end translation with a source language encoder and a target language decoder, making promising translation performance. However, as a newly emerged approach, the method has some limitations. An NMT system …

  6. An End-to-End Model for Question Answering over Knowledge Base with Cross-Attention Combining Global Knowledge

    2017

    With the rapid growth of knowledge bases (KBs) on the web, how to take full advantage of them becomes increasingly important. Question answering over knowledge base (KB-QA) is one of the promising approaches to access …

  7. Multi-Turn Response Selection for Chatbots with Deep Attention Matching Network

    2018

    Xiangyang Zhou, Lu Li, Daxiang Dong, Yi Liu, Ying Chen, Wayne Xin Zhao, Dianhai Yu, Hua Wu. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018.

  8. ERNIE: Enhanced Representation through Knowledge Integration

    2019 · arXiv (Cornell University)

    We present a novel language representation model enhanced by knowledge called ERNIE (Enhanced Representation through kNowledge IntEgration). Inspired by the masking strategy of BERT, ERNIE is designed to learn language representation enhanced by knowledge masking …

  9. Proactive Human-Machine Conversation with Explicit Conversation Goal

    2019

    Though great progress has been made for human-machine conversation, current dialogue system is still in its infancy: it usually converses passively and utters words more as a matter of response, rather than on its own …

  10. Minimum Risk Training for Neural Machine Translation

    2016

    We propose minimum risk training for end-to-end neural machine translation. Unlike conventional maximum likelihood estimation, minimum risk training is capable of optimizing model parameters directly with respect to arbitrary evaluation metrics, which are not necessarily …

  11. Multi-Passage Machine Reading Comprehension with Cross-Passage Answer Verification

    2018

    Yizhong Wang, Kai Liu, Jing Liu, Wei He, Yajuan Lyu, Hua Wu, Sujian Li, Haifeng Wang. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018.

  12. DuReader: a Chinese Machine Reading Comprehension Dataset from Real-world Applications

    2018

    Wei He, Kai Liu, Jing Liu, Yajuan Lyu, Shiqi Zhao, Xinyan Xiao, Yuan Liu, Yizhong Wang, Hua Wu, Qiaoqiao She, Xuan Liu, Tian Wu, Haifeng Wang. Proceedings of the Workshop on Machine Reading for Question …

  13. Learning to Select Knowledge for Response Generation in Dialog Systems

    2019

    End-to-end neural models for intelligent dialogue systems suffer from the problem of generating uninformative responses. Various methods were proposed to generate more informative responses by leveraging external knowledge. However, few previous work has focused on …

  14. ERNIE 2.0: A Continual Pre-training Framework for Language Understanding

    2019 · arXiv (Cornell University)

    Recently, pre-trained models have achieved state-of-the-art results in various language understanding tasks, which indicates that pre-training on large-scale corpora may play a crucial role in natural language processing. Current pre-training procedures usually focus on training …

  15. Knowledge Aware Conversation Generation with Explainable Reasoning over Augmented Graphs

    2019

    Zhibin Liu, Zheng-Yu Niu, Hua Wu, Haifeng Wang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.

  16. End-to-End Speech Translation with Knowledge Distillation

    2019

    End-to-end speech translation (ST), which directly translates from source language speech into target language text, has attracted intensive attentions in recent years.Compared to conventional pipepine systems, end-to-end ST models have advantages of lower latency, smaller …

  17. ERNIE 2.0: A Continual Pre-Training Framework for Language Understanding

    2020 · Proceedings of the AAAI Conference on Artificial Intelligence

    Recently pre-trained models have achieved state-of-the-art results in various language understanding tasks. Current pre-training procedures usually focus on training the model with several simple tasks to grasp the co-occurrence of words or sentences. However, besides …

  18. ERNIE-M: Enhanced Multilingual Representation by Aligning Cross-lingual Semantics with Monolingual Corpora

    2021 · Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing

    Recent studies have demonstrated that pretrained cross-lingual models achieve impressive performance in downstream cross-lingual tasks. This improvement benefits from learning a large amount of monolingual and parallel corpora. Although it is generally acknowledged that parallel …

  19. RocketQA: An Optimized Training Approach to Dense Passage Retrieval for Open-Domain Question Answering

    2021

    Yingqi Qu, Yuchen Ding, Jing Liu, Kai Liu, Ruiyang Ren, Wayne Xin Zhao, Daxiang Dong, Hua Wu, Haifeng Wang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: …

  20. ERNIE 3.0: Large-scale Knowledge Enhanced Pre-training for Language Understanding and Generation

    2021 · arXiv (Cornell University)

    Pre-trained models have achieved state-of-the-art results in various Natural Language Processing (NLP) tasks. Recent works such as T5 and GPT-3 have shown that scaling up pre-trained language models can improve their generalization abilities. Particularly, the …

  21. RocketQAv2: A Joint Training Method for Dense Passage Retrieval and Passage Re-ranking

    2021 · Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing

    In various natural language processing tasks, passage retrieval and passage re-ranking are two key procedures in finding and ranking relevant information. Since both the two procedures contribute to the final performance, it is important to …

  22. Unified Structure Generation for Universal Information Extraction

    2022 · Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)

    Yaojie Lu, Qing Liu, Dai Dai, Xinyan Xiao, Hongyu Lin, Xianpei Han, Le Sun, Hua Wu. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.