Pascale Fung
14 ورقة في مجموعة PaperMetrix
أوراق هذا المؤلف
-
One-step and Two-step Classification for Abusive Language Detection on Twitter
2017
Automatic abusive language detection is a difficult but important task for online social media. Our research explores a twostep approach of performing classification on abusive language and then classifying into specific types and compares it …
-
A Novel Repetition Normalized Adversarial Reward for Headline Generation
2019
While reinforcement learning can effectively improve language generation models, it often suffers from generating incoherent and repetitive phrases [1]. In this paper, we propose a novel repetition normalized adversarial reward to mitigate these problems. Our …
-
Hierarchical Meta-Embeddings for Code-Switching Named Entity Recognition
2019
Genta Indra Winata, Zhaojiang Lin, Jamin Shin, Zihan Liu, Pascale Fung. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
-
End-to-End Question Answering Models for Goal-Oriented Dialog Learning
2019
The task of Next Utterance Classification in dialog learning highly resembles that of Question Answering, but there has not been much attention to applying models across the two fields, especially not in more practical dialog …
-
MinTL: Minimalist Transfer Learning for Task-Oriented Dialogue Systems
2020 · arXiv (Cornell University)
In this paper, we propose Minimalist Transfer Learning (MinTL) to simplify the system design process of task-oriented dialogue systems and alleviate the over-dependency on annotated data. MinTL is a simple yet effective transfer learning framework, …
-
Plug-and-Play Conversational Models
2020
There has been considerable progress made towards conversational models that generate coherent and fluent responses; however, this often involves training large language models on large dialogue datasets, such as Reddit. These large conversational models provide …
-
NusaCrowd: Open Source Initiative for Indonesian NLP Resources
2022 · arXiv (Cornell University)
We present NusaCrowd, a collaborative initiative to collect and unify existing resources for Indonesian languages, including opening access to previously non-public resources. Through this initiative, we have brought together 137 datasets and 118 standardized data …
-
Which One Are You Referring To? Multimodal Object Identification in Situated Dialogue
2023 · arXiv (Cornell University)
The demand for multimodal dialogue systems has been rising in various domains, emphasizing the importance of interpreting multimodal inputs from conversational and situational contexts. We explore three methods to tackle this problem and evaluate them …
-
Improving Query-Focused Meeting Summarization with Query-Relevant Knowledge
2023
Query-Focused Meeting Summarization (QFMS) aims to generate a summary of a given meeting transcript conditioned upon a query. The main challenges for QFMS are the long input text length and sparse query-relevant information in the …
-
Mem2Seq: Effectively Incorporating Knowledge Bases into End-to-End Task-Oriented Dialog Systems
2018
End-to-end task-oriented dialog systems usually suffer from the challenge of incorporating knowledge bases. In this paper, we propose a novel yet simple end-toend differentiable model called memoryto-sequence (Mem2Seq) to address this issue. Mem2Seq is the …
-
Learning Multilingual Meta-Embeddings for Code-Switching Named Entity Recognition
2019
In this paper, we propose Multilingual Meta-Embeddings (MME), an effective method to learn multilingual representations by leveraging monolingual pre-trained embeddings. MME learns to utilize information from these embeddings via a self-attention mechanism without explicit language …
-
Exploring Versatile Generative Language Model Via Parameter-Efficient Transfer Learning
2020
Fine-tuning pre-trained generative language models to down-stream language generation tasks has shown promising results. However, this comes with the cost of having a single, large model for each task, which is not ideal in low-memory/power …
-
Survey of Hallucination in Natural Language Generation
2022 · ACM Computing Surveys
Natural Language Generation (NLG) has improved exponentially in recent years thanks to the development of sequence-to-sequence deep learning technologies such as Transformer-based language models. This advancement has led to more fluent and coherent NLG, leading …
-
A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity
2023 · arXiv (Cornell University)
This paper proposes a framework for quantitatively evaluating interactive LLMs such as ChatGPT using publicly available data sets. We carry out an extensive technical evaluation of ChatGPT using 23 data sets covering 8 different common …