Researcher profile

Angela Fan

16 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Prior matters: simple and general methods for evaluating and improving topic quality in topic modeling

    2017 · arXiv (Cornell University)

    Latent Dirichlet Allocation (LDA) models trained without stopword removal often produce topics with high posterior probabilities on uninformative words, obscuring the underlying corpus content. Even when canonical stopwords are manually removed, uninformative words common in …

  2. A Few Thousand Translations Go a Long Way! Leveraging Pre-trained Models for African News Translation

    2022 · Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies

    David Adelani, Jesujoba Alabi, Angela Fan, Julia Kreutzer, Xiaoyu Shen, Machel Reid, Dana Ruiter, Dietrich Klakow, Peter Nabende, Ernie Chang, Tajuddeen Gwadabe, Freshia Sackey, Bonaventure F. P. Dossou, Chris Emezue, Colin Leong, Michael Beukman, Shamsuddeen …

  3. Language Modeling with Gated Convolutional Networks

    2016 · arXiv (Cornell University)

    The pre-dominant approach to language modeling to date is based on recurrent neural networks. Their success on this task is often linked to their ability to capture unbounded context. In this paper we develop a …

  4. Hierarchical Neural Story Generation

    2018 · arXiv (Cornell University)

    We explore story generation: creative systems that can build coherent and fluent passages of text about a topic. We collect a large dataset of 300K human-written stories paired with writing prompts from an online forum. …

  5. Wizard of Wikipedia: Knowledge-Powered Conversational agents

    2018 · arXiv (Cornell University)

    In open-domain dialogue intelligent agents should exhibit the use of knowledge, however there are few convincing demonstrations of this to date. The most popular sequence to sequence models typically "generate and hope" generic utterances that …

  6. Pay Less Attention with Lightweight and Dynamic Convolutions

    2019 · arXiv (Cornell University)

    Self-attention is a useful mechanism to build generative models for language and images. It determines the importance of context elements by comparing each element to the current time step. In this paper, we show that …

  7. fairseq: A Fast, Extensible Toolkit for Sequence Modeling

    2019

    Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, Michael Auli. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (Demonstrations). 2019.

  8. ELI5: Long Form Question Answering

    2019

    We introduce the first large-scale corpus for long-form question answering, a task requiring elaborate and in-depth answers to openended questions. The dataset comprises 270K threads from the Reddit forum "Explain Like I'm Five" (ELI5) where …

  9. Controllable Abstractive Summarization

    2018

    Current models for document summarization disregard user preferences such as the desired length, style, the entities that the user might be interested in, or how much of the document the user has already read. We …

  10. Language modeling with gated convolutional networks

    2017 · International Conference on Machine Learning

    The pre-dominant approach to language modeling to date is based on recurrent neural networks. Their success on this task is often linked to their ability to capture unbounded context. In this paper we develop a …

  11. Reducing Transformer Depth on Demand with Structured Dropout

    2019 · arXiv (Cornell University)

    Overparameterized transformer networks have obtained state of the art results in various natural language processing tasks, such as machine translation, language modeling, and question answering. These models contain hundreds of millions of parameters, necessitating a …

  12. Reducing Transformer Depth on Demand with Structured Dropout

    2020 · arXiv (Cornell University)

    Overparametrized transformer networks have obtained state of the art results in various natural language processing tasks, such as machine translation, language modeling, and question answering. These models contain hundreds of millions of parameters, necessitating a …

  13. Beyond English-Centric Multilingual Machine Translation

    2020 · arXiv (Cornell University)

    Existing work in translation demonstrated the potential of massively multilingual machine translation by training a single model able to translate between any pair of languages. However, much of this work is English-Centric by training only …

  14. The <scp>Flores-101</scp> Evaluation Benchmark for Low-Resource and Multilingual Machine Translation

    2022 · Transactions of the Association for Computational Linguistics

    Abstract One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the lack of good evaluation benchmarks. Current evaluation benchmarks either lack good coverage of low-resource languages, consider only restricted domains, …

  15. No Language Left Behind: Scaling Human-Centered Machine Translation

    2022 · arXiv (Cornell University)

    Driven by the goal of eradicating language barriers on a global scale, machine translation has solidified itself as a key focus of artificial intelligence research today. However, such efforts have coalesced around a small subset …

  16. Llama 2: Open Foundation and Fine-Tuned Chat Models

    2023 · arXiv (Cornell University)

    In this work, we develop and release Llama 2, a collection of pretrained and fine-tuned large language models (LLMs) ranging in scale from 7 billion to 70 billion parameters. Our fine-tuned LLMs, called Llama 2-Chat, …