Researcher profile

Mark Dredze

10 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Harmonic Grammar, Optimality Theory, and Syntax Learnability: An Empirical Exploration of Czech Word Order

    2017 · arXiv (Cornell University)

    This work presents a systematic theoretical and empirical comparison of the major algorithms that have been proposed for learning Harmonic and Optimality Theory grammars (HG and OT, respectively). By comparing learning algorithms, we are also …

  2. Are All Languages Created Equal in Multilingual BERT?

    2020

    Multilingual BERT (mBERT) However, these evaluations have focused on cross-lingual transfer with highresource languages, covering only a third of the languages covered by mBERT. We explore how mBERT performs on a much wider set of …

  3. Demographic Representation and Collective Storytelling in the Me Too Twitter Hashtag Activism Movement

    2021 · Proceedings of the ACM on Human-Computer Interaction

    The #MeToo movement on Twitter has drawn attention to the pervasive nature of sexual harassment and violence. While #MeToo has been praised for providing support for self-disclosures of harassment or violence and shifting societal response, …

  4. Improving Zero-Shot Multi-Lingual Entity Linking

    2021 · arXiv (Cornell University)

    Entity linking -- the task of identifying references in free text to relevant knowledge base representations -- often focuses on single languages. We consider multilingual entity linking, where a single model is trained to link …

  5. Updated Headline Generation: Creating Updated Summaries for Evolving News Stories

    2022 · Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)

    We propose the task of updated headline generation, in which a system generates a headline for an updated article, considering both the previous article and headline. The system must identify the novel information in the …

  6. Are Clinical T5 Models Better for Clinical Text?

    2024 · arXiv (Cornell University)

    Large language models with a transformer-based encoder/decoder architecture, such as T5, have become standard platforms for supervised tasks. To bring these technologies to the clinical domain, recent work has trained new or adapted existing models …

  7. Named Entity Recognition for Chinese Social Media with Jointly Trained Embeddings

    2015

    We consider the task of named entity recognition for Chinese social media. The long line of work in Chinese NER has fo-cused on formal domains, and NER for social media has been largely restricted to …

  8. Improving Named Entity Recognition for Chinese Social Media with Word Segmentation Representation Learning

    2016

    Named entity recognition, and other information extraction tasks, frequently use linguistic features such as part of speech tags or chunkings. For languages where word boundaries are not readily identified in text, word segmentation is a …

  9. Beto, Bentz, Becas: The Surprising Cross-Lingual Effectiveness of BERT

    2019

    Shijie Wu, Mark Dredze. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.

  10. BloombergGPT: A Large Language Model for Finance

    2023 · arXiv (Cornell University)

    The use of NLP in the realm of financial technology is broad and complex, with applications ranging from sentiment analysis and named entity recognition to question answering. Large Language Models (LLMs) have been shown to …