Noah A. Smith
43 papers in the PaperMetrix corpus
Papers by this author
-
Frame-Semantic Role Labeling with Heterogeneous Annotations
2015
Meghana Kshirsagar, Sam Thomson, Nathan Schneider, Jaime Carbonell, Noah A. Smith, Chris Dyer. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing …
-
Segmental Recurrent Neural Networks for End-to-end Speech Recognition
2016 · arXiv (Cornell University)
We study the segmental recurrent neural network for end-to-end acoustic modelling. This model connects the segmental conditional random field (CRF) with a recurrent neural network (RNN) used for feature extraction. Compared to most previous CRF-based …
-
Low-Resource Parsing with Crosslingual Contextualized Representations
2019
Despite advances in dependency parsing, languages with small treebanks still present challenges.
-
Improving Transformer Models by Reordering their Sublayers
2020
Multilayer transformer networks consist of interleaved self-attention and feedforward sublayers. Could ordering the sublayers in a different pattern lead to better performance? We generate randomly ordered transformers and train them with the language modeling objective. …
-
Plug and Play Autoencoders for Conditional Text Generation
2020
Text autoencoders are commonly used for conditional generation tasks such as style transfer.We propose methods which are plug and play, where any pretrained autoencoder can be used, and only require learning a mapping within the …
-
Deep Encoder, Shallow Decoder: Reevaluating Non-autoregressive Machine Translation
2020 · arXiv (Cornell University)
Much recent effort has been invested in non-autoregressive neural machine translation, which appears to be an efficient alternative to state-of-the-art autoregressive machine translation on modern GPUs. In contrast to the latter, where generation is sequential, …
-
Scarecrow: A Framework for Scrutinizing Machine Text.
2021 · arXiv (Cornell University)
Modern neural text generation systems can produce remarkably fluent and grammatical texts. While earlier language models suffered from repetition and syntactic errors, the errors made by contemporary models are often semantic, narrative, or discourse failures. …
-
Expected Validation Performance and Estimation of a Random Variable's Maximum
2021 · arXiv (Cornell University)
Research in NLP is often supported by experimental results, and improved reporting of such results can lead to better understanding and more reproducible science. In this paper we analyze three statistical estimators for expected validation …
-
Self-Instruct: Aligning Language Models with Self-Generated Instructions
2022 · arXiv (Cornell University)
Large "instruction-tuned" language models (i.e., finetuned to respond to instructions) have demonstrated a remarkable ability to generalize zero-shot to new tasks. Nevertheless, they depend heavily on human-written instruction data that is often limited in quantity, …
-
Modeling Context With Linear Attention for Scalable Document-Level Translation
2022
Document-level machine translation leverages inter-sentence dependencies to produce more coherent and consistent translations. However, these models, predominantly based on transformers, are difficult to scale to long documents as their attention layers have quadratic complexity in …
-
One Embedder, Any Task: Instruction-Finetuned Text Embeddings
2023
Hongjin Su, Weijia Shi, Jungo Kasai, Yizhong Wang, Yushi Hu, Mari Ostendorf, Wen-tau Yih, Noah A. Smith, Luke Zettlemoyer, Tao Yu. Findings of the Association for Computational Linguistics: ACL 2023. 2023.
-
Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2
2023 · arXiv (Cornell University)
Since the release of TÜLU [Wang et al., 2023b], open resources for instruction tuning have developed quickly, from better base models to new finetuning techniques. We test and incorporate a number of these advances into …
-
Time is Encoded in the Weights of Finetuned Language Models
2024
We present time vectors, a simple tool to customize language models to new time periods.Time vectors are created by finetuning a language model on data from a single time (e.g., a year or month), and …
-
2 OLMo 2 Furious
2024 · arXiv (Cornell University)
We present OLMo 2, the next generation of our fully open language models. OLMo 2 includes a family of dense autoregressive language models at 7B, 13B and 32B scales with fully released artifacts -- model …
-
OLMoTrace: Tracing Language Model Outputs Back to Trillions of Training Tokens
2025 · arXiv (Cornell University)
We present OLMoTrace, the first system that traces the outputs of language models back to their full, multi-trillion-token training data in real time. OLMoTrace finds and shows verbatim matches between segments of language model output …
-
Retrofitting Word Vectors to Semantic Lexicons
2015
Manaal Faruqui, Jesse Dodge, Sujay Kumar Jauhar, Chris Dyer, Eduard Hovy, Noah A. Smith. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
-
Tree Edit Models for Recognizing Textual Entailments, Paraphrases, and Answers to Questions
2018 · KiltHub Repository
We describe tree edit models for representing sequences of tree transformations involving complex reordering phenomena and demonstrate that they offer a simple, intuitive, and effective method for modeling pairs of semantically related sentences. To efficiently …
-
Good Question! Statistical Ranking for Question Generation
2018 · Figshare
We address the challenge of automatically generating questions from reading materials for educational practice and assessment. Our approach is to overgenerate questions, then rank them. We use manually written rules to perform a sequence of …
-
Improved Transition-based Parsing by Modeling Characters instead of Words with LSTMs
2015 · RECERCAT (Consorci de Serveis Universitaris de Catalunya)
We present extensions to a continuousstate dependency parsing method that makes it applicable to morphologically rich languages. Starting with a highperformance transition-based parser that uses long short-term memory (LSTM) recurrent neural networks to learn representations …
-
Turning on the Turbo: Fast Third-Order Non-Projective Turbo Parsers
2018 · INDIGO (University of Illinois at Chicago)
We present fast, accurate, direct nonprojective dependency parsers with thirdorder features. Our approach uses AD3 , an accelerated dual decomposition algorithm which we extend to handle specialized head automata and sequential head bigram models. Experiments …
-
Transition-Based Dependency Parsing with Stack Long Short-Term Memory
2015 · RECERCAT (Consorci de Serveis Universitaris de Catalunya)
Chris Dyer, Miguel Ballesteros, Wang Ling, Austin Matthews, Noah A. Smith. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: …
-
Better Hypothesis Testing for Statistical Machine Translation: Controlling for Optimizer Instability
2018 · Figshare
In statistical machine translation, a researcher seeks to determine whether some innovation (e.g., a new feature, model, or inference algorithm) improves translation quality in comparison to a baseline system. To answer this question, he runs …
-
A Simple, Fast, and Effective Reparameterization of IBM Model 2
2018 · KiltHub Repository
We present a simple log-linear reparameterization of IBM Model 2 that overcomes problems arising from Model 1’s strong assumptions and Model 2’s overparameterization. Efficient inference, likelihood evaluation, and parameter estimation algorithms are provided. Training the …
-
Part-of-Speech Tagging for Twitter: Annotation, Features, and Experiments
2018 · Figshare
We address the problem of part-of-speech tagging for English data from the popular microblogging service Twitter. We develop a tagset, annotate data, develop features, and report tagging results nearing 90% accuracy. The data and tools …
-
Extractive Summarization by Maximizing Semantic Volume
2015
The most successful approaches to extractive text summarization seek to maximize bigram coverage subject to a budget constraint. In this work, we propose instead to maximize semantic volume. We embed each sentence in a semantic …
-
The Media Frames Corpus: Annotations of Frames Across Issues
2015
Dallas Card, Amber E. Boydstun, Justin H. Gross, Philip Resnik, Noah A. Smith. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing …
-
Massively Multilingual Word Embeddings
2016 · arXiv (Cornell University)
We introduce new methods for estimating and evaluating embeddings of words in more than fifty languages in a single shared embedding space. Our estimation methods, multiCluster and multiCCA, use dictionaries and monolingual data; they do …
-
Segmental Recurrent Neural Networks
2015 · arXiv (Cornell University)
We introduce segmental recurrent neural networks (SRNNs) which define, given an input sequence, a joint probability distribution over segmentations of the input and labelings of the segments. Representations of the input segments (i.e., contiguous subsequences …
-
Training with Exploration Improves a Greedy Stack LSTM Parser
2016
We adapt the greedy Stack-LSTM dependency parser of Dyer et al. (2015) to support a training-with-exploration procedure using dynamic oracles(Goldberg and Nivre, 2013) instead of cross-entropy minimization. This form of training, which accounts for model …
-
Generation from Abstract Meaning Representation using Tree Transducers
2016
Jeffrey Flanigan, Chris Dyer, Noah A. Smith, Jaime Carbonell. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
-
What Do Recurrent Neural Network Grammars Learn About Syntax?
2017
Adhiguna Kuncoro, Miguel Ballesteros, Lingpeng Kong, Chris Dyer, Graham Neubig, Noah A. Smith. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017.
-
Annotation Artifacts in Natural Language Inference Data
2018
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, Noah A. Smith. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short …
-
Neural Text Generation in Stories Using Entity Representations as Context
2018
Elizabeth Clark, Yangfeng Ji, Noah A. Smith. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
-
Linguistic Knowledge and Transferability of Contextual Representations
2019
Nelson F. Liu, Matt Gardner, Yonatan Belinkov, Matthew E. Peters, Noah A. Smith. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long …
-
Segmental Recurrent Neural Networks
2016 · International Conference on Learning Representations
Abstract: We introduce segmental recurrent neural networks (SRNNs) which define, given an input sequence, a joint probability distribution over segmentations of the input and labelings of the segments. Representations of the input segments (i.e., contiguous …
-
ATOMIC: An Atlas of Machine Commonsense for If-Then Reasoning
2019
We present ATOMIC, an atlas of everyday commonsense reasoning, organized through 877k textual descriptions of inferential knowledge. Compared to existing resources that center around taxonomic knowledge, ATOMIC focuses on inferential knowledge organized as typed if-then …
-
Show Your Work: Improved Reporting of Experimental Results
2019
Jesse Dodge, Suchin Gururangan, Dallas Card, Roy Schwartz, Noah A. Smith. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
-
Knowledge Enhanced Contextual Word Representations
2019
Matthew E. Peters, Mark Neumann, Robert Logan, Roy Schwartz, Vidur Joshi, Sameer Singh, Noah A. Smith. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on …
-
Fine-Tuning Pretrained Language Models: Weight Initializations, Data Orders, and Early Stopping
2020 · arXiv (Cornell University)
Fine-tuning pretrained contextual word embedding models to supervised downstream tasks has become commonplace in natural language processing. This process, however, is often brittle: even with the same hyperparameter values, distinct random seeds can lead to …
-
Improved Part-of-Speech Tagging for Online Conversational Text with Word Clusters
2018 · Figshare
We consider the problem of part-of-speech tagging for informal, online conversational text. We systematically evaluate the use of large-scale unsupervised word clustering and new lexical features to improve tagging accuracy. With these features, our system …
-
Self-Instruct: Aligning Language Models with Self-Generated Instructions
2023
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, Hannaneh Hajishirzi. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
-
UnifiedSKG: Unifying and Multi-Tasking Structured Knowledge Grounding with Text-to-Text Language Models
2022
Tianbao Xie, Chen Henry Wu, Peng Shi, Ruiqi Zhong, Torsten Scholak, Michihiro Yasunaga, Chien-Sheng Wu, Ming Zhong, Pengcheng Yin, Sida I. Wang, Victor Zhong, Bailin Wang, Chengzu Li, Connor Boyle, Ansong Ni, Ziyu Yao, Dragomir …
-
Demystifying Prompts in Language Models via Perplexity Estimation
2023
Language models can be prompted to perform a wide variety of tasks with zero- and few-shot in-context learning. However, performance varies significantly with the choice of prompt, and we do not yet understand why this …