Marco Cognetta
3 papers in the PaperMetrix corpus
Papers by this author
-
Online Infix Probability Computation for Probabilistic Finite Automata
2019
Probabilistic finite automata (PFAs) are common statistical language model in natural language and speech processing. A typical task for PFAs is to compute the probability of all strings that match a query pattern. An important …
-
SoftRegex: Generating Regex from Natural Language Descriptions using Softened Regex Equivalence
2019
Jun-U Park, Sang-Ki Ko, Marco Cognetta, Yo-Sub Han. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
-
Distributional Properties of Subword Regularization
2024 · arXiv (Cornell University)
Subword regularization, used widely in NLP, improves model performance by reducing the dependency on exact tokenizations, augmenting the training corpus, and exposing the model to more unique contexts during training. BPE and MaxMatch, two popular …