ملف الباحث

Vilém Zouhar

6 أوراق في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Fusing Sentence Embeddings Into LSTM-based Autoregressive Language Models

    2022 · arXiv (Cornell University)

    Although masked language models are highly performant and widely adopted by NLP practitioners, they can not be easily used for autoregressive language modelling (next word prediction and sequence probability estimation). We present an LSTM-based autoregressive …

  2. Sentence Ambiguity, Grammaticality and Complexity Probes

    2022 · arXiv (Cornell University)

    It is unclear whether, how and where large pre-trained language models capture subtle linguistic traits like ambiguity, grammaticality and sentence complexity. We present results of automatic classification of these traits and compare their viability and …

  3. Revisiting Automated Topic Model Evaluation with Large Language Models

    2023

    Topic models help us make sense of large text collections. Automatically evaluating their output and determining the optimal number of topics are both longstanding challenges, with no effective automated solutions to date. This paper proposes …

  4. Quality and Quantity of Machine Translation References for Automatic Metrics

    2024 · arXiv (Cornell University)

    Automatic machine translation metrics typically rely on human translations to determine the quality of system translations. Common wisdom in the field dictates that the human references should be of very high quality. However, there are …

  5. Error Span Annotation: A Balanced Approach for Human Evaluation of Machine Translation

    2024 · arXiv (Cornell University)

    High-quality Machine Translation (MT) evaluation relies heavily on human judgments. Comprehensive error classification methods, such as Multidimensional Quality Metrics (MQM), are expensive as they are time-consuming and can only be done by experts, whose availability …

  6. Distributional Properties of Subword Regularization

    2024 · arXiv (Cornell University)

    Subword regularization, used widely in NLP, improves model performance by reducing the dependency on exact tokenizations, augmenting the training corpus, and exposing the model to more unique contexts during training. BPE and MaxMatch, two popular …