Timothy Baldwin
17 papers in the PaperMetrix corpus
Papers by this author
-
A Probabilistic Rating Auto-encoder for Personalized Recommender Systems
2015
User profiling is a key component of personalized recommender systems, and is used to generate user profiles that describe individual user interests and preferences. The increasing availability of big data is driving the urgent need …
-
Shared Tasks of the 2015 Workshop on Noisy User-generated Text: Twitter Lexical Normalization and Named Entity Recognition
2015 · The Association for Computational Linguistics
This paper presents the results of the two shared tasks associated with W-NUT 2015: (1) a text normalization task with 10 participants; and (2) a named entity tagging task with 8 participants. We outline the …
-
Predicting Online Islamophobic Behavior after #ParisAttacks
2018 · The Journal of Web Science
The tragic Paris terrorist attacks of November 13, 2015 sparked a massive global discussion on Twitter and other social media, with millions of tweets in the first few hours after the attacks. Most of these …
-
A preliminary comparison of job, talent, and web search
2018 · CEUR Workshop Proceedings
Copyright held by the author(s). This paper presents an initial comparison of user behavior in job, talent, and web search using query and click logs from a popular employment marketplace. The observations suggest that the …
-
Balancing out Bias: Achieving Fairness Through Training Reweighting.
2021 · arXiv (Cornell University)
Bias in natural language processing arises primarily from models learning characteristics of the author such as gender and race when modelling tasks such as sentiment and syntactic parsing. This problem manifests as disparities in error …
-
One Country, 700+ Languages: NLP Challenges for Underrepresented Languages and Dialects in Indonesia
2022 · arXiv (Cornell University)
NLP research is impeded by a lack of resources and awareness of the challenges presented by underrepresented languages and dialects. Focusing on the languages spoken in Indonesia, the second most linguistically diverse and the fourth …
-
NusaCrowd: Open Source Initiative for Indonesian NLP Resources
2022 · arXiv (Cornell University)
We present NusaCrowd, a collaborative initiative to collect and unify existing resources for Indonesian languages, including opening access to previously non-public resources. Through this initiative, we have brought together 137 datasets and 118 standardized data …
-
LM-Polygraph: Uncertainty Estimation for Language Models
2023 · arXiv (Cornell University)
Recent advancements in the capabilities of large language models (LLMs) have paved the way for a myriad of groundbreaking applications in various fields. However, a significant challenge arises as these models often "hallucinate", i.e., fabricate …
-
Robustness Tests for Automatic Machine Translation Metrics with Adversarial Attacks
2023
We investigate MT evaluation metric performance on adversarially-synthesized texts, to shed light on metric robustness. We experiment with word- and character-level attacks on three popular machine translation metrics: BERTScore, BLEURT, and COMET. Our human experiments …
-
Fact-Checking the Output of Large Language Models via Token-Level Uncertainty Quantification
2024 · arXiv (Cornell University)
Large language models (LLMs) are notorious for hallucinating, i.e., producing erroneous claims in their output. Such hallucinations can be dangerous, as occasional factual inaccuracies in the generated text might be obscured by the rest of …
-
Benchmarking Uncertainty Quantification Methods for Large Language Models with LM-Polygraph
2024 · arXiv (Cornell University)
The rapid proliferation of large language models (LLMs) has stimulated researchers to seek effective and efficient approaches to deal with LLM hallucinations and low-quality outputs. Uncertainty quantification (UQ) is a key element of machine learning …
-
Demystifying Instruction Mixing for Fine-tuning Large Language Models
2024
Renxi Wang, Haonan Li, Minghao Wu, Yuxia Wang, Xudong Han, Chiyu Zhang, Timothy Baldwin. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop). 2024.
-
Language Bias in Multilingual Information Retrieval: The Nature of the Beast and Mitigation Methods
2024
Language fairness in multilingual information retrieval (MLIR) systems is crucial for ensuring equitable access to information across diverse languages.This paper sheds light on the issue, based on the assumption that queries in different languages, but …
-
CQADupStack
2015
This paper presents a benchmark dataset, CQADupStack, for use in community question-answering (cQA) research. It contains threads from twelve StackExchange subforums, annotated with duplicate question information. We provide pre-defined training and test splits, both for …
-
Can machine translation systems be evaluated by the crowd alone
2015 · Natural Language Engineering
Abstract Crowd-sourced assessments of machine translation quality allow evaluations to be carried out cheaply and on a large scale. It is essential, however, that the crowd's work be filtered to avoid contamination of results through …
-
Accurate Evaluation of Segment-level Machine Translation Metrics
2015
Evaluation of segment-level machine translation metrics is currently hampered by: (1) low inter-annotator agreement levels in human assessments; (2) lack of an effective mechanism for evaluation of translations of equal quality; and (3) lack of …
-
SemEval-2017 Task 3: Community Question Answering
2017
Preslav Nakov, Doris Hoogeveen, Lluís Màrquez, Alessandro Moschitti, Hamdy Mubarak, Timothy Baldwin, Karin Verspoor. Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval-2017). 2017.