Roy Ka-Wei Lee
7 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
Text Style Transfer: A Review and Experiment Evaluation.
2020 · arXiv (Cornell University)
The stylistic properties of text have intrigued computational linguistics researchers in recent years. Specifically, researchers have investigated the Text Style Transfer (TST) task, which aims to change the stylistic properties of the text while retaining …
-
GitHub and Stack Overflow: Analyzing Developer Interests Across Multiple Social Collaborative Platforms
2017 · arXiv (Cornell University)
Increasingly, software developers are using a wide array of social collaborative platforms for software development and learning. In this work, we examined the similarities in developer's interests within and across GitHub and Stack Overflow. Our …
-
On Explaining Multimodal Hateful Meme Detection Models
2022 · arXiv (Cornell University)
Hateful meme detection is a new multimodal task that has gained significant traction in academic and industry research communities. Recently, researchers have applied pre-trained visual-linguistic models to perform the multimodal classification task, and some of …
-
Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models
2024 · arXiv (Cornell University)
Large language models (LLMs) have demonstrated impressive reasoning capabilities, particularly in textual mathematical problem-solving. However, existing open-source image instruction fine-tuning datasets, containing limited question-answer pairs per image, do not fully exploit visual information to enhance …
-
Toxicity Red-Teaming: Benchmarking LLM Safety in Singapore's Low-Resource Languages
2025 · arXiv (Cornell University)
The advancement of Large Language Models (LLMs) has transformed natural language processing; however, their safety mechanisms remain under-explored in low-resource, multilingual settings. Here, we aim to bridge this gap. In particular, we introduce \textsf{SGToxicGuard}, a …
-
Lost in Localization: Building RabakBench with Human-in-the-Loop Validation to Measure Multilingual Safety Gaps
2025 · arXiv (Cornell University)
Large language models (LLMs) often fail to maintain safety in low-resource language varieties, such as code-mixed vernaculars and regional dialects. We introduce RabakBench, a multilingual safety benchmark and scalable pipeline localized to Singapore's unique linguistic …
-
Graph-to-Tree Learning for Solving Math Word Problems
2020
While the recent tree-based neural models have demonstrated promising results in generating solution expression for the math word problem (MWP), most of these models do not capture the relationships and order information among the quantities …