Genta Indra Winata
9 papers in the PaperMetrix corpus
Papers by this author
-
Handling imbalanced dataset in multi-label text categorization using Bagging and Adaptive Boosting
2015
Imbalanced dataset is occurred due to uneven distribution of data available in the real world such as disposition of complaints on government offices in Bandung. Consequently, multi-label text categorization algorithms may not produce the best …
-
Hierarchical Meta-Embeddings for Code-Switching Named Entity Recognition
2019
Genta Indra Winata, Zhaojiang Lin, Jamin Shin, Zihan Liu, Pascale Fung. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
-
MinTL: Minimalist Transfer Learning for Task-Oriented Dialogue Systems
2020 · arXiv (Cornell University)
In this paper, we propose Minimalist Transfer Learning (MinTL) to simplify the system design process of task-oriented dialogue systems and alleviate the over-dependency on annotated data. MinTL is a simple yet effective transfer learning framework, …
-
One Country, 700+ Languages: NLP Challenges for Underrepresented Languages and Dialects in Indonesia
2022 · arXiv (Cornell University)
NLP research is impeded by a lack of resources and awareness of the challenges presented by underrepresented languages and dialects. Focusing on the languages spoken in Indonesia, the second most linguistically diverse and the fourth …
-
NusaCrowd: Open Source Initiative for Indonesian NLP Resources
2022 · arXiv (Cornell University)
We present NusaCrowd, a collaborative initiative to collect and unify existing resources for Indonesian languages, including opening access to previously non-public resources. Through this initiative, we have brought together 137 datasets and 118 standardized data …
-
Linguistics Theory Meets LLM: Code-Switched Text Generation via Equivalence Constrained Large Language Models
2026
Code-switching is a common practice for millions of multilingual speakers but remains challenging for Large Language Models (LLMs).This paper investigates LLM capabilities in generating code-switched text, conducting extensive experiments across five diverse language pairs: English …
-
MetaMetrics-MT: Tuning Meta-Metrics for Machine Translation via Human Preference Calibration
2024 · arXiv (Cornell University)
We present MetaMetrics-MT, an innovative metric designed to evaluate machine translation (MT) tasks by aligning closely with human preferences through Bayesian optimization with Gaussian Processes. MetaMetrics-MT enhances existing MT metrics by optimizing their correlation with …
-
Entropy2Vec: Crosslingual Language Modeling Entropy as End-to-End Learnable Language Representations
2025 · arXiv (Cornell University)
We introduce Entropy2Vec, a novel framework for deriving cross-lingual language representations by leveraging the entropy of monolingual language models. Unlike traditional typological inventories that suffer from feature sparsity and static snapshots, Entropy2Vec uses the inherent …
-
Learning Multilingual Meta-Embeddings for Code-Switching Named Entity Recognition
2019
In this paper, we propose Multilingual Meta-Embeddings (MME), an effective method to learn multilingual representations by leveraging monolingual pre-trained embeddings. MME learns to utilize information from these embeddings via a self-attention mechanism without explicit language …