Alham Fikri Aji
8 papers in the PaperMetrix corpus
Papers by this author
-
One Country, 700+ Languages: NLP Challenges for Underrepresented Languages and Dialects in Indonesia
2022 · arXiv (Cornell University)
NLP research is impeded by a lack of resources and awareness of the challenges presented by underrepresented languages and dialects. Focusing on the languages spoken in Indonesia, the second most linguistically diverse and the fourth …
-
Which Student is Best? A Comprehensive Knowledge Distillation Exam for Task-Specific BERT Models
2022 · arXiv (Cornell University)
We perform knowledge distillation (KD) benchmark from task-specific BERT-base teacher models to various student models: BiLSTM, CNN, BERT-Tiny, BERT-Mini, and BERT-Small. Our experiment involves 12 datasets grouped in two tasks: text classification and sequence labeling …
-
NusaCrowd: Open Source Initiative for Indonesian NLP Resources
2022 · arXiv (Cornell University)
We present NusaCrowd, a collaborative initiative to collect and unify existing resources for Indonesian languages, including opening access to previously non-public resources. Through this initiative, we have brought together 137 datasets and 118 standardized data …
-
Data Laundering: Artificially Boosting Benchmark Results through Knowledge Distillation
2024 · arXiv (Cornell University)
In this paper, we show that knowledge distillation can be subverted to manipulate language model benchmark scores, revealing a critical vulnerability in current evaluation practices. We introduce "Data Laundering," a process that enables the covert …
-
Data Laundering: Artificially Boosting Benchmark Results through Knowledge Distillation
2025
In this paper, we show that knowledge distillation can be subverted to manipulate language model benchmark scores, revealing a critical vulnerability in current evaluation practices.We introduce "Data Laundering," a process that enables the covert transfer …
-
Entropy2Vec: Crosslingual Language Modeling Entropy as End-to-End Learnable Language Representations
2025 · arXiv (Cornell University)
We introduce Entropy2Vec, a novel framework for deriving cross-lingual language representations by leveraging the entropy of monolingual language models. Unlike traditional typological inventories that suffer from feature sparsity and static snapshots, Entropy2Vec uses the inherent …
-
BLOOM: A 176B-Parameter Open-Access Multilingual Language Model
2022 · arXiv (Cornell University)
Large language models (LLMs) have been shown to be able to perform new tasks based on a few demonstrations or natural language instructions. While these capabilities have led to widespread adoption, most LLMs are developed …
-
Crosslingual Generalization through Multitask Finetuning
2023
Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng Xin Yong, Hailey Schoelkopf, Xiangru Tang, Dragomir Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid …