Xu Chu
5 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
PIClean
2019
With the dramatic increasing interest in data analysis, ensuring data quality becomes one of the most important topics in data science. Data Cleaning, the process of ensuring data quality, is composed of two stages: error …
-
DiffPrep: Differentiable Data Preprocessing Pipeline Search for Learning over Tabular Data
2023 · Proceedings of the ACM on Management of Data
Data preprocessing is a crucial step in the machine learning process that transforms raw data into a more usable format for downstream ML models. However, it can be costly and time-consuming, often requiring the expertise …
-
KnowPO: Knowledge-Aware Preference Optimization for Controllable Knowledge Selection in Retrieval-Augmented Language Models
2025 · Proceedings of the AAAI Conference on Artificial Intelligence
By integrating external knowledge, Retrieval-Augmented Generation (RAG) has become an effective strategy for mitigating the hallucination problems that large language models (LLMs) encounter when dealing with knowledge-intensive tasks. However, in the process of integrating external …
-
Domaino1s: Guiding LLM Reasoning for Explainable Answers in High-Stakes Domains
2025
Large Language Models (LLMs) are widely applied to downstream domains.However, current LLMs for high-stakes domain tasks, such as financial investment and legal QA, typically generate brief answers without reasoning processes and explanations.This limits users' confidence …
-
AIMMerging: Adaptive Iterative Model Merging Using Training Trajectories for Language Model Continual Learning
2025 · arXiv (Cornell University)
Continual learning (CL) is essential for deploying large language models (LLMs) in dynamic real-world environments without the need for costly retraining. Recent model merging-based methods have attracted significant attention, but they still struggle to effectively …