See-Kiong Ng
5 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models
2024 · arXiv (Cornell University)
Large language models (LLMs) have demonstrated impressive reasoning capabilities, particularly in textual mathematical problem-solving. However, existing open-source image instruction fine-tuning datasets, containing limited question-answer pairs per image, do not fully exploit visual information to enhance …
-
Efficient Inference for Large Language Model-based Generative Recommendation
2024 · arXiv (Cornell University)
Large Language Model (LLM)-based generative recommendation has achieved notable success, yet its practical deployment is costly particularly due to excessive inference latency caused by autoregressive decoding. For lossless LLM decoding acceleration, Speculative Decoding (SD) has …
-
Multi-Modal One-Shot Federated Ensemble Learning for Medical Data with Vision Large Language Model
2025 · arXiv (Cornell University)
Federated learning (FL) has attracted considerable interest in the medical domain due to its capacity to facilitate collaborative model training while maintaining data privacy. However, conventional FL methods typically necessitate multiple communication rounds, leading to …
-
Balancing Truthfulness and Informativeness with Uncertainty-Aware Instruction Fine-Tuning
2025 · arXiv (Cornell University)
Instruction fine-tuning (IFT) can increase the informativeness of large language models (LLMs), but may reduce their truthfulness. This trade-off arises because IFT steers LLMs to generate responses containing long-tail knowledge that was not well covered …
-
MCM-DPO: Multifaceted Cross-Modal Direct Preference Optimization for Alt-text Generation
2025 · arXiv (Cornell University)
The alt-text generation task produces concise, context-relevant descriptions of images, enabling blind and low-vision users to access online images. Despite the capabilities of large vision-language models, alt-text generation performance remains limited due to noisy user …