Fei, Hao
4 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
PanoSent: A Panoptic Sextuple Extraction Benchmark for Multimodal Conversational Aspect-based Sentiment Analysis
2024 · arXiv (Cornell University)
While existing Aspect-based Sentiment Analysis (ABSA) has received extensive effort and advancement, there are still gaps in defining a more holistic research target seamlessly integrating multimodality, conversation context, fine-granularity, and also covering the changing sentiment …
-
Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark
2025 · arXiv (Cornell University)
Empathetic Response Generation (ERG) is one of the key tasks of the affective computing area, which aims to produce emotionally nuanced and compassionate responses to user's queries. However, existing ERG research is predominantly confined to …
-
MCM-DPO: Multifaceted Cross-Modal Direct Preference Optimization for Alt-text Generation
2025 · arXiv (Cornell University)
The alt-text generation task produces concise, context-relevant descriptions of images, enabling blind and low-vision users to access online images. Despite the capabilities of large vision-language models, alt-text generation performance remains limited due to noisy user …
-
Watch Out Your Album! On the Inadvertent Privacy Memorization in Multi-Modal Large Language Models
2025 · arXiv (Cornell University)
Multi-Modal Large Language Models (MLLMs) have exhibited remarkable performance on various vision-language tasks such as Visual Question Answering (VQA). Despite accumulating evidence of privacy concerns associated with task-relevant content, it remains unclear whether MLLMs inadvertently …