Researcher profile

Soumi Maiti

1 paper in the PaperMetrix corpus

Publications

Papers by this author

  1. SynesLM: A Unified Approach for Audio-visual Speech Recognition and Translation via Language Model and Synthetic Data

    2024 · arXiv (Cornell University)

    In this work, we present SynesLM, an unified model which can perform three multimodal language understanding tasks: audio-visual automatic speech recognition(AV-ASR) and visual-aided speech/machine translation(VST/VMT). Unlike previous research that focused on lip motion as visual …