Researcher profile

Bin Ma

4 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Orthogonal Temporal Interpolation for Zero-Shot Video Recognition

    2023 · arXiv (Cornell University)

    Zero-shot video recognition (ZSVR) is a task that aims to recognize video categories that have not been seen during the model training process. Recently, vision-language models (VLMs) pre-trained on large-scale image-text pairs have demonstrated impressive …

  2. Are Soft Prompts Good Zero-Shot Learners for Speech Recognition?

    2024

    Large self-supervised pre-trained speech models require computationally expensive fine-tuning for downstream tasks. Soft prompt tuning offers a simple parameter-efficient alternative by utilizing minimal soft prompt guidance, enhancing portability while also maintaining competitive performance. However, not …

  3. Emotional Dimension Control in Language Model-Based Text-To-Speech: Spanning a Broad Spectrum of Human Emotions

    2026

    Emotional text-to-speech (TTS) systems struggle to capture the full spectrum of human emotions due to the inherent complexity of emotional expressions and the limited coverage of existing emotion labels. To address this, we propose a …

  4. Constrained Output Embeddings for End-to-End Code-Switching Speech Recognition with Only Monolingual Data

    2019

    The lack of code-switch training data is one of the major concerns in the development of end-to-end code-switching automatic speech recognition (ASR) models. In this work, we propose a method to train an improved end-to-end …