Jingdong Chen
3 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
Partial AUC optimization based deep speaker embeddings with class-center learning for text-independent speaker verification
2019 · arXiv (Cornell University)
Deep embedding based text-independent speaker verification has demonstrated superior performance to traditional methods in many challenging scenarios. Its loss functions can be generally categorized into two classes, i.e., verification and identification. The verification loss functions …
-
Accelerating Pre-training of Multimodal LLMs via Chain-of-Sight
2024 · arXiv (Cornell University)
This paper introduces Chain-of-Sight, a vision-language bridge module that accelerates the pre-training of Multimodal Large Language Models (MLLMs). Our approach employs a sequence of visual resamplers that capture visual details at various spacial scales. This …
-
Deep Speech 2: End-to-End Speech Recognition in English and Mandarin
2015 · arXiv (Cornell University)
We show that an end-to-end deep learning approach can be used to recognize either English or Mandarin Chinese speech--two vastly different languages. Because it replaces entire pipelines of hand-engineered components with neural networks, end-to-end learning …