Herman Kamper
9 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
Phoneme Based Embedded Segmental K-Means for Unsupervised Term Discovery
2018
Identifying and grouping the frequently occurring word-like patterns from raw acoustic waveforms is an important task in the zero resource speech processing. Embedded segmental K-means (ES-KMeans) discovers both the word boundaries and the word types …
-
Almost Zero-Resource ASR-free Keyword Spotting using Multilingual Bottleneck Features and Correspondence Autoencoders.
2018 · arXiv (Cornell University)
We compare features for dynamic time warping based keyword spotting in an almost zero-resource setting. The objective is to support United Nations (UN) humanitarian relief efforts in parts of Africa with severely under-resourced languages. As …
-
Unsupervised vs. Transfer Learning for Multimodal One-Shot Matching of Speech and Images
2020
We consider the task of multimodal one-shot speech-image matching.An agent is shown a picture along with a spoken word describing the object in the picture, e.g.cookie, broccoli and ice-cream.After observing one paired speech-image example per …
-
A Temporal Extension of Latent Dirichlet Allocation for Unsupervised Acoustic Unit Discovery
2022 · arXiv (Cornell University)
Latent Dirichlet allocation (LDA) is widely used for unsupervised topic modelling on sets of documents. No temporal information is used in the model. However, there is often a relationship between the corresponding topics of consecutive …
-
Word Segmentation on Discovered Phone Units With Dynamic Programming and Self-Supervised Scoring
2022 · IEEE/ACM Transactions on Audio Speech and Language Processing
Recent work on unsupervised speech segmentation has used self-supervised models with phone and word segmentation modules that are trained jointly. This paper instead revisits an older approach to word segmentation: bottom-up phone-like unit discovery is …
-
Leveraging multilingual transfer for unsupervised semantic acoustic word embeddings
2023 · arXiv (Cornell University)
Acoustic word embeddings (AWEs) are fixed-dimensional vector representations of speech segments that encode phonetic content so that different realisations of the same word have similar embeddings. In this paper we explore semantic AWE modelling. These …
-
Unsupervised Word Discovery: Boundary Detection with Clustering vs. Dynamic Programming
2024 · arXiv (Cornell University)
We look at the long-standing problem of segmenting unlabeled speech into word-like segments and clustering these into a lexicon. Several previous methods use a scoring model coupled with dynamic programming to find an optimal segmentation. …
-
LinearVC: Linear transformations of self-supervised features through the lens of voice conversion
2025 · arXiv (Cornell University)
We introduce LinearVC, a simple voice conversion method that sheds light on the structure of self-supervised representations. First, we show that simple linear transformations of self-supervised features effectively convert voices. Next, we probe the geometry …
-
Unsupervised Word Segmentation and Lexicon Discovery Using Acoustic Word Embeddings
2016 · IEEE/ACM Transactions on Audio Speech and Language Processing
In settings where only unlabeled speech data is available, speech technology needs to be developed without transcriptions, pronunciation dictionaries, or language modelling text. A similar problem is faced when modeling infant language acquisition. In these …