Researcher profile

Herman Kamper

9 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Phoneme Based Embedded Segmental K-Means for Unsupervised Term Discovery

    2018

    Identifying and grouping the frequently occurring word-like patterns from raw acoustic waveforms is an important task in the zero resource speech processing. Embedded segmental K-means (ES-KMeans) discovers both the word boundaries and the word types …

  2. Almost Zero-Resource ASR-free Keyword Spotting using Multilingual Bottleneck Features and Correspondence Autoencoders.

    2018 · arXiv (Cornell University)

    We compare features for dynamic time warping based keyword spotting in an almost zero-resource setting. The objective is to support United Nations (UN) humanitarian relief efforts in parts of Africa with severely under-resourced languages. As …

  3. Unsupervised vs. Transfer Learning for Multimodal One-Shot Matching of Speech and Images

    2020

    We consider the task of multimodal one-shot speech-image matching.An agent is shown a picture along with a spoken word describing the object in the picture, e.g.cookie, broccoli and ice-cream.After observing one paired speech-image example per …

  4. A Temporal Extension of Latent Dirichlet Allocation for Unsupervised Acoustic Unit Discovery

    2022 · arXiv (Cornell University)

    Latent Dirichlet allocation (LDA) is widely used for unsupervised topic modelling on sets of documents. No temporal information is used in the model. However, there is often a relationship between the corresponding topics of consecutive …

  5. Word Segmentation on Discovered Phone Units With Dynamic Programming and Self-Supervised Scoring

    2022 · IEEE/ACM Transactions on Audio Speech and Language Processing

    Recent work on unsupervised speech segmentation has used self-supervised models with phone and word segmentation modules that are trained jointly. This paper instead revisits an older approach to word segmentation: bottom-up phone-like unit discovery is …

  6. Leveraging multilingual transfer for unsupervised semantic acoustic word embeddings

    2023 · arXiv (Cornell University)

    Acoustic word embeddings (AWEs) are fixed-dimensional vector representations of speech segments that encode phonetic content so that different realisations of the same word have similar embeddings. In this paper we explore semantic AWE modelling. These …

  7. Unsupervised Word Discovery: Boundary Detection with Clustering vs. Dynamic Programming

    2024 · arXiv (Cornell University)

    We look at the long-standing problem of segmenting unlabeled speech into word-like segments and clustering these into a lexicon. Several previous methods use a scoring model coupled with dynamic programming to find an optimal segmentation. …

  8. LinearVC: Linear transformations of self-supervised features through the lens of voice conversion

    2025 · arXiv (Cornell University)

    We introduce LinearVC, a simple voice conversion method that sheds light on the structure of self-supervised representations. First, we show that simple linear transformations of self-supervised features effectively convert voices. Next, we probe the geometry …

  9. Unsupervised Word Segmentation and Lexicon Discovery Using Acoustic Word Embeddings

    2016 · IEEE/ACM Transactions on Audio Speech and Language Processing

    In settings where only unlabeled speech data is available, speech technology needs to be developed without transcriptions, pronunciation dictionaries, or language modelling text. A similar problem is faced when modeling infant language acquisition. In these …