Hiroshi Saruwatari
5 papers in the PaperMetrix corpus
Papers by this author
-
J-MAC: Japanese multi-speaker audiobook corpus for speech synthesis
2022 · arXiv (Cornell University)
In this paper, we construct a Japanese audiobook speech corpus called "J-MAC" for speech synthesis research. With the success of reading-style speech synthesis, the research target is shifting to tasks that use complicated contexts. Audiobook …
-
Improving Speech Prosody of Audiobook Text-to-Speech Synthesis with Acoustic and Textual Contexts
2022 · arXiv (Cornell University)
We present a multi-speaker Japanese audiobook text-to-speech (TTS) system that leverages multimodal context information of preceding acoustic context and bilateral textual context to improve the prosody of synthetic speech. Previous work either uses unilateral or …
-
Diversity-based core-set selection for text-to-speech with linguistic and acoustic features
2023 · arXiv (Cornell University)
This paper proposes a method for extracting a lightweight subset from a text-to-speech (TTS) corpus ensuring synthetic speech quality. In recent years, methods have been proposed for constructing large-scale TTS corpora by collecting diverse data …
-
Coco-Nut: Corpus of Japanese Utterance and Voice Characteristics Description for Prompt-based Control
2023 · arXiv (Cornell University)
In text-to-speech, controlling voice characteristics is important in achieving various-purpose speech synthesis. Considering the success of text-conditioned generation, such as text-to-image, free-form text instruction should be useful for intuitive and complicated control of voice characteristics. …
-
JSUT corpus: free large-scale Japanese speech corpus for end-to-end speech synthesis
2017 · arXiv (Cornell University)
Thanks to improvements in machine learning techniques including deep learning, a free large-scale speech corpus that can be shared between academic institutions and commercial companies has an important role. However, such a corpus for Japanese …