conference-paper Open access

English-to-Japanese Multimodal Machine Translation Based on Image-Text Matching of Lecture Videos

Research footprint

At a glance

Citations
1
References
0
Comments
0
Paper overview

Abstract

We work on a multimodal machine translation of the audio contained in English lecture videos to generate Japanese subtitles.Image-guided multimodal machine translation is promising for error correction in speech recognition and for text disambiguation.In our situation, lecture videos provide a variety of images.Images of presentation materials can complement information not available from audio and may help improve translation quality.However, images of speakers or audiences would not directly affect the translation quality.We construct a multimodal parallel corpus with automatic speech recognition text and multiple images for a transcribed parallel corpus of lecture videos, and propose a method to select the most relevant ones from the multiple images with the speech text for improving the performance of image-guided multimodal machine translation.Experimental results on translating automatic speech recognition or transcribed English text into Japanese show the effectiveness of our method to select a relevant image.

Record transparency

Publication details

DOI
10.18653/v1/2024.alvr-1.7
OpenAlex
W4402683909
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.