Laurent Besacier
14 ورقة في مجموعة PaperMetrix
أوراق هذا المؤلف
-
MultiVec: a Multilingual and Multilevel Representation Learning Toolkit for NLP
2016
We present MultiVec, a new toolkit for computing continuous representations for text at different granularity levels (word-level or sequences of words).MultiVec includes Mikolov et al. [2013b]'s word2vec features, Le and Mikolov [2014]'s paragraph vector (batch …
-
Automatic Quality Assessment for Speech Translation Using Joint ASR and MT Features
2016 · arXiv (Cornell University)
This paper addresses automatic quality assessment of spoken language translation (SLT). This relatively new task is defined and formalized as a sequence labeling problem where each word in the SLT hypothesis is tagged as good …
-
Deep Investigation of Cross-Language Plagiarism Detection Methods
2017
This paper is a deep investigation of cross-language plagiarism detection methods on a new recently introduced open dataset, which contains parallel and comparable collections of documents with multiple characteristics (different genres, languages and sizes of …
-
End-to-End Automatic Speech Translation of Audiobooks
2018
We investigate end-to-end speech-to-text translation on a corpus of audiobooks specifically augmented for this task. Previous works investigated the extreme case where source language transcription is not available during learning nor decoding, but we also …
-
The Zero Resource Speech Challenge 2019: TTS Without T
2019
We present the Zero Resource Speech Challenge 2019, which proposes to build a\nspeech synthesizer without any text or phonetic labels: hence, TTS without T\n(text-to-speech without text). We provide raw audio for a target voice in …
-
A small Griko-Italian speech translation corpus
2018 · White Rose Research Online (University of Leeds, The University of Sheffield, University of York)
This paper presents an extension to a very low-resource parallel corpus collected in an endangered language, Griko, making it useful for computational research. The corpus consists of 330 utterances (about 20 minutes of speech) which …
-
Character-based NMT with Transformer
2019 · arXiv (Cornell University)
Character-based translation has several appealing advantages, but its performance is in general worse than a carefully tuned BPE baseline. In this paper we study the impact of character-based input and output with the Transformer architecture. …
-
Visualizing Cross‐Lingual Discourse Relations in Multilingual TED Corpora
2021
This paper presents an interactive data dashboard that provides users with an overview of the preservation of discourse relations among 28 language pairs. We display a graph network depicting the cross-lingual discourse relations between a …
-
ASR-Generated Text for Language Model Pre-training Applied to Speech Tasks
2022 · arXiv (Cornell University)
We aim at improving spoken language modeling (LM) using very large amount of automatically transcribed speech. We leverage the INA (French National Audiovisual Institute) collection and obtain 19GB of text after applying ASR on 350,000 …
-
Enabling Interactive Transcription in an Indigenous Community
2020 · arXiv (Cornell University)
We propose a novel transcription workflow which combines spoken term detection and human-in-the-loop, together with a pilot experiment. This work is grounded in an almost zero-resource scenario where only a few terms have so far …
-
Findings of the Third Automatic Minuting (AutoMin) Challenge
2025 · arXiv (Cornell University)
This paper presents the third edition of AutoMin, a shared task on automatic meeting summarization into minutes. In 2025, AutoMin featured the main task of minuting, the creation of structured meeting minutes, as well as …
-
Breaking the Unwritten Language Barrier: The BULB Project
2016 · Procedia Computer Science
The project Breaking the Unwritten Language Barrier (BULB), which brings together linguists and computer scientists, aims at supporting linguists in documenting unwritten languages. In order to achieve this we develop tools tailored to the needs …
-
Listen and Translate: A Proof of Concept for End-to-End Speech-to-Text Translation
2016 · arXiv (Cornell University)
This paper proposes a first attempt to build an end-to-end speech-to-text translation system, which does not use source language transcription during learning or decoding. We propose a model for direct speech-to-text translation, which gives promising …
-
A Very Low Resource Language Speech Corpus for Computational Language Documentation Experiments
2018
Most speech and language technologies are trained with massive amounts of speech and text information.However, most of the world languages do not have such resources and some even lack a stable orthography.Building systems under these …