preprint Open access

EARSHOT: A minimal neural network model of incremental human speech recognition

Research footprint

At a glance

Citations
1
References
21
Comments
0
Paper overview

Abstract

Despite the “lack of invariance problem” (multiple acoustic patterns map to the same phoneme, and one acoustic pattern can map to different phonemes), humans experience phonetic constancy: we typically perceive what the speaker intended despite this variability. Computational models of human speech recognition have deferred the problem, working with idealized inputs rather than speech. Deep learning models have allowed automatic speech recognition to become robust, but are too complex to relate directly to theories of human speech recognition. We report results from a simple network using long short-term memory nodes, which allow it to learn to recognize real speech with high accuracy, with minimal complexity. Representations emerge in the model that resemble those observed in human cortical responses to speech.

Record transparency

Publication details

DOI
10.31234/osf.io/h7a4n
OpenAlex
W4240369743
Document type
preprint
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.