EARSHOT: A minimal neural network model of incremental human speech recognition
At a glance
- Citations
- 1
- References
- 21
- Comments
- 0
Abstract
Despite the “lack of invariance problem” (multiple acoustic patterns map to the same phoneme, and one acoustic pattern can map to different phonemes), humans experience phonetic constancy: we typically perceive what the speaker intended despite this variability. Computational models of human speech recognition have deferred the problem, working with idealized inputs rather than speech. Deep learning models have allowed automatic speech recognition to become robust, but are too complex to relate directly to theories of human speech recognition. We report results from a simple network using long short-term memory nodes, which allow it to learn to recognize real speech with high accuracy, with minimal complexity. Representations emerge in the model that resemble those observed in human cortical responses to speech.
Publication details
- DOI
- 10.31234/osf.io/h7a4n
- OpenAlex
- W4240369743
- Document type
- preprint
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.