conference-paper Open access

Learning Text Similarity with Siamese Recurrent Networks

Research footprint

At a glance

Citations
344
References
55
Comments
0
Paper overview

Öz

This paper presents a deep architecture for learning a similarity metric on variablelength character sequences. The model combines a stack of character-level bidirectional LSTM's with a Siamese architecture. It learns to project variablelength strings into a fixed-dimensional embedding space by using only information about the similarity between pairs of strings. This model is applied to the task of job title normalization based on a manually annotated taxonomy. A small data set is incrementally expanded and augmented with new sources of variance. The model learns a representation that is selective to differences in the input that reflect semantic differences (e.g., "Java developer" vs. "HR manager") but also invariant to nonsemantic string differences (e.g., "Java developer" vs. "Java programmer").

Record transparency

Publication details

DOI
10.18653/v1/w16-1617
OpenAlex
W2510940142
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.