Researcher profile

Guolin Ke

3 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Less is More: Pre-train a Strong Text Encoder for Dense Retrieval Using a Weak Decoder

    2021 · arXiv (Cornell University)

    Dense retrieval requires high-quality text sequence embeddings to support effective search in the representation space. Autoencoder-based language models are appealing in dense retrieval as they train the encoder to output high-quality embedding that can reconstruct …

  2. METRO: Efficient Denoising Pretraining of Large Scale Autoencoding Language Models with Model Generated Signals

    2022 · arXiv (Cornell University)

    We present an efficient method of pretraining large-scale autoencoding language models using training signals generated by an auxiliary model. Originated in ELECTRA, this training strategy has demonstrated sample-efficiency to pretrain models at the scale of …

  3. Rethinking Positional Encoding in Language Pre-training

    2020 · arXiv (Cornell University)

    In this work, we investigate the positional encoding methods used in language pre-training (e.g., BERT) and identify several problems in the existing formulations. First, we show that in the absolute positional encoding, the addition operation …