Researcher profile

Bo Zheng

6 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Improving Pretrained Cross-Lingual Language Models via Self-Labeled Word Alignment

    2021 · arXiv (Cornell University)

    The cross-lingual language models are typically pretrained with masked language modeling on multilingual text or parallel sentences. In this paper, we introduce denoising word alignment as a new cross-lingual pre-training task. Specifically, the model first …

  2. Unified Visual Preference Learning for User Intent Understanding

    2024

    In the world of E-Commerce, the core task is to understand the personalized preference from various kinds of heterogeneous information, such as textual reviews, item images and historical behaviors. In current systems, these heterogeneous information …

  3. Small Models are LLM Knowledge Triggers on Medical Tabular Prediction

    2024 · arXiv (Cornell University)

    Recent development in large language models (LLMs) has demonstrated impressive domain proficiency on unstructured textual or multi-modal tasks. However, despite with intrinsic world knowledge, their application on structured tabular data prediction still lags behind, primarily …

  4. SEGMENT+: Long Text Processing with Short-Context Language Models

    2024 · arXiv (Cornell University)

    There is a growing interest in expanding the input capacity of language models (LMs) across various domains. However, simply increasing the context window does not guarantee robust performance across diverse long-input processing tasks, such as …

  5. GraphReader: Building Graph-based Agent to Enhance Long-Context Abilities of Large Language Models

    2024

    Shilong Li, Yancheng He, Hangyu Guo, Xingyuan Bu, Ge Bai, Jie Liu, Jiaheng Liu, Xingwei Qu, Yangguang Li, Wanli Ouyang, Wenbo Su, Bo Zheng. Findings of the Association for Computational Linguistics: EMNLP 2024. 2024.

  6. Towards Better

    2018 · Proceedings of the

    This paper describes our system (HIT-SCIR) submitted to the CoNLL 2018 shared task on Multilingual Parsing from Raw Text to Universal Dependencies. We base our submission on Stanford's winning system for the CoNLL 2017 shared …