Researcher profile

Guo, Zhihui

1 paper in the PaperMetrix corpus

Publications

Papers by this author

  1. Do we really have to filter out random noise in pre-training data for language models?

    2025 · arXiv (Cornell University)

    Web-scale pre-training datasets are the cornerstone of LLMs' success. However, text data curated from the Internet inevitably contains random noise caused by decoding errors or unregulated web content. In contrast to previous works that focus …