ملف الباحث

Ren, Qianli

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Do we really have to filter out random noise in pre-training data for language models?

    2025 · arXiv (Cornell University)

    Web-scale pre-training datasets are the cornerstone of LLMs' success. However, text data curated from the Internet inevitably contains random noise caused by decoding errors or unregulated web content. In contrast to previous works that focus …