Rangan Majumder
4 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
MS MARCO: A Human Generated MAchine Reading COmprehension Dataset.
2016 · Neural Information Processing Systems
This paper presents our recent work on the design and development of a new, large scale dataset, which we name MS MARCO, for MAchine Reading COmprehension. This new dataset is aimed to overcome a number …
-
Analysis of Points of Interests Recommended for Leisure Walk Descriptions
2024 · arXiv (Cornell University)
Data for Sub-Task 1 of the Advertisement in Retrieval-Augmented Generation task at Touché 2025. The dataset contains segments retrieved from the segmented version of MS MARCO V2.1. The queries used in retrieval are taken from …
-
XGLUE: A New Benchmark Dataset for Cross-lingual Pre-training, Understanding and Generation
2020
Yaobo Liang, Nan Duan, Yeyun Gong, Ning Wu, Fenfei Guo, Weizhen Qi, Ming Gong, Linjun Shou, Daxin Jiang, Guihong Cao, Xiaodong Fan, Ruofei Zhang, Rahul Agrawal, Edward Cui, Sining Wei, Taroon Bharti, Ying Qiao, Jiun-Hung …
-
Text Embeddings by Weakly-Supervised Contrastive Pre-training
2022 · arXiv (Cornell University)
This paper presents E5, a family of state-of-the-art text embeddings that transfer well to a wide range of tasks. The model is trained in a contrastive manner with weak supervision signals from our curated large-scale …