Yu Sun
15 ورقة في مجموعة PaperMetrix
أوراق هذا المؤلف
-
Learning Effective Embeddings for Machine Generated Emails with Applications to Email Category Prediction
2018
Machine generated business-to-consumer (B2C) emails such as receipts, newsletters, and promotions constitute a large portion of users' inboxes today. These emails reflect the users' interests and often are sequentially correlated, e.g., users interested in relocating …
-
Multi-feature automatic abstract based on LDA model and redundant control
2020 · Journal of Physics Conference Series
Abstract With the continuous popularization of computer application technology and the rapid development of Internet technology, there has been an explosion of information in every field, and more and more information is transmitted and stored …
-
A Real-time Multiplayer FPS Game using 3D Modeling and AI Machine Learning
2022 · Computer Science and Information Technology
AIs have been a key component in the gaming industry throughout its history. Developers have had multiple ways of creating new AI models that best suit their game in order to enhance the playing experience. …
-
High Precision ≠ High Cost: Temporal Data Fusion for Multiple Low-Precision Sensors
2024 · Proceedings of the ACM on Management of Data
High-quality data are crucial for practical applications, but obtaining them through high-precision sensors comes at a high cost. To guarantee the trade-off between cost and precision, we may use multiple low-precision sensors to obtain the …
-
Proxy Prompt: Endowing SAM and SAM 2 with Auto-Interactive-Prompt for Medical Segmentation
2025 · arXiv (Cornell University)
In this paper, we aim to address the unmet demand for automated prompting and enhanced human-model interactions of SAM and SAM2 for the sake of promoting their widespread clinical adoption. Specifically, we propose Proxy Prompt …
-
LLMGuard : Safeguarding Real-Time Inference for Large Language Models on Edge Devices
2026 · ACM Transactions on Software Engineering and Methodology
TEE-shielded secure inference offers an efficient solution to protect valuable edge-deployed models from potential thefts. Nevertheless, existing methods are lack of theoretical security analysis, failing to achieve the optimal security. Furthermore, while feasible for small …
-
Supervised word mover's distance
2016 · PolyPublie (École Polytechnique de Montréal)
Recently, a new document metric called the word mover’s distance (WMD) has been proposed with unprecedented results on kNN-based document classification. The WMD elevates high-quality word embeddings to a document metric by formulating the distance …
-
Field-weighted Factorization Machines for Click-Through Rate Prediction in Display Advertising
2018
Click-through rate (CTR) prediction is a critical task in online display advertising. The data involved in CTR prediction are typically multi-field categorical data, i.e., every feature is categorical and belongs to one and only one …
-
ERNIE: Enhanced Representation through Knowledge Integration
2019 · arXiv (Cornell University)
We present a novel language representation model enhanced by knowledge called ERNIE (Enhanced Representation through kNowledge IntEgration). Inspired by the masking strategy of BERT, ERNIE is designed to learn language representation enhanced by knowledge masking …
-
Adversarial Deep Averaging Networks for Cross-Lingual Sentiment Classification
2018 · Transactions of the Association for Computational Linguistics
In recent years great success has been achieved in sentiment classification for English, thanks in part to the availability of copious annotated resources. Unfortunately, most languages do not enjoy such an abundance of labeled data. …
-
ERNIE 2.0: A Continual Pre-training Framework for Language Understanding
2019 · arXiv (Cornell University)
Recently, pre-trained models have achieved state-of-the-art results in various language understanding tasks, which indicates that pre-training on large-scale corpora may play a crucial role in natural language processing. Current pre-training procedures usually focus on training …
-
ERNIE 2.0: A Continual Pre-Training Framework for Language Understanding
2020 · Proceedings of the AAAI Conference on Artificial Intelligence
Recently pre-trained models have achieved state-of-the-art results in various language understanding tasks. Current pre-training procedures usually focus on training the model with several simple tasks to grasp the co-occurrence of words or sentences. However, besides …
-
ERNIE-M: Enhanced Multilingual Representation by Aligning Cross-lingual Semantics with Monolingual Corpora
2021 · Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing
Recent studies have demonstrated that pretrained cross-lingual models achieve impressive performance in downstream cross-lingual tasks. This improvement benefits from learning a large amount of monolingual and parallel corpora. Although it is generally acknowledged that parallel …
-
ERNIE 3.0: Large-scale Knowledge Enhanced Pre-training for Language Understanding and Generation
2021 · arXiv (Cornell University)
Pre-trained models have achieved state-of-the-art results in various Natural Language Processing (NLP) tasks. Recent works such as T5 and GPT-3 have shown that scaling up pre-trained language models can improve their generalization abilities. Particularly, the …
-
From Word Embeddings To Document Distances
2015 · PolyPublie (École Polytechnique de Montréal)
We present the Word Mover's Distance (WMD), a novel distance function between text documents. Our work is based on recent results in word embeddings that learn semantically meaningful representations for words from local cooccurrences in …