Y. Chen
3 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
DeepServe: Serverless Large Language Model Serving at Scale
2025 · arXiv (Cornell University)
In this paper, we propose DEEPSERVE, a scalable and serverless AI platform designed to efficiently serve large language models (LLMs) at scale in cloud environments. DEEPSERVE addresses key challenges such as resource allocation, serving efficiency, …
-
Predicting 30-Day Postoperative Mortality and American Society of Anesthesiologists Physical Status Using Retrieval-Augmented Large Language Models: Development and Validation Study
2025 · Journal of Medical Internet Research
Background Accurately assessing perioperative risk is critical for informed surgical planning and patient safety. However, current prediction models often rely on structured data and overlook the nuanced clinical reasoning embedded in free-text preoperative notes. Recent …
-
OLMoTrace: Tracing Language Model Outputs Back to Trillions of Training Tokens
2025 · arXiv (Cornell University)
We present OLMoTrace, the first system that traces the outputs of language models back to their full, multi-trillion-token training data in real time. OLMoTrace finds and shows verbatim matches between segments of language model output …