Yue Guan
3 papers in the PaperMetrix corpus
Papers by this author
-
Accelerating Sparse DNN Models without Hardware-Support via Tile-Wise Sparsity
2020
Network pruning can reduce the high computation cost of deep neural network (DNN) models. However, to maintain their accuracies, sparse models often carry randomly-distributed weights, leading to irregular computations. Consequently, sparse models cannot achieve meaningful …
-
On the Adversarial Convex Body Chasing Problem
2022 · arXiv (Cornell University)
In this work, we extend the convex bodies chasing problem (CBC) to an adversarial setting, where an agent (the Player) is tasked with chasing a sequence of convex bodies generated adversarially by another agent (the …
-
KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows
2025
Large language model (LLM) based agentic workflows have become a popular paradigm for coordinating multiple specialized agents to solve complex tasks. To improve serving efficiency, existing LLM systems employ prefix caching to reuse key-value (KV) …