Researcher profile
Chaoqun Yang
1 paper in the PaperMetrix corpus
Publications
Papers by this author
-
Efficient Inference for Large Language Model-based Generative Recommendation
2024 · arXiv (Cornell University)
Large Language Model (LLM)-based generative recommendation has achieved notable success, yet its practical deployment is costly particularly due to excessive inference latency caused by autoregressive decoding. For lossless LLM decoding acceleration, Speculative Decoding (SD) has …