Researcher profile

X. Zhang

1 paper in the PaperMetrix corpus

Publications

Papers by this author

  1. Collaborative Speculative Inference for Efficient LLM Inference Serving

    2025 · arXiv (Cornell University)

    Speculative inference is a promising paradigm employing small speculative models (SSMs) as drafters to generate draft tokens, which are subsequently verified in parallel by the target large language model (LLM). This approach enhances the efficiency …