Researcher profile
Muhammad Taha Cheema
1 paper in the PaperMetrix corpus
Publications
Papers by this author
-
TweakLLM: A Routing Architecture for Dynamic Tailoring of Cached Responses
2025 · arXiv (Cornell University)
Large Language Models (LLMs) process millions of queries daily, making efficient response caching a compelling optimization for reducing cost and latency. However, preserving relevance to user queries using this approach proves difficult due to the …