Jingyuan Zhang
3 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction
2024 · arXiv (Cornell University)
The attention operator is arguably the key distinguishing factor of transformer architectures, which have demonstrated state-of-the-art performance on a variety of tasks. However, transformer attention operators often impose a significant computational burden, with the computational …
-
Knowledge Graph Embedding Based Question Answering
2019
Question answering over knowledge graph (QA-KG) aims to use facts in the knowledge graph (KG) to answer natural language questions. It helps end users more efficiently and more easily access the substantial and valuable knowledge …
-
AIBox
2019
As one of the major search engines in the world, Baidu's Sponsored Search has long adopted the use of deep neural network (DNN) models for Ads click-through rate (CTR) predictions, as early as in 2013. …