Rui Wang
34 ورقة في مجموعة PaperMetrix
أوراق هذا المؤلف
-
Converting Continuous-Space Language Models into <i>N</i> -gram Language Models with Efficient Bilingual Pruning for Statistical Machine Translation
2016 · ACM Transactions on Asian and Low-Resource Language Information Processing
The Language Model (LM) is an essential component of Statistical Machine Translation (SMT). In this article, we focus on developing efficient methods for LM construction. Our main contribution is that we propose a Natural N …
-
English to Chinese Translation: How Chinese Character Matters
2015
Word segmentation is helpful in Chinese nat-ural language processing in many aspects. However it is showed that different word seg-mentation strategies do not affect the per-formance of Statistical Machine Translation (SMT) from English to Chinese …
-
Research of Database Encryption Technology and Its Application
2017
With the continuous improvement of the national level of information technology, especially the rapid development of Internet technology, IT database system is more and more widely. However, the database system data storage and extensive sharing …
-
Research on key technology of data mining based on hospital information system
2017
At present, the processing of medical information is mostly based on the level of the operation of database technology, it is the specific application service. Hospital information system is a branch of medical informatics, which …
-
Bilingual Education Transformation on Textiles Material Experiment
2017 · DEStech Transactions on Social Science Education and Human Science
Bilingual education is a great attempt of high education reform in our country, and becomes the new question facing by many teachers. This study discuss a bilingual education of Textile Material Experiment on teaching efficiency …
-
Instance Weighting for Neural Machine Translation Domain Adaptation
2017
Instance weighting has been widely applied to phrase-based machine translation domain adaptation. However, it is challenging to be applied to Neural Machine Translation (NMT) directly, because NMT is not a linear model.
-
Optimization Design of Electromagnetic Relay Based on Improved Particle Swarm Optimization Algorithm
2018
The optimal design of the relay is in ensuring reliable of electromagnetic relay absorbed and released by changing the premise of structural parameters and materials to achieve energy saving purposes and reduce the energy of …
-
An Atari Model Zoo for Analyzing, Visualizing, and Comparing Deep Reinforcement Learning Agents
2019
Much human and computational effort has aimed to improve how deep reinforcement learning (DRL) algorithms perform on benchmarks such as the Atari Learning Environment. Comparatively less effort has focused on understanding what has been learned …
-
SUDA-Alibaba at MRP 2019: Graph-Based Models with BERT
2019
Yue Zhang, Wei Jiang, Qingrong Xia, Junjie Cao, Rui Wang, Zhenghua Li, Min Zhang. Proceedings of the Shared Task on Cross-Framework Meaning Representation Parsing at the 2019 Conference on Natural Language Learning. 2019.
-
Weighted Focus-Attention Deep Network for Fine-grained Image Classification
2019
Fine-Grained Visual Classification (FGVC) is a challenging task, due to the small variation of visual representations from different categories. An effective solution is utilizing the bounding boxes centering the object parts to extract the discriminative …
-
Information Fusion-Based Deep Neural Attentive Matrix Factorization Recommendation
2021 · Algorithms
The emergence of the recommendation system has effectively alleviated the information overload problem. However, traditional recommendation systems either ignore the rich attribute information of users and items, such as the user’s social-demographic features, the item’s …
-
Representer Theorems in Banach Spaces: Minimum Norm Interpolation,\n Regularized Learning and Semi-Discrete Inverse Problems
2020 · arXiv (Cornell University)
Constructing or learning a function from a finite number of sampled data\npoints (measurements) is a fundamental problem in science and engineering. This\nis often formulated as a minimum norm interpolation problem, regularized\nlearning problem or, in general, …
-
Wasserstein Cross-Lingual Alignment For Named Entity Recognition
2022 · ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
Supervised training of Named Entity Recognition (NER) models generally require large amounts of annotations, which are hardly available for less widely used (low resource) languages, e.g., Armenian and Dutch. Therefore, it will be desirable if …
-
Amer: A New Attribute-Missing Network Embedding Approach
2022 · IEEE Transactions on Cybernetics
Network embedding which aims to learn a low dimensional representation of nodes is a powerful technique for network analysis. While network embedding for networks with complete attributes has been widely investigated, in many real-world applications …
-
Understanding Chinese secondary school students’ perceptions of mobile-assisted language learning
2022 · Interactive Learning Environments
Given the paucity of research on mobile-assisted language learning (MALL) in secondary schools in China, this retrospective case study explored the psychological processes underlying the non-voluntary MALL experiences of Chinese secondary school students during a …
-
LightHuBERT: Lightweight and Configurable Speech Representation Learning with Once-for-All Hidden-Unit BERT
2022 · arXiv (Cornell University)
Self-supervised speech representation learning has shown promising results in various speech processing tasks. However, the pre-trained models, e.g., HuBERT, are storage-intensive Transformers, limiting their scope of applications under low-resource settings. To this end, we propose …
-
Nearest Neighbor Machine Translation is Meta-Optimizer on Output Projection Layer
2023 · arXiv (Cornell University)
Nearest Neighbor Machine Translation ($k$NN-MT) has achieved great success in domain adaptation tasks by integrating pre-trained Neural Machine Translation (NMT) models with domain-specific token-level retrieval. However, the reasons underlying its success have not been thoroughly …
-
A Multisubobject Approach to Dynamic Formation Target Tracking Using Random Matrices
2023 · IEEE Transactions on Aerospace and Electronic Systems
Bird flocks are typical group targets with various linear formations and high dynamics due to swarm intelligence. This leads to several problems in traditional multisubobject group target tracking such as shape model mismatch and false …
-
What are Public Concerns about ChatGPT? A Novel Self-Supervised Neural Topic Model Tells You
2023 · arXiv (Cornell University)
The recently released artificial intelligence conversational agent, ChatGPT, has gained significant attention in academia and real life. A multitude of early ChatGPT users eagerly explore its capabilities and share their opinions on it via social …
-
Rethinking Word-Level Auto-Completion in Computer-Aided Translation
2023 · arXiv (Cornell University)
Word-Level Auto-Completion (WLAC) plays a crucial role in Computer-Assisted Translation. It aims at providing word-level auto-completion suggestions for human translators. While previous studies have primarily focused on designing complex model architectures, this paper takes a …
-
UniMS-RAG: A Unified Multi-source Retrieval-Augmented Generation for Personalized Dialogue Systems
2024 · arXiv (Cornell University)
Large Language Models (LLMs) has shown exceptional capabilities in many natual language understanding and generation tasks. However, the personalization issue still remains a much-coveted property, especially when it comes to the multiple sources involved in …
-
Logit Standardization in Knowledge Distillation
2024 · arXiv (Cornell University)
Knowledge distillation involves transferring soft labels from a teacher to a student using a shared temperature-based softmax function. However, the assumption of a shared temperature between teacher and student implies a mandatory exact match between …
-
Three-dimensional visualization of thyroid ultrasound images based on multi-scale features fusion and hierarchical attention
2024 · BioMedical Engineering OnLine
BACKGROUND: Ultrasound three-dimensional visualization, a cutting-edge technology in medical imaging, enhances diagnostic accuracy by providing a more comprehensive and readable portrayal of anatomical structures compared to traditional two-dimensional ultrasound. Crucial to this visualization is the …
-
Short-Term Photovoltaic System Output Power Prediction Based on Integrated Deep Learning Algorithms in the Clean Energy Sector
2024 · International Journal of e-Collaboration
Photovoltaic power generation system plays an important role in renewable energy. Therefore, accurately predicting the short-term output power of photovoltaic system has become a key challenge for real-time power grid management. This study focuses on …
-
Personalized Federated Learning for Text Classification with Gradient-Free Prompt Tuning
2024
In this paper, we study personalized federated learning for text classification with Pretrained Language Models (PLMs).We identify two challenges in efficiently leveraging PLMs for personalized federated learning: 1) Communication.PLMs are usually large in size, inducing …
-
Improving Machine Translation with Human Feedback: An Exploration of Quality Estimation as a Reward Model
2024
Zhiwei He, Xing Wang, Wenxiang Jiao, Zhuosheng Zhang, Rui Wang, Shuming Shi, Zhaopeng Tu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: …
-
Study on the Effect of Lightning Wave and Its Characteristic Parameters on the Level of Lightning Resistance of Transmission Lines
2024
The traditional method of calculating the lightning resistance level of transmission lines ignores the influence of lightning current waveform and characteristic parameters on the lightning resistance level, resulting in poor accuracy of the calculation results. …
-
A Correlation Manifold Self-Attention Network for EEG Decoding
2025
Riemannian neural networks, which generalize the deep learning paradigm to non-Euclidean geometries, have garnered widespread attention across diverse applications in artificial intelligence. Among these, the representative attention models have been studied on various non-Euclidean spaces …
-
AndesVL Technical Report: An Efficient Mobile-side Multimodal Large Language Model
2025 · arXiv (Cornell University)
In recent years, while cloud-based MLLMs such as QwenVL, InternVL, GPT-4o, Gemini, and Claude Sonnet have demonstrated outstanding performance with enormous model sizes reaching hundreds of billions of parameters, they significantly surpass the limitations in …
-
CachePrune: Teaching LLMs What Not to Follow via KV-Cache Editing
2026
Rui Wang, Junda Wu, Yu Xia, Tong Yu, Ruiyi Zhang, Ryan A. Rossi, Subrata Mitra, Lina Yao, Julian McAuley. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). …
-
Sentence Embedding for Neural Machine Translation Domain Adaptation
2017
Although new corpora are becoming increasingly available for machine translation, only those that belong to the same or similar domains are typically able to improve translation performance. Recently Neural Machine Translation (NMT) has become prominent …
-
Syntax-Directed Attention for Neural Machine Translation
2018 · Proceedings of the AAAI Conference on Artificial Intelligence
Attention mechanism, including global attention and local attention, plays a key role in neural machine translation (NMT). Global attention attends to all source words for word prediction. In comparison, local attention selectively looks at fixed-window …
-
A Survey of Domain Adaptation for Neural Machine Translation
2018 · arXiv (Cornell University)
Neural machine translation (NMT) is a deep learning based approach for machine translation, which yields the state-of-the-art translation performance in scenarios where large-scale parallel corpora are available. Although the high-quality and domain-specific translation is crucial …
-
SG-Net: Syntax-Guided Machine Reading Comprehension
2020 · Proceedings of the AAAI Conference on Artificial Intelligence
For machine reading comprehension, the capacity of effectively modeling the linguistic knowledge from the detail-riddled and lengthy passages and getting ride of the noises is essential to improve its performance. Traditional attentive models attend to …