Wei Xu
26 papers in the PaperMetrix corpus
Papers by this author
-
Application of K-Means Cluster and Rough Set in Classified Real-Time Flood Forecasting
2015 · Advanced materials research
A new classified real-time flood forecasting framework was presented. Firstly, the historical floods were classified by K-means cluster, according to the hydrological factors. Then rough set was used to extract operation rules for flood forecasting. …
-
Shared Tasks of the 2015 Workshop on Noisy User-generated Text: Twitter Lexical Normalization and Named Entity Recognition
2015 · The Association for Computational Linguistics
This paper presents the results of the two shared tasks associated with W-NUT 2015: (1) a text normalization task with 10 participants; and (2) a named entity tagging task with 8 participants. We outline the …
-
High-speed railway signal system information security analysis and countermeasures
2017
For a long time, the railway signal system adopts a dedicated closed network without problems such as external intrusion and network virus transmission. However, because of the high-speed and real-time data sharing requirements between the …
-
Research and realization of rapid development architecture based on WeChat platform
2017 · 2017 IEEE 2nd Information Technology, Networking, Electronic and Automation Control Conference (ITNEC)
With the rapid development of WeChat, there are more and more the applications of WeChat public platform, but now developing efficiency is low. After studying the related technology of WeChat development, this paper puts forward …
-
PrivPy: Enabling Scalable and General Privacy-Preserving Machine Learning
2018 · arXiv (Cornell University)
We introduce PrivPy, a practical privacy-preserving collaborative computation framework, especially optimized for machine learning tasks. PrivPy provides an easy-to-use and highly compatible Python programming front-end which supports high-level array operations and different secure computation engines …
-
PrivPy
2019
Privacy is a big hurdle for collaborative data mining across multiple parties. We present multi-party computation (MPC) framework designed for large-scale data mining tasks. PrivPy combines an easy-to-use and highly flexible Python programming interface with …
-
Character-Based Neural Networks for Sentence Pair Modeling
2018
Sentence pair modeling is critical for many NLP tasks, such as paraphrase identification, semantic textual similarity, and natural language inference. Most state-of-the-art neural models for these tasks rely on pretrained word embedding and compose sentence-level …
-
CFO: Conditional Focused Neural Question Answering with Large-scale Knowledge Bases
2016
How can we enable computers to automatically answer questions like "Who created the character Harry Potter"? Carefully built knowledge bases provide rich sources of facts. However, it remains a challenge to answer factoid questions raised …
-
Semi-Supervised Learning for Neural Machine Translation
2016 · arXiv (Cornell University)
While end-to-end neural machine translation (NMT) has made remarkable progress recently, NMT systems only rely on parallel corpora for parameter estimation. Since parallel corpora are usually limited in quantity, quality, and coverage, especially for low-resource …
-
Meta-Tuning LLMs to Leverage Lexical Knowledge for Generalizable Language Style Understanding
2023 · arXiv (Cornell University)
Language style is often used by writers to convey their intentions, identities, and mastery of language. In this paper, we show that current large language models struggle to capture some language styles without fine-tuning. To …
-
Thresh: A Unified, Customizable and Deployable Platform for Fine-Grained Text Evaluation
2023 · arXiv (Cornell University)
Fine-grained, span-level human evaluation has emerged as a reliable and robust method for evaluating text generation tasks such as summarization, simplification, machine translation and news generation, and the derived annotations have been useful for training …
-
Deep CSI Compression for Dual-Polarized Massive MIMO Channels with Disentangled Representation Learning
2024 · arXiv (Cornell University)
Channel state information (CSI) feedback is critical for achieving the promised advantages of enhancing spectral and energy efficiencies in massive multiple-input multiple-output (MIMO) wireless communication systems. Deep learning (DL)-based methods have been proven effective in …
-
Leveraging Large Language Models to Enhance Personalized Recommendations in E-commerce
2024 · arXiv (Cornell University)
This study deeply explores the application of large language model (LLM) in personalized recommendation system of e-commerce. Aiming at the limitations of traditional recommendation algorithms in processing large-scale and multi-dimensional data, a recommendation system framework …
-
A Multi-Task Based Clustering Personalized Federated Learning Method
2024 · Big Data Mining and Analytics
Federated Learning (FL) is a framework for machine learning on a large-scale distributed dataset, enabling the training of a collaborative model across multiple parties while preserving the privacy of user data. However, in cases where …
-
Ultrafast & Low-power Consumption 2D Floating-Gate Devices for Opto-electronic Hybrid Neural Networks
2025
Optoelectronic hybrid neural networks combine the advantages of electronical and optical neural networks, enabling next-generation neuromorphic computing systems with nanosecond processing speeds, fJ-level energy efficiency, and wafer-scale integration density. Here, we demonstrate a MoS2/h-BN/graphene based …
-
Task-Oriented Low-Label Semantic Communication With Self-Supervised Learning
2025 · IEEE Transactions on Wireless Communications
Task-oriented semantic communication enhances transmission efficiency by conveying semantic information rather than exact messages. Deep learning (DL)-based semantic communication can effectively cultivate the essential semantic knowledge for semantic extraction, transmission, and interpretation by leveraging massive …
-
A Generalized Bisimulation Metric of State Similarity between Markov Decision Processes: From Theoretical Propositions to Applications
2025
The bisimulation metric (BSM) is a powerful tool for computing state similarities within a Markov decision process (MDP), revealing that states closer in BSM have more similar optimal value functions. While BSM has been successfully …
-
SemEval-2015 Task 1: Paraphrase and Semantic Similarity in Twitter (PIT)
2015
In this shared task, we present evaluations on two related tasks Paraphrase Identification (PI) and Semantic Textual Similarity (SS) systems for the Twitter data. Given a pair of sentences, participants are asked to produce a …
-
End-to-end learning of semantic role labeling using recurrent neural networks
2015
Jie Zhou, Wei Xu. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
-
Optimizing Statistical Machine Translation for Text Simplification
2016 · Transactions of the Association for Computational Linguistics
Most recent sentence simplification systems use basic machine translation models to learn lexical and syntactic paraphrases from a manually simplified parallel corpus. These methods are limited by the quality and quantity of manually simplified corpora, …
-
Results of the WNUT16 Named Entity Recognition Shared Task
2016 · Digital Access to Libraries
This paper presents the results of the Twitter Named Entity Recognition shared task associated with W-NUT 2016: a named entity tagging task with 10 teams participating. We outline the shared task, annotation process and dataset …
-
Neural Network Models for Paraphrase Identification, Semantic Textual Similarity, Natural Language Inference, and Question Answering
2018 · International Conference on Computational Linguistics
In this paper, we analyze several neural network designs (and their variations) for sentence pair modeling and compare their performance extensively across eight datasets, including paraphrase identification, semantic textual similarity, natural language inference, and question …
-
A Continuously Growing Dataset of Sentential Paraphrases
2017
A major challenge in paraphrase research is the lack of parallel corpora. In this paper, we present a new method to collect large-scale sentential paraphrases from Twitter by linking tweets through shared URLs. The main …
-
Deep Recurrent Models with Fast-Forward Connections for Neural Machine Translation
2016 · Transactions of the Association for Computational Linguistics
Neural machine translation (NMT) aims at solving machine translation (MT) problems using neural networks and has exhibited promising results in recent years. However, most of the existing NMT models are shallow and there is still …
-
An Empirical Study of Pre-trained Transformers for Arabic Information Extraction
2020
Multilingual pre-trained Transformers, such as mBERT However, their performance on Arabic information extraction (IE) tasks is not very well studied. In this paper, we pre-train a customized bilingual BERT, dubbed GigaBERT, that is designed specifically …
-
The GEM Benchmark: Natural Language Generation, its Evaluation and Metrics
2021
Sebastian Gehrmann, Tosin Adewumi, Karmanya Aggarwal, Pawan Sasanka Ammanamanchi, Anuoluwapo Aremu, Antoine Bosselut, Khyathi Raghavi Chandu, Miruna-Adriana Clinciu, Dipanjan Das, Kaustubh Dhole, Wanyu Du, Esin Durmus, Ondřej Dušek, Chris Chinenye Emezue, Varun Gangal, Cristina Garbacea, …