Xiaojun Wan
16 ورقة في مجموعة PaperMetrix
أوراق هذا المؤلف
-
User Embedding for Scholarly Microblog Recommendation
2016 · Annual Meeting of the Association for Computational Linguistics
Nowadays, many scholarly messages are posted on Chinese microblogs and more and more researchers tend to find scholarly information on microblogs. In order to exploit microblogging to benefit scientific research, we propose a scholarly microblog …
-
SentiGAN: Generating Sentimental Texts via Mixture Adversarial Networks
2018
Generating texts of different sentiment labels is getting more and more attention in the area of natural language generation. Recently, Generative Adversarial Net (GAN) has shown promising results in text generation. However, the texts generated …
-
Adapting Neural Single-Document Summarization Model for Abstractive Multi-Document Summarization: A Pilot Study
2018
Till now, neural abstractive summarization methods have achieved great success for single document summarization (SDS). However, due to the lack of large scale multi-document summaries, such methods can be hardly applied to multi-document summarization (MDS). …
-
Controllable Unsupervised Text Attribute Transfer via Editing Entangled Latent Representation
2019 · arXiv (Cornell University)
Unsupervised text attribute transfer automatically transforms a text to alter a specific attribute (e.g. sentiment) without using any parallel data, while simultaneously preserving its attribute-independent content. The dominant approaches are trying to model the content-independent …
-
IGSQL: Database Schema Interaction Graph Based Neural Model for Context-Dependent Text-to-SQL Generation
2020
Context-dependent text-to-SQL task has drawn much attention in recent years. Previous models on context-dependent text-to-SQL task only concentrate on utilizing historical user inputs. In this work, in addition to using encoders to capture historical information …
-
DivGAN: Towards Diverse Paraphrase Generation via Diversified Generative Adversarial Network
2020
Paraphrases refer to texts that convey the same meaning with different expression forms. Traditional seq2seq-based models on paraphrase generation mainly focus on the fidelity while ignoring the diversity of outputs. In this paper, we propose …
-
Comparing Knowledge-Intensive and Data-Intensive Models for English Resource Semantic Parsing
2021 · Computational Linguistics
Abstract In this work, we present a phenomenon-oriented comparative analysis of the two dominant approaches in English Resource Semantic (ERS) parsing: classic, knowledge-intensive and neural, data-intensive models. To reflect state-of-the-art neural NLP technologies, a factorization-based …
-
WIND: Weighting Instances Differentially for Model-Agnostic Domain Adaptation
2021
Domain Adaptation is a fundamental problem in machine learning and natural language processing. In this paper, we study the domain adaptation problem from the perspective of instance weighting. Conventional instance weighting approaches cannot learn the …
-
BiRdQA: A Bilingual Dataset for Question Answering on Tricky Riddles
2022 · Proceedings of the AAAI Conference on Artificial Intelligence
A riddle is a question or statement with double or veiled meanings, followed by an unexpected answer. Solving riddle is a challenging task for both machine and human, testing the capability of understanding figurative, creative …
-
A New Dataset and Empirical Study for Sentence Simplification in Chinese
2023
Sentence Simplification is a valuable technique that can benefit language learners and children a lot. However, current research focuses more on English sentence simplification. The development of Chinese sentence simplification is relatively slow due to …
-
New Datasets and Controllable Iterative Data Augmentation Method for Code-switching ASR Error Correction
2023
With the wide use of automatic speech recognition(ASR) systems, researchers pay more attention to the ASR error correction task to improve the quality of recognition results. In particular, ASR in bilingual or multilingual settings, namely …
-
Automated Similarity Metric Generation for Recommendation
2024 · arXiv (Cornell University)
The embedding-based architecture has become the dominant approach in modern recommender systems, mapping users and items into a compact vector space. It then employs predefined similarity metrics, such as the inner product, to calculate similarity …
-
Better than Random: Reliable NLG Human Evaluation with Constrained Active Sampling
2024 · arXiv (Cornell University)
Human evaluation is viewed as a reliable evaluation method for NLG which is expensive and time-consuming. To save labor and costs, researchers usually perform human evaluation on a small subset of data sampled from the …
-
Exploring the Multilingual NLG Evaluation Abilities of LLM-Based Evaluators
2025 · arXiv (Cornell University)
Previous research has shown that LLMs have potential in multilingual NLG evaluation tasks. However, existing research has not fully explored the differences in the evaluation capabilities of LLMs across different languages. To this end, this …
-
Accurate SHRG-Based Semantic Parsing
2018
We demonstrate that an SHRG-based parser can produce semantic graphs much more accurately than previously shown, by relating synchronous production rules to the syntacto-semantic composition process. Our parser achieves an accuracy of 90.35 for EDS …
-
AMR-To-Text Generation with Graph Transformer
2020 · Transactions of the Association for Computational Linguistics
Abstract meaning representation (AMR)-to-text generation is the challenging task of generating natural language texts from AMR graphs, where nodes represent concepts and edges denote relations. The current state-of-the-art methods use graph-to-sequence models; however, they still …