Researcher profile

Xiaojun Wan

16 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. User Embedding for Scholarly Microblog Recommendation

    2016 · Annual Meeting of the Association for Computational Linguistics

    Nowadays, many scholarly messages are posted on Chinese microblogs and more and more researchers tend to find scholarly information on microblogs. In order to exploit microblogging to benefit scientific research, we propose a scholarly microblog …

  2. SentiGAN: Generating Sentimental Texts via Mixture Adversarial Networks

    2018

    Generating texts of different sentiment labels is getting more and more attention in the area of natural language generation. Recently, Generative Adversarial Net (GAN) has shown promising results in text generation. However, the texts generated …

  3. Adapting Neural Single-Document Summarization Model for Abstractive Multi-Document Summarization: A Pilot Study

    2018

    Till now, neural abstractive summarization methods have achieved great success for single document summarization (SDS). However, due to the lack of large scale multi-document summaries, such methods can be hardly applied to multi-document summarization (MDS). …

  4. Controllable Unsupervised Text Attribute Transfer via Editing Entangled Latent Representation

    2019 · arXiv (Cornell University)

    Unsupervised text attribute transfer automatically transforms a text to alter a specific attribute (e.g. sentiment) without using any parallel data, while simultaneously preserving its attribute-independent content. The dominant approaches are trying to model the content-independent …

  5. IGSQL: Database Schema Interaction Graph Based Neural Model for Context-Dependent Text-to-SQL Generation

    2020

    Context-dependent text-to-SQL task has drawn much attention in recent years. Previous models on context-dependent text-to-SQL task only concentrate on utilizing historical user inputs. In this work, in addition to using encoders to capture historical information …

  6. DivGAN: Towards Diverse Paraphrase Generation via Diversified Generative Adversarial Network

    2020

    Paraphrases refer to texts that convey the same meaning with different expression forms. Traditional seq2seq-based models on paraphrase generation mainly focus on the fidelity while ignoring the diversity of outputs. In this paper, we propose …

  7. Comparing Knowledge-Intensive and Data-Intensive Models for English Resource Semantic Parsing

    2021 · Computational Linguistics

    Abstract In this work, we present a phenomenon-oriented comparative analysis of the two dominant approaches in English Resource Semantic (ERS) parsing: classic, knowledge-intensive and neural, data-intensive models. To reflect state-of-the-art neural NLP technologies, a factorization-based …

  8. WIND: Weighting Instances Differentially for Model-Agnostic Domain Adaptation

    2021

    Domain Adaptation is a fundamental problem in machine learning and natural language processing. In this paper, we study the domain adaptation problem from the perspective of instance weighting. Conventional instance weighting approaches cannot learn the …

  9. BiRdQA: A Bilingual Dataset for Question Answering on Tricky Riddles

    2022 · Proceedings of the AAAI Conference on Artificial Intelligence

    A riddle is a question or statement with double or veiled meanings, followed by an unexpected answer. Solving riddle is a challenging task for both machine and human, testing the capability of understanding figurative, creative …

  10. A New Dataset and Empirical Study for Sentence Simplification in Chinese

    2023

    Sentence Simplification is a valuable technique that can benefit language learners and children a lot. However, current research focuses more on English sentence simplification. The development of Chinese sentence simplification is relatively slow due to …

  11. New Datasets and Controllable Iterative Data Augmentation Method for Code-switching ASR Error Correction

    2023

    With the wide use of automatic speech recognition(ASR) systems, researchers pay more attention to the ASR error correction task to improve the quality of recognition results. In particular, ASR in bilingual or multilingual settings, namely …

  12. Automated Similarity Metric Generation for Recommendation

    2024 · arXiv (Cornell University)

    The embedding-based architecture has become the dominant approach in modern recommender systems, mapping users and items into a compact vector space. It then employs predefined similarity metrics, such as the inner product, to calculate similarity …

  13. Better than Random: Reliable NLG Human Evaluation with Constrained Active Sampling

    2024 · arXiv (Cornell University)

    Human evaluation is viewed as a reliable evaluation method for NLG which is expensive and time-consuming. To save labor and costs, researchers usually perform human evaluation on a small subset of data sampled from the …

  14. Exploring the Multilingual NLG Evaluation Abilities of LLM-Based Evaluators

    2025 · arXiv (Cornell University)

    Previous research has shown that LLMs have potential in multilingual NLG evaluation tasks. However, existing research has not fully explored the differences in the evaluation capabilities of LLMs across different languages. To this end, this …

  15. Accurate SHRG-Based Semantic Parsing

    2018

    We demonstrate that an SHRG-based parser can produce semantic graphs much more accurately than previously shown, by relating synchronous production rules to the syntacto-semantic composition process. Our parser achieves an accuracy of 90.35 for EDS …

  16. AMR-To-Text Generation with Graph Transformer

    2020 · Transactions of the Association for Computational Linguistics

    Abstract meaning representation (AMR)-to-text generation is the challenging task of generating natural language texts from AMR graphs, where nodes represent concepts and edges denote relations. The current state-of-the-art methods use graph-to-sequence models; however, they still …