ملف الباحث

Lu Wang

21 ورقة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Improving Agreement and Disagreement Identification in Online Discussions with A Socially-Tuned Sentiment Lexicon

    2016 · arXiv (Cornell University)

    We study the problem of agreement and disagreement detection in online discussions. An isotonic Conditional Random Fields (isotonic CRF) based sequential model is proposed to make predictions on sentence- or segment-level. We automatically construct a …

  2. Leveraging Semantic Web Search and Browse Sessions for Multi-Turn Spoken Dialog Systems

    2016 · arXiv (Cornell University)

    Training statistical dialog models in spoken dialog systems (SDS) requires large amounts of annotated data. The lack of scalable methods for data mining and annotation poses a significant hurdle for state-of-the-art statistical dialog managers. This …

  3. Summary of Association Rules

    2019 · IOP Conference Series Earth and Environmental Science

    In recent years, scholars at home and abroad have conducted a large number of researches on association rules, in order to deeply understand the mining technology of association rules, and master its research status and …

  4. Power Saving and Secure Text Input for Commodity Smart Watches

    2020 · IEEE Transactions on Mobile Computing

    Smart wristband has become a dominant device in the wearable ecosystem, providing versatile functions such as fitness tracking, mobile payment, and transport ticketing. However, the small form-factor, low-profile hardware interfaces and computational resources limit their …

  5. Spanning Attack: Reinforce Black-box Attacks with Unlabeled Data

    2020 · arXiv (Cornell University)

    Adversarial black-box attacks aim to craft adversarial perturbations by querying input-output pairs of machine learning models. They are widely used to evaluate the robustness of pre-trained models. However, black-box attacks often suffer from the issue …

  6. Knowledge Graph-Augmented Abstractive Summarization with Semantic-Driven Cloze Reward

    2020 · arXiv (Cornell University)

    Sequence-to-sequence models for abstractive summarization have been studied extensively, yet the generated summaries commonly suffer from fabricated content, and are often found to be near-extractive. We argue that, to address these issues, the summarizer should …

  7. Scalable Uncertainty Quantification via GenerativeBootstrap Sampler

    2020 · arXiv (Cornell University)

    It has been believed that the virtue of using statistical procedures is on uncertainty quantification in statistical decisions, and the bootstrap method has been commonly used for this purpose. However, nowadays as the size of …

  8. Efficient Argument Structure Extraction with Transfer Learning and Active Learning

    2022 · Zenodo (CERN European Organization for Nuclear Research)

    This repository contains AMPERE++, the dataset associated with the paper “Efficient Argument Structure Extraction with Transfer Learning and Active Learning. Xinyu Hua and Lu Wang, Findings of ACL 2022”. It contains 400 academic peer reviews …

  9. Parallel Ranking of Ads and Creatives in Real-Time Advertising Systems

    2023 · arXiv (Cornell University)

    "Creativity is the heart and soul of advertising services". Effective creatives can create a win-win scenario: advertisers can reach target users and achieve marketing objectives more effectively, users can more quickly find products of interest, …

  10. A Mallows-like Criterion for Anomaly Detection with Random Forest Implementation

    2024 · arXiv (Cornell University)

    The effectiveness of anomaly signal detection can be significantly undermined by the inherent uncertainty of relying on one specified model. Under the framework of model average methods, this paper proposes a novel criterion to select …

  11. Surface Anomaly Detection and Localization with Diffusion-based Reconstruction

    2024

    Surface anomaly detection aims to detect and locate anomalies in the product surface images. Anomalies are mostly rare and difficult to find, so the unsupervised method which trained using anomaly-free training samples is popular. One …

  12. On Many-Shot In-Context Learning for Long-Context Evaluation

    2024 · arXiv (Cornell University)

    Many-shot in-context learning (ICL) has emerged as a unique setup to both utilize and test the ability of large language models to handle long context. This paper delves into long-context language model (LCLM) evaluation through …

  13. Research on Interpretable Recommendation Methods for Civil Aviation Ancillary Services

    2024

    In the modern civil aviation industry, with the escalating market competition and the increasing diversification of airline profit models, revenue from ancillary services has become a key factor for airlines to enhance their overall revenue …

  14. Efficient Ensemble for Fine-tuning Language Models on Multiple Datasets

    2025

    This paper develops an ensemble method for fine-tuning a language model to multiple datasets.Existing methods, such as quantized LoRA (QLoRA), are efficient when adapting to a single dataset.When training on multiple datasets of different tasks, …

  15. Thyroid pathology image classification via multi-scale feature fusion and multi-instance learning

    2025 · Diagnostic Pathology

    BACKGROUND: The global incidence of thyroid cancer has significantly increased, while traditional pathological diagnosis remains time-consuming and expert-dependent. This study develops an auxiliary diagnostic tool designed to reduce the workload of pathologists and improve diagnostic …

  16. Logit Arithmetic Elicits Long Reasoning Capabilities Without Training

    2025 · arXiv (Cornell University)

    Large reasoning models exhibit long chain-of-thought reasoning with complex strategies such as backtracking and self-verification. Yet, these capabilities typically require resource-intensive post-training. We investigate whether such behaviors can be elicited in large models without any …

  17. Argument Mining for Understanding Peer Reviews

    2019

    Xinyu Hua, Mitko Nikolov, Nikhil Badugu, Lu Wang. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). 2019.

  18. BIGPATENT: A Large-Scale Dataset for Abstractive and Coherent Summarization

    2019

    Most existing text summarization datasets are compiled from the news domain, where summaries have a flattened discourse structure. In such datasets, summary-worthy content often appears in the beginning of input articles. Moreover, large segments from …

  19. Semi-Supervised Learning for Neural Keyphrase Generation

    2018

    We study the problem of generating keyphrases that summarize the key points for a given document. While sequence-to-sequence (seq2seq) models have achieved remarkable performance on this task In this paper, we propose semi-supervised keyphrase generation …

  20. Neural Network-Based Abstract Generation for Opinions and Arguments

    2016

    We study the problem of generating abstractive summaries for opinionated text. We propose an attention-based neural network model that is able to absorb information from multiple text units to construct informative, concise, and fluent summaries. …

  21. Efficient Attentions for Long Document Summarization

    2021

    Luyang Huang, Shuyang Cao, Nikolaus Parulian, Heng Ji, Lu Wang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.