Researcher profile

Yang Liu

84 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Learning Cross-lingual Word Embeddings via Matrix Co-factorization

    2015

    Tianze Shi, Zhiyuan Liu, Yang Liu, Maosong Sun. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 2: Short Papers). 2015.

  2. Generalized Agreement for Bidirectional Word Alignment

    2015

    While agreement-based joint training has proven to deliver state-of-the-art alignment accuracy, the produced word alignments are usually restricted to one-toone mappings because of the hard constraint on agreement. We propose a general framework to allow …

  3. Construction of Smart Campus Based on Situational Awareness in the Era Of Big Data

    2016

    Big data as a new data management technology, based on the current Internet of things and cloud computing "smart campus" has an important role in the construction. "Smart campus" as an information construction of educational …

  4. Particle flow for sequential Monte Carlo implementation of probability hypothesis density

    2017

    Target tracking is a challenging task and generally no analytical solution is available, especially for the multi-target tracking systems. To address this problem, probability hypothesis density (PHD) filter is used by propagating the PHD instead …

  5. Brand key asset discovery via cluster-wise biased discriminant projection

    2017 · Proceedings of the International Conference on Web Intelligence

    Accurate and effective discovery of a brand's key assets, namely, Key Opinion Leaders (KOLs) and potential customers, plays an essential role in marketing campaigns. In a massive online social network, brands are challenged with identifying …

  6. A hierarchical classification approach for tor anonymous traffic

    2017

    Tor is an anonymous communication system that can protect our privacy, but it also provides a haven for criminals to avoid network tracing. Therefore, anonymous traffic analysis and classification is an important part of maintaining …

  7. Radar signal sorting algorithm of k-means clustering based on data field

    2017

    Radar signal sorting is one of the essential technologies in radar countermeasures reconnaissance system. Non-cooperative radar signal sorting without prior information has been a great challenge for radar countermeasures. This paper presents a k-means clustering …

  8. A prototype simulator for the simulation of complicated hadoop framework behaviors

    2017

    Distributed computing and parallel computing have become the most effective tools for solving complex problems in amount of academia and industrial fields. Among a number of distributed and parallel computing technologies, MapReduce has been proved …

  9. ReCDroid: Automatically Reproducing Android Application Crashes from Bug Reports

    2019

    The large demand of mobile devices creates significant concerns about the quality of mobile applications (apps). Developers heavily rely on bug reports in issue tracking systems to reproduce failures (e.g., crashes). However, the process of …

  10. A Multi-Goal Oriented Approach for Adaptation Rules Generation

    2018

    Modern software runs in a dynamic, uncertain environment, and should satisfy multiple goals simultaneously. In order to allow software to respond to changes in the environment or user requirements and meet user goals continuously, an …

  11. Experimental test of error-disturbance uncertainty relation with continuous variables

    2019 · arXiv (Cornell University)

    Uncertainty relation is one of the fundamental principle in quantum mechanics and plays an important role in quantum information science. We experimentally test the error-disturbance uncertainty relation (EDR) with continuous variables for Gaussian states. Two …

  12. aLeak: Context-Free Side-Channel from Your Smart Watch Leaks Your Typing Privacy

    2019 · IEEE Transactions on Mobile Computing

    We revisit a crucial privacy problem in this paper - can the sensitive information, like the numeric passwords and personal data, frequently typed by user on mobile devices be inferred through the motion sensors of …

  13. Multi-info Fusion Based Video Recommendation System

    2019 · Journal of Physics Conference Series

    Abstract The great progress in recommendation system help users discover more interesting items that satisfy their appetites. Considering the video recommendation is an increasing popular sub-field of recommendation, but the traditional recommendation techniques such as …

  14. A Countermeasure Against Statistical Ineffective Fault Analysis

    2020 · IEEE Transactions on Circuits & Systems II Express Briefs

    Current state-of-the-art countermeasures against Fault Injection Attacks (FIA) provide good protection against analysis methods that require the differences in the correct and faulty ciphertext to derive the secret information, such as Differential Fault Analysis (DFA) …

  15. Abnormal Client Behavior Detection in Federated Learning

    2019 · arXiv (Cornell University)

    In federated learning systems, clients are autonomous in that their behaviors are not fully governed by the server. Consequently, a client may intentionally or unintentionally deviate from the prescribed course of federated model training, resulting …

  16. Stealing Deep Reinforcement Learning Models for Fun and Profit

    2020 · arXiv (Cornell University)

    This paper presents the first model extraction attack against Deep Reinforcement Learning (DRL), which enables an external adversary to precisely recover a black-box DRL model only from its interaction with the environment. Model extraction attacks …

  17. Generating Behavior-Diverse Game AIs with Evolutionary Multi-Objective Deep Reinforcement Learning

    2020

    Generating diverse behaviors for game artificial intelligence (Game AI) has been long recognized as a challenging task in the game industry. Designing a Game AI with a satisfying behavioral characteristic (style) heavily depends on the …

  18. Shortened Linear Codes over Finite Fields

    2020 · arXiv (Cornell University)

    The puncturing and shortening technique are two important approaches to constructing new linear codes from old ones. In the past 70 years, a lot of progress on the puncturing technique has been made, and many …

  19. BatchCrypt: Efficient homomorphic encryption for cross-silo federated learning

    2020 · Rare & Special e-Zone (The Hong Kong University of Science and Technology)

    Cross-silo federated learning (FL) enables organizations (e.g., financial or medical) to collaboratively train a machine learning model by aggregating local gradient updates from each client without sharing privacy-sensitive data. To ensure no update is revealed …

  20. Learning to Expand: Reinforced Pseudo-relevance Feedback Selection for Information-seeking Conversations

    2020 · arXiv (Cornell University)

    Information-seeking conversation systems are increasingly popular in real-world applications, especially for e-commerce companies. To retrieve appropriate responses for users, it is necessary to compute the matching degrees between candidate responses and users' queries with historical …

  21. Neural Machine Translation With Explicit Phrase Alignment

    2021 · IEEE/ACM Transactions on Audio Speech and Language Processing

    While neural machine translation has achieved state-of-the-art translation performance, it is unable to capture the alignment between the input and output during the translation process. The lack of alignment in neural machine translation models leads …

  22. Linear Classifiers that Encourage Constructive Adaptation

    2020 · arXiv (Cornell University)

    Machine learning systems are often used in settings where individuals adapt their features to obtain a desired outcome. In such settings, strategic behavior leads to a sharp loss in model performance in deployment. In this …

  23. Synthetic Benchmarks for Scientific Research in Explainable Machine Learning

    2021 · arXiv (Cornell University)

    As machine learning models grow more complex and their applications become more high-stakes, tools for explaining model predictions have become increasingly important. This has spurred a flurry of research in model explainability and has given …

  24. Optimization and Simulation of Labor Resource Management Information Platform Based on Internet of Things

    2021 · Wireless Communications and Mobile Computing

    This paper conducts an in‐depth analysis and research on the optimization of the labor resource management information platform through the Internet of Things (IoT) technology; through the collection, classification, and data search functions of this …

  25. Fine-tuning Is Not Enough: A Simple yet Effective Watermark Removal Attack for DNN Models

    2021

    Watermarking has become the tendency in protecting the intellectual property of DNN models. Recent works, from the adversary's perspective, attempted to subvert watermarking mechanisms by designing watermark removal attacks. However, these attacks mainly adopted sophisticated …

  26. A Large-Scale Empirical Study of Real-Life Performance Issues in Open Source Projects

    2022 · Zenodo (CERN European Organization for Nuclear Research)

    1. The spreadsheet "Perf Issue Empirical Data Package.xlsx" contains the details of data extraction and annotation of the performance issues. The three tabs in the above spreadsheet, i.e., “Java Projects Issues”, “Python Projects Issues”, and …

  27. Neighboring Backdoor Attacks on Graph Convolutional Network

    2022 · arXiv (Cornell University)

    Backdoor attacks have been widely studied to hide the misclassification rules in the normal models, which are only activated when the model is aware of the specific inputs (i.e., the trigger). However, despite their success …

  28. MAGAN: Mask Attention Generative Adversarial Network for Liver Tumor CT Image Synthesis

    2020 · Research Square (Research Square)

    Abstract Background : For deep learning, the size of the dataset greatly affects the final training effect. However, in the field of computer-aided diagnosis, medical image datasets are often limited and even scarce. Methods : …

  29. A Template-based Method for Constrained Neural Machine Translation

    2022 · arXiv (Cornell University)

    Machine translation systems are expected to cope with various types of constraints in many practical scenarios. While neural machine translation (NMT) has achieved strong performance in unconstrained cases, it is non-trivial to impose pre-specified constraints …

  30. Defense against Backdoor Attacks via Identifying and Purifying Bad Neurons

    2022 · arXiv (Cornell University)

    The opacity of neural networks leads their vulnerability to backdoor attacks, where hidden attention of infected neurons is triggered to override normal predictions to the attacker-chosen ones. In this paper, we propose a novel backdoor …

  31. Improving Radiology Summarization with Radiograph and Anatomy Prompts

    2022 · arXiv (Cornell University)

    The impression is crucial for the referring physicians to grasp key information since it is concluded from the findings and reasoning of radiologists. To alleviate the workload of radiologists and reduce repetitive human labor in …

  32. Semi-Supervised Learning for Neural Machine Translation

    2016 · arXiv (Cornell University)

    While end-to-end neural machine translation (NMT) has made remarkable progress recently, NMT systems only rely on parallel corpora for parameter estimation. Since parallel corpora are usually limited in quantity, quality, and coverage, especially for low-resource …

  33. Listening to Users' Voice: Automatic Summarization of Helpful App Reviews

    2022 · IEEE Transactions on Reliability

    App reviews are crowdsourcing knowledge of user experience with the apps, providing valuable information for app release planning, such as major bugs to fix and important features to add. There exist prior explorations on app …

  34. From "Law + Engineering" to Engineering Law: Current Situation, Goals and Value Implications

    2020 · JOURNAL OF ENGINEERING STUDIES

    The study of legal issues in the productivity development of engineering activities at its core is a response to the orderly needs and risk regulations of complex engineering fields. It is a realistic need for …

  35. Using In-Context Learning to Improve Dialogue Safety

    2023 · arXiv (Cornell University)

    While large neural-based conversational models have become increasingly proficient dialogue agents, recent work has highlighted safety issues with these systems. For example, these systems can be goaded into generating toxic content, which often perpetuates social …

  36. G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment

    2023 · arXiv (Cornell University)

    The quality of texts generated by natural language generation (NLG) systems is hard to measure automatically. Conventional reference-based metrics, such as BLEU and ROUGE, have been shown to have relatively low correlation with human judgments, …

  37. Casualty on the Titanic based on Machine Learning Methods

    2023 · Highlights in Science Engineering and Technology

    The Titanic sank on April 15, 1914, with 2224 people on board, and only 32% survived. The survivors are somewhat random, but they are somewhat the same. Studying the types of people who are more …

  38. MSN-net: Multi-Scale Normality Network for Video Anomaly Detection

    2023

    Existing unsupervised video anomaly detection methods often suffer from performance degradation due to the overgeneralization of deep models. In this paper, we propose a simple yet effective Multi-Scale Normality network (MSN-net) that uses hierarchical memories …

  39. Who is the Real Hero? Measuring Developer Contribution via Multi-dimensional Data Integration

    2023 · arXiv (Cornell University)

    Proper incentives are important for motivating developers in open-source communities, which is crucial for maintaining the development of open-source software healthy. To provide such incentives, an accurate and objective developer contribution measurement method is needed. …

  40. Privacy-Preserving Techniques in Cloud/Fog and Internet of Things

    2023 · Cryptography

    Recently, wireless networks have been developed using cloud infrastructure and software-based networks [...]

  41. ASTER: Automatic Speech Recognition System Accessibility Testing for Stutterers

    2023

    The popularity of automatic speech recognition (ASR) systems nowadays leads to an increasing need for improving their accessibility. Handling stuttering speech is an important feature for accessible ASR systems. To improve the accessibility of ASR …

  42. MERCY: Multiple Response Ranking Concurrently in Realistic Open-Domain Conversational Systems

    2023

    Automatic Evaluation (AE) and Response Selection (RS) models assign quality scores to various candidate responses and rank them in conversational setups. Prior response ranking research compares various models’ performance on synthetically generated test sets. In …

  43. Data-Efficient Alignment of Large Language Models with Human Feedback Through Natural Language

    2023 · arXiv (Cornell University)

    Learning from human feedback is a prominent technique to align the output of large language models (LLMs) with human expectations. Reinforcement learning from human feedback (RLHF) leverages human preference signals that are in the form …

  44. A balanced allocation of network teaching resources in higher vocational colleges based on demand prediction

    2023 · International Journal of Continuing Engineering Education and Life-Long Learning

    Because the traditional teaching resource allocation method has the problems of low accuracy of resource demand prediction and low balance of resource allocation, this paper studies a new balanced allocation method based on demand prediction. …

  45. Human-Instruction-Free LLM Self-Alignment with Limited Samples

    2024 · arXiv (Cornell University)

    Aligning large language models (LLMs) with human values is a vital task for LLM practitioners. Current alignment techniques have several limitations: (1) requiring a large amount of annotated data; (2) demanding heavy human involvement; (3) …

  46. BMLP: Behavior-aware MLP for Heterogeneous Sequential Recommendation

    2024 · arXiv (Cornell University)

    In real recommendation scenarios, users often have different types of behaviors, such as clicking and buying. Existing research methods show that it is possible to capture the heterogeneous interests of users through different types of …

  47. Datasets for Large Language Models: A Comprehensive Survey

    2024 · arXiv (Cornell University)

    This paper embarks on an exploration into the Large Language Model (LLM) datasets, which play a crucial role in the remarkable advancements of LLMs. The datasets serve as the foundational infrastructure analogous to a root …

  48. ToolRerank: Adaptive and Hierarchy-Aware Reranking for Tool Retrieval

    2024 · arXiv (Cornell University)

    Tool learning aims to extend the capabilities of large language models (LLMs) with external tools. A major challenge in tool learning is how to support a large number of tools, including unseen tools. To address …

  49. Adversarial Learning for Coordinate Regression Through k-Layer Penetrating Representation

    2024 · IEEE Transactions on Dependable and Secure Computing

    Adversarial attack is a crucial step when evaluating the reliability and robustness of deep neural networks (DNNs) models. Most existing attack approaches apply an end-to-end gradient update strategy to generate adversarial examples for a classification …

  50. Towards Enabling DPOAE Estimation on Single-Speaker Earbuds

    2024

    Distortion Product OtoAcoustic Emissions (DPOAEs) represents faint cochlear responses to dual-frequency stimuli, commonly employed in hearing screening. This paper introduces an innovative approach to trigger DPOAEs using single-speaker earbuds. Due to their compact size, the …

  51. Revisiting a Pain in the Neck: Semantic Phrase Processing Benchmark for Language Models

    2024 · arXiv (Cornell University)

    We introduce LexBench, a comprehensive evaluation suite enabled to test language models (LMs) on ten semantic phrase processing tasks. Unlike prior studies, it is the first work to propose a framework from the comparative perspective …

  52. Membership Inference on Text-to-Image Diffusion Models via Conditional Likelihood Discrepancy

    2024 · arXiv (Cornell University)

    Text-to-image diffusion models have achieved tremendous success in the field of controllable image generation, while also coming along with issues of privacy leakage and data copyrights. Membership inference arises in these contexts as a potential …

  53. Understanding the Effectiveness of Coverage Criteria for Large Language Models: A Special Angle from Jailbreak Attacks

    2024 · arXiv (Cornell University)

    Large language models (LLMs) have revolutionized artificial intelligence, but their increasing deployment across critical domains has raised concerns about their abnormal behaviors when faced with malicious attacks. Such vulnerability alerts the widespread inadequacy of pre-release …

  54. Bridging the Gap: Unpacking the Hidden Challenges in Knowledge Distillation for Online Ranking Systems

    2024

    Knowledge Distillation (KD) is a powerful approach for compressing a large model into a smaller, more efficient model, particularly beneficial for latency-sensitive applications like recommender systems. However, current KD research predominantly focuses on Computer Vision …

  55. SelfBC: Self Behavior Cloning for Offline Reinforcement Learning

    2024 · arXiv (Cornell University)

    Policy constraint methods in offline reinforcement learning employ additional regularization techniques to constrain the discrepancy between the learned policy and the offline dataset. However, these methods tend to result in overly conservative policies that resemble …

  56. On the Role of Attention Heads in Large Language Model Safety

    2024 · arXiv (Cornell University)

    Large language models (LLMs) achieve state-of-the-art performance on multiple language tasks, yet their safety guardrails can be circumvented, leading to harmful generations. In light of this, recent research on safety mechanisms has emerged, revealing that …

  57. Concept-Aware Graph Convolutional Network for Compositional Zero-Shot Learning

    2025 · IEEE Transactions on Neural Networks and Learning Systems

    Compositional zero-shot learning (CZSL) aims to identify unobservable compositional concepts with prior knowledge of known primitives (attributes and objects). Due to distribution differences between seen and unseen components, existing methods for CZSL often ignore intrinsic …

  58. Automated Runtime Verification of Security for E-Commerce Smart Contracts

    2025 · Journal of theoretical and applied electronic commerce research

    As a novel decentralized computing paradigm, blockchain is expected to disrupt the existing e-commerce architecture and process. Secure smart contracts are the crucial foundation for e-commerce based on blockchain. However, vulnerabilities in smart contracts occur …

  59. Incorporating Pre-Training Data Matters in Unsupervised Domain Adaptation

    2025 · IEEE Transactions on Pattern Analysis and Machine Intelligence

    In deep learning, initializing models with pre-trained weights has become the de facto practice for various downstream tasks. Many unsupervised domain adaptation (UDA) methods typically adopt a backbone pre-trained on ImageNet, and focus on reducing …

  60. SAP: Privacy-Preserving Fine-Tuning on Language Models with Split-and-Privatize Framework

    2025

    Pre-trained Language Models (PLM) have enabled a cost-effective approach to handling various downstream applications via Parameter-Efficient-Fine-Tuning (PEFT) techniques. In this context, service providers have introduced a popular fine-tuning-based product service known as Model-as-a-Service (MaaS). This …

  61. COSMIC: Generalized Refusal Direction Identification in LLM Activations

    2025 · arXiv (Cornell University)

    Large Language Models (LLMs) encode behaviors such as refusal within their activation space, yet identifying these behaviors remains a significant challenge. Existing methods often rely on predefined refusal templates detectable in output tokens or require …

  62. Dual-Priv Pruning : Efficient Differential Private Fine-Tuning in Multimodal Large Language Models

    2025 · arXiv (Cornell University)

    Differential Privacy (DP) is a widely adopted technique, valued for its effectiveness in protecting the privacy of task-specific datasets, making it a critical tool for large language models. However, its effectiveness in Multimodal Large Language …

  63. Hidden in the Noise: Unveiling Backdoors in Audio LLMs Alignment Through Latent Acoustic Pattern Triggers

    2026 · Proceedings of the AAAI Conference on Artificial Intelligence

    As Audio Large Language Models (ALLMs) emerge as powerful tools for speech processing, their safety implications demand urgent attention. While considerable research has explored textual and vision safety, audio’s distinct characteristics present significant challenges. This …

  64. Topical Word Embeddings

    2015 · Proceedings of the AAAI Conference on Artificial Intelligence

    Most word embedding models typically represent each word using a single vector, which makes these models indiscriminative for ubiquitous homonymy and polysemy. In order to enhance discriminativeness, we employ latent topic models to assign topics …

  65. Learning Tag Embeddings and Tag-specific Composition Functions in Recursive Neural Network

    2015

    Qiao Qian, Bo Tian, Minlie Huang, Yang Liu, Xuan Zhu, Xiaoyan Zhu. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume …

  66. Modeling Coverage for Neural Machine Translation

    2016 · arXiv (Cornell University)

    Attention mechanism has enhanced state-of-the-art Neural Machine Translation (NMT) by jointly learning to align and translate. It tends to ignore past alignment information, however, which often leads to over-translation and under-translation. To address this problem, …

  67. Learning Natural Language Inference using Bidirectional LSTM model and Inner-Attention

    2016 · arXiv (Cornell University)

    In this paper, we proposed a sentence encoding-based model for recognizing text entailment. In our approach, the encoding of sentence is a two-stage process. Firstly, average pooling was used over word-level bidirectional LSTM (biLSTM) to …

  68. THUMT: An Open Source Toolkit for Neural Machine Translation

    2017 · arXiv (Cornell University)

    This paper introduces THUMT, an open-source toolkit for neural machine translation (NMT) developed by the Natural Language Processing Group at Tsinghua University. THUMT implements the standard attention-based encoder-decoder framework on top of Theano and supports …

  69. Adversarial Training for Unsupervised Bilingual Lexicon Induction

    2017

    Word embeddings are well known to capture linguistic regularities of the language on which they are trained. Researchers also observe that these regularities can transfer across languages. However, previous endeavors to connect separate monolingual word …

  70. Visualizing and Understanding Neural Machine Translation

    2017

    While neural machine translation (NMT) has made remarkable progress in recent years, it is hard to interpret its internal workings due to the continuous representations and non-linearity of neural networks. In this work, we propose …

  71. Using Context Information for Dialog Act Classification in DNN Framework

    2017

    Previous work on dialog act (DA) classification has investigated different methods, such as hidden Markov models, maximum entropy, conditional random fields, graphical models, and support vector machines. A few recent studies explored using deep learning …

  72. Towards Conversational Search and Recommendation

    2018

    Conversational search and recommendation based on user-system dialogs exhibit major differences from conventional search and recommendation tasks in that 1) the user and system can interact for multiple semantically coherent rounds on a task through …

  73. Hierarchical Transformers for Multi-Document Summarization

    2019

    In this paper, we develop a neural summarization model which can effectively process multiple input documents and distill Transformer architecture with the ability to encode documents in a hierarchical manner. We represent cross-document relationships via …

  74. Reducing Word Omission Errors in Neural Machine Translation: A Contrastive Learning Approach

    2019

    While neural machine translation (NMT) has achieved remarkable success, NMT systems are prone to make word omission errors. In this work, we propose a contrastive learning approach to reducing word omission errors in NMT. The …

  75. Improving the Transformer Translation Model with Document-Level Context

    2018

    Although the Transformer translation model In this work, we extend the Transformer model with a new context encoder to represent document-level context, which is then incorporated into the original encoder and decoder. As large-scale document-level …

  76. A Dependency-Based Neural Network for Relation Classification

    2015

    Yang Liu, Furu Wei, Sujian Li, Heng Ji, Ming Zhou, Houfeng Wang. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume …

  77. Minimum Risk Training for Neural Machine Translation

    2016

    We propose minimum risk training for end-to-end neural machine translation. Unlike conventional maximum likelihood estimation, minimum risk training is capable of optimizing model parameters directly with respect to arbitrary evaluation metrics, which are not necessarily …

  78. Learning to Remember Translation History with a Continuous Cache

    2018 · Transactions of the Association for Computational Linguistics

    Existing neural machine translation (NMT) models generally translate sentences in isolation, missing the opportunity to take advantage of document-level information. In this work, we propose to augment NMT models with a very light-weight cache-like memory …

  79. Text Summarization with Pretrained Encoders

    2019 · arXiv (Cornell University)

    Bidirectional Encoder Representations from Transformers (BERT) represents the latest incarnation of pretrained language models which have recently advanced a wide range of natural language processing tasks. In this paper, we showcase how BERT can be …

  80. Iterative Dual Domain Adaptation for Neural Machine Translation

    2019

    Jiali Zeng, Yang Liu, Jinsong Su, Yubing Ge, Yaojie Lu, Yongjing Yin, Jiebo Luo. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language …

  81. MediaSum: A Large-scale Media Interview Dataset for Dialogue Summarization

    2021

    This paper introduces MEDIASUM 1 , a largescale media interview dataset consisting of 463.6K transcripts with abstractive summaries. To create this dataset, we collect interview transcripts from NPR and CNN and employ the overview and …

  82. QMSum: A New Benchmark for Query-based Multi-domain Meeting Summarization

    2021

    Ming Zhong, Da Yin, Tao Yu, Ahmad Zaidi, Mutethia Mutuma, Rahul Jha, Ahmed Hassan Awadallah, Asli Celikyilmaz, Yang Liu, Xipeng Qiu, Dragomir Radev. Proceedings of the 2021 Conference of the North American Chapter of the …

  83. DialogSum: A Real-Life Scenario Dialogue Summarization Dataset

    2021

    Proposal of large-scale datasets has facilitated research on deep neural models for news summarization. Deep learning can also be potentially useful for spoken dialogue summarization, which can benefit a range of reallife scenarios including customer …

  84. Parameter-efficient fine-tuning of large-scale pre-trained language models

    2023 · Nature Machine Intelligence

    Abstract With the prevalence of pre-trained language models (PLMs) and the pre-training–fine-tuning paradigm, it has been continuously shown that larger models tend to yield better performance. However, as PLMs scale up, fine-tuning and storing all …