Yang Liu
84 papers in the PaperMetrix corpus
Papers by this author
-
Learning Cross-lingual Word Embeddings via Matrix Co-factorization
2015
Tianze Shi, Zhiyuan Liu, Yang Liu, Maosong Sun. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 2: Short Papers). 2015.
-
Generalized Agreement for Bidirectional Word Alignment
2015
While agreement-based joint training has proven to deliver state-of-the-art alignment accuracy, the produced word alignments are usually restricted to one-toone mappings because of the hard constraint on agreement. We propose a general framework to allow …
-
Construction of Smart Campus Based on Situational Awareness in the Era Of Big Data
2016
Big data as a new data management technology, based on the current Internet of things and cloud computing "smart campus" has an important role in the construction. "Smart campus" as an information construction of educational …
-
Particle flow for sequential Monte Carlo implementation of probability hypothesis density
2017
Target tracking is a challenging task and generally no analytical solution is available, especially for the multi-target tracking systems. To address this problem, probability hypothesis density (PHD) filter is used by propagating the PHD instead …
-
Brand key asset discovery via cluster-wise biased discriminant projection
2017 · Proceedings of the International Conference on Web Intelligence
Accurate and effective discovery of a brand's key assets, namely, Key Opinion Leaders (KOLs) and potential customers, plays an essential role in marketing campaigns. In a massive online social network, brands are challenged with identifying …
-
A hierarchical classification approach for tor anonymous traffic
2017
Tor is an anonymous communication system that can protect our privacy, but it also provides a haven for criminals to avoid network tracing. Therefore, anonymous traffic analysis and classification is an important part of maintaining …
-
Radar signal sorting algorithm of k-means clustering based on data field
2017
Radar signal sorting is one of the essential technologies in radar countermeasures reconnaissance system. Non-cooperative radar signal sorting without prior information has been a great challenge for radar countermeasures. This paper presents a k-means clustering …
-
A prototype simulator for the simulation of complicated hadoop framework behaviors
2017
Distributed computing and parallel computing have become the most effective tools for solving complex problems in amount of academia and industrial fields. Among a number of distributed and parallel computing technologies, MapReduce has been proved …
-
ReCDroid: Automatically Reproducing Android Application Crashes from Bug Reports
2019
The large demand of mobile devices creates significant concerns about the quality of mobile applications (apps). Developers heavily rely on bug reports in issue tracking systems to reproduce failures (e.g., crashes). However, the process of …
-
A Multi-Goal Oriented Approach for Adaptation Rules Generation
2018
Modern software runs in a dynamic, uncertain environment, and should satisfy multiple goals simultaneously. In order to allow software to respond to changes in the environment or user requirements and meet user goals continuously, an …
-
Experimental test of error-disturbance uncertainty relation with continuous variables
2019 · arXiv (Cornell University)
Uncertainty relation is one of the fundamental principle in quantum mechanics and plays an important role in quantum information science. We experimentally test the error-disturbance uncertainty relation (EDR) with continuous variables for Gaussian states. Two …
-
aLeak: Context-Free Side-Channel from Your Smart Watch Leaks Your Typing Privacy
2019 · IEEE Transactions on Mobile Computing
We revisit a crucial privacy problem in this paper - can the sensitive information, like the numeric passwords and personal data, frequently typed by user on mobile devices be inferred through the motion sensors of …
-
Multi-info Fusion Based Video Recommendation System
2019 · Journal of Physics Conference Series
Abstract The great progress in recommendation system help users discover more interesting items that satisfy their appetites. Considering the video recommendation is an increasing popular sub-field of recommendation, but the traditional recommendation techniques such as …
-
A Countermeasure Against Statistical Ineffective Fault Analysis
2020 · IEEE Transactions on Circuits & Systems II Express Briefs
Current state-of-the-art countermeasures against Fault Injection Attacks (FIA) provide good protection against analysis methods that require the differences in the correct and faulty ciphertext to derive the secret information, such as Differential Fault Analysis (DFA) …
-
Abnormal Client Behavior Detection in Federated Learning
2019 · arXiv (Cornell University)
In federated learning systems, clients are autonomous in that their behaviors are not fully governed by the server. Consequently, a client may intentionally or unintentionally deviate from the prescribed course of federated model training, resulting …
-
Stealing Deep Reinforcement Learning Models for Fun and Profit
2020 · arXiv (Cornell University)
This paper presents the first model extraction attack against Deep Reinforcement Learning (DRL), which enables an external adversary to precisely recover a black-box DRL model only from its interaction with the environment. Model extraction attacks …
-
Generating Behavior-Diverse Game AIs with Evolutionary Multi-Objective Deep Reinforcement Learning
2020
Generating diverse behaviors for game artificial intelligence (Game AI) has been long recognized as a challenging task in the game industry. Designing a Game AI with a satisfying behavioral characteristic (style) heavily depends on the …
-
Shortened Linear Codes over Finite Fields
2020 · arXiv (Cornell University)
The puncturing and shortening technique are two important approaches to constructing new linear codes from old ones. In the past 70 years, a lot of progress on the puncturing technique has been made, and many …
-
BatchCrypt: Efficient homomorphic encryption for cross-silo federated learning
2020 · Rare & Special e-Zone (The Hong Kong University of Science and Technology)
Cross-silo federated learning (FL) enables organizations (e.g., financial or medical) to collaboratively train a machine learning model by aggregating local gradient updates from each client without sharing privacy-sensitive data. To ensure no update is revealed …
-
Learning to Expand: Reinforced Pseudo-relevance Feedback Selection for Information-seeking Conversations
2020 · arXiv (Cornell University)
Information-seeking conversation systems are increasingly popular in real-world applications, especially for e-commerce companies. To retrieve appropriate responses for users, it is necessary to compute the matching degrees between candidate responses and users' queries with historical …
-
Neural Machine Translation With Explicit Phrase Alignment
2021 · IEEE/ACM Transactions on Audio Speech and Language Processing
While neural machine translation has achieved state-of-the-art translation performance, it is unable to capture the alignment between the input and output during the translation process. The lack of alignment in neural machine translation models leads …
-
Linear Classifiers that Encourage Constructive Adaptation
2020 · arXiv (Cornell University)
Machine learning systems are often used in settings where individuals adapt their features to obtain a desired outcome. In such settings, strategic behavior leads to a sharp loss in model performance in deployment. In this …
-
Synthetic Benchmarks for Scientific Research in Explainable Machine Learning
2021 · arXiv (Cornell University)
As machine learning models grow more complex and their applications become more high-stakes, tools for explaining model predictions have become increasingly important. This has spurred a flurry of research in model explainability and has given …
-
Optimization and Simulation of Labor Resource Management Information Platform Based on Internet of Things
2021 · Wireless Communications and Mobile Computing
This paper conducts an in‐depth analysis and research on the optimization of the labor resource management information platform through the Internet of Things (IoT) technology; through the collection, classification, and data search functions of this …
-
Fine-tuning Is Not Enough: A Simple yet Effective Watermark Removal Attack for DNN Models
2021
Watermarking has become the tendency in protecting the intellectual property of DNN models. Recent works, from the adversary's perspective, attempted to subvert watermarking mechanisms by designing watermark removal attacks. However, these attacks mainly adopted sophisticated …
-
A Large-Scale Empirical Study of Real-Life Performance Issues in Open Source Projects
2022 · Zenodo (CERN European Organization for Nuclear Research)
1. The spreadsheet "Perf Issue Empirical Data Package.xlsx" contains the details of data extraction and annotation of the performance issues. The three tabs in the above spreadsheet, i.e., “Java Projects Issues”, “Python Projects Issues”, and …
-
Neighboring Backdoor Attacks on Graph Convolutional Network
2022 · arXiv (Cornell University)
Backdoor attacks have been widely studied to hide the misclassification rules in the normal models, which are only activated when the model is aware of the specific inputs (i.e., the trigger). However, despite their success …
-
MAGAN: Mask Attention Generative Adversarial Network for Liver Tumor CT Image Synthesis
2020 · Research Square (Research Square)
Abstract Background : For deep learning, the size of the dataset greatly affects the final training effect. However, in the field of computer-aided diagnosis, medical image datasets are often limited and even scarce. Methods : …
-
A Template-based Method for Constrained Neural Machine Translation
2022 · arXiv (Cornell University)
Machine translation systems are expected to cope with various types of constraints in many practical scenarios. While neural machine translation (NMT) has achieved strong performance in unconstrained cases, it is non-trivial to impose pre-specified constraints …
-
Defense against Backdoor Attacks via Identifying and Purifying Bad Neurons
2022 · arXiv (Cornell University)
The opacity of neural networks leads their vulnerability to backdoor attacks, where hidden attention of infected neurons is triggered to override normal predictions to the attacker-chosen ones. In this paper, we propose a novel backdoor …
-
Improving Radiology Summarization with Radiograph and Anatomy Prompts
2022 · arXiv (Cornell University)
The impression is crucial for the referring physicians to grasp key information since it is concluded from the findings and reasoning of radiologists. To alleviate the workload of radiologists and reduce repetitive human labor in …
-
Semi-Supervised Learning for Neural Machine Translation
2016 · arXiv (Cornell University)
While end-to-end neural machine translation (NMT) has made remarkable progress recently, NMT systems only rely on parallel corpora for parameter estimation. Since parallel corpora are usually limited in quantity, quality, and coverage, especially for low-resource …
-
Listening to Users' Voice: Automatic Summarization of Helpful App Reviews
2022 · IEEE Transactions on Reliability
App reviews are crowdsourcing knowledge of user experience with the apps, providing valuable information for app release planning, such as major bugs to fix and important features to add. There exist prior explorations on app …
-
From "Law + Engineering" to Engineering Law: Current Situation, Goals and Value Implications
2020 · JOURNAL OF ENGINEERING STUDIES
The study of legal issues in the productivity development of engineering activities at its core is a response to the orderly needs and risk regulations of complex engineering fields. It is a realistic need for …
-
Using In-Context Learning to Improve Dialogue Safety
2023 · arXiv (Cornell University)
While large neural-based conversational models have become increasingly proficient dialogue agents, recent work has highlighted safety issues with these systems. For example, these systems can be goaded into generating toxic content, which often perpetuates social …
-
G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment
2023 · arXiv (Cornell University)
The quality of texts generated by natural language generation (NLG) systems is hard to measure automatically. Conventional reference-based metrics, such as BLEU and ROUGE, have been shown to have relatively low correlation with human judgments, …
-
Casualty on the Titanic based on Machine Learning Methods
2023 · Highlights in Science Engineering and Technology
The Titanic sank on April 15, 1914, with 2224 people on board, and only 32% survived. The survivors are somewhat random, but they are somewhat the same. Studying the types of people who are more …
-
MSN-net: Multi-Scale Normality Network for Video Anomaly Detection
2023
Existing unsupervised video anomaly detection methods often suffer from performance degradation due to the overgeneralization of deep models. In this paper, we propose a simple yet effective Multi-Scale Normality network (MSN-net) that uses hierarchical memories …
-
Who is the Real Hero? Measuring Developer Contribution via Multi-dimensional Data Integration
2023 · arXiv (Cornell University)
Proper incentives are important for motivating developers in open-source communities, which is crucial for maintaining the development of open-source software healthy. To provide such incentives, an accurate and objective developer contribution measurement method is needed. …
-
Privacy-Preserving Techniques in Cloud/Fog and Internet of Things
2023 · Cryptography
Recently, wireless networks have been developed using cloud infrastructure and software-based networks [...]
-
ASTER: Automatic Speech Recognition System Accessibility Testing for Stutterers
2023
The popularity of automatic speech recognition (ASR) systems nowadays leads to an increasing need for improving their accessibility. Handling stuttering speech is an important feature for accessible ASR systems. To improve the accessibility of ASR …
-
MERCY: Multiple Response Ranking Concurrently in Realistic Open-Domain Conversational Systems
2023
Automatic Evaluation (AE) and Response Selection (RS) models assign quality scores to various candidate responses and rank them in conversational setups. Prior response ranking research compares various models’ performance on synthetically generated test sets. In …
-
Data-Efficient Alignment of Large Language Models with Human Feedback Through Natural Language
2023 · arXiv (Cornell University)
Learning from human feedback is a prominent technique to align the output of large language models (LLMs) with human expectations. Reinforcement learning from human feedback (RLHF) leverages human preference signals that are in the form …
-
A balanced allocation of network teaching resources in higher vocational colleges based on demand prediction
2023 · International Journal of Continuing Engineering Education and Life-Long Learning
Because the traditional teaching resource allocation method has the problems of low accuracy of resource demand prediction and low balance of resource allocation, this paper studies a new balanced allocation method based on demand prediction. …
-
Human-Instruction-Free LLM Self-Alignment with Limited Samples
2024 · arXiv (Cornell University)
Aligning large language models (LLMs) with human values is a vital task for LLM practitioners. Current alignment techniques have several limitations: (1) requiring a large amount of annotated data; (2) demanding heavy human involvement; (3) …
-
BMLP: Behavior-aware MLP for Heterogeneous Sequential Recommendation
2024 · arXiv (Cornell University)
In real recommendation scenarios, users often have different types of behaviors, such as clicking and buying. Existing research methods show that it is possible to capture the heterogeneous interests of users through different types of …
-
Datasets for Large Language Models: A Comprehensive Survey
2024 · arXiv (Cornell University)
This paper embarks on an exploration into the Large Language Model (LLM) datasets, which play a crucial role in the remarkable advancements of LLMs. The datasets serve as the foundational infrastructure analogous to a root …
-
ToolRerank: Adaptive and Hierarchy-Aware Reranking for Tool Retrieval
2024 · arXiv (Cornell University)
Tool learning aims to extend the capabilities of large language models (LLMs) with external tools. A major challenge in tool learning is how to support a large number of tools, including unseen tools. To address …
-
Adversarial Learning for Coordinate Regression Through k-Layer Penetrating Representation
2024 · IEEE Transactions on Dependable and Secure Computing
Adversarial attack is a crucial step when evaluating the reliability and robustness of deep neural networks (DNNs) models. Most existing attack approaches apply an end-to-end gradient update strategy to generate adversarial examples for a classification …
-
Towards Enabling DPOAE Estimation on Single-Speaker Earbuds
2024
Distortion Product OtoAcoustic Emissions (DPOAEs) represents faint cochlear responses to dual-frequency stimuli, commonly employed in hearing screening. This paper introduces an innovative approach to trigger DPOAEs using single-speaker earbuds. Due to their compact size, the …
-
Revisiting a Pain in the Neck: Semantic Phrase Processing Benchmark for Language Models
2024 · arXiv (Cornell University)
We introduce LexBench, a comprehensive evaluation suite enabled to test language models (LMs) on ten semantic phrase processing tasks. Unlike prior studies, it is the first work to propose a framework from the comparative perspective …
-
Membership Inference on Text-to-Image Diffusion Models via Conditional Likelihood Discrepancy
2024 · arXiv (Cornell University)
Text-to-image diffusion models have achieved tremendous success in the field of controllable image generation, while also coming along with issues of privacy leakage and data copyrights. Membership inference arises in these contexts as a potential …
-
Understanding the Effectiveness of Coverage Criteria for Large Language Models: A Special Angle from Jailbreak Attacks
2024 · arXiv (Cornell University)
Large language models (LLMs) have revolutionized artificial intelligence, but their increasing deployment across critical domains has raised concerns about their abnormal behaviors when faced with malicious attacks. Such vulnerability alerts the widespread inadequacy of pre-release …
-
Bridging the Gap: Unpacking the Hidden Challenges in Knowledge Distillation for Online Ranking Systems
2024
Knowledge Distillation (KD) is a powerful approach for compressing a large model into a smaller, more efficient model, particularly beneficial for latency-sensitive applications like recommender systems. However, current KD research predominantly focuses on Computer Vision …
-
SelfBC: Self Behavior Cloning for Offline Reinforcement Learning
2024 · arXiv (Cornell University)
Policy constraint methods in offline reinforcement learning employ additional regularization techniques to constrain the discrepancy between the learned policy and the offline dataset. However, these methods tend to result in overly conservative policies that resemble …
-
On the Role of Attention Heads in Large Language Model Safety
2024 · arXiv (Cornell University)
Large language models (LLMs) achieve state-of-the-art performance on multiple language tasks, yet their safety guardrails can be circumvented, leading to harmful generations. In light of this, recent research on safety mechanisms has emerged, revealing that …
-
Concept-Aware Graph Convolutional Network for Compositional Zero-Shot Learning
2025 · IEEE Transactions on Neural Networks and Learning Systems
Compositional zero-shot learning (CZSL) aims to identify unobservable compositional concepts with prior knowledge of known primitives (attributes and objects). Due to distribution differences between seen and unseen components, existing methods for CZSL often ignore intrinsic …
-
Automated Runtime Verification of Security for E-Commerce Smart Contracts
2025 · Journal of theoretical and applied electronic commerce research
As a novel decentralized computing paradigm, blockchain is expected to disrupt the existing e-commerce architecture and process. Secure smart contracts are the crucial foundation for e-commerce based on blockchain. However, vulnerabilities in smart contracts occur …
-
Incorporating Pre-Training Data Matters in Unsupervised Domain Adaptation
2025 · IEEE Transactions on Pattern Analysis and Machine Intelligence
In deep learning, initializing models with pre-trained weights has become the de facto practice for various downstream tasks. Many unsupervised domain adaptation (UDA) methods typically adopt a backbone pre-trained on ImageNet, and focus on reducing …
-
SAP: Privacy-Preserving Fine-Tuning on Language Models with Split-and-Privatize Framework
2025
Pre-trained Language Models (PLM) have enabled a cost-effective approach to handling various downstream applications via Parameter-Efficient-Fine-Tuning (PEFT) techniques. In this context, service providers have introduced a popular fine-tuning-based product service known as Model-as-a-Service (MaaS). This …
-
COSMIC: Generalized Refusal Direction Identification in LLM Activations
2025 · arXiv (Cornell University)
Large Language Models (LLMs) encode behaviors such as refusal within their activation space, yet identifying these behaviors remains a significant challenge. Existing methods often rely on predefined refusal templates detectable in output tokens or require …
-
Dual-Priv Pruning : Efficient Differential Private Fine-Tuning in Multimodal Large Language Models
2025 · arXiv (Cornell University)
Differential Privacy (DP) is a widely adopted technique, valued for its effectiveness in protecting the privacy of task-specific datasets, making it a critical tool for large language models. However, its effectiveness in Multimodal Large Language …
-
Hidden in the Noise: Unveiling Backdoors in Audio LLMs Alignment Through Latent Acoustic Pattern Triggers
2026 · Proceedings of the AAAI Conference on Artificial Intelligence
As Audio Large Language Models (ALLMs) emerge as powerful tools for speech processing, their safety implications demand urgent attention. While considerable research has explored textual and vision safety, audio’s distinct characteristics present significant challenges. This …
-
Topical Word Embeddings
2015 · Proceedings of the AAAI Conference on Artificial Intelligence
Most word embedding models typically represent each word using a single vector, which makes these models indiscriminative for ubiquitous homonymy and polysemy. In order to enhance discriminativeness, we employ latent topic models to assign topics …
-
Learning Tag Embeddings and Tag-specific Composition Functions in Recursive Neural Network
2015
Qiao Qian, Bo Tian, Minlie Huang, Yang Liu, Xuan Zhu, Xiaoyan Zhu. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume …
-
Modeling Coverage for Neural Machine Translation
2016 · arXiv (Cornell University)
Attention mechanism has enhanced state-of-the-art Neural Machine Translation (NMT) by jointly learning to align and translate. It tends to ignore past alignment information, however, which often leads to over-translation and under-translation. To address this problem, …
-
Learning Natural Language Inference using Bidirectional LSTM model and Inner-Attention
2016 · arXiv (Cornell University)
In this paper, we proposed a sentence encoding-based model for recognizing text entailment. In our approach, the encoding of sentence is a two-stage process. Firstly, average pooling was used over word-level bidirectional LSTM (biLSTM) to …
-
THUMT: An Open Source Toolkit for Neural Machine Translation
2017 · arXiv (Cornell University)
This paper introduces THUMT, an open-source toolkit for neural machine translation (NMT) developed by the Natural Language Processing Group at Tsinghua University. THUMT implements the standard attention-based encoder-decoder framework on top of Theano and supports …
-
Adversarial Training for Unsupervised Bilingual Lexicon Induction
2017
Word embeddings are well known to capture linguistic regularities of the language on which they are trained. Researchers also observe that these regularities can transfer across languages. However, previous endeavors to connect separate monolingual word …
-
Visualizing and Understanding Neural Machine Translation
2017
While neural machine translation (NMT) has made remarkable progress in recent years, it is hard to interpret its internal workings due to the continuous representations and non-linearity of neural networks. In this work, we propose …
-
Using Context Information for Dialog Act Classification in DNN Framework
2017
Previous work on dialog act (DA) classification has investigated different methods, such as hidden Markov models, maximum entropy, conditional random fields, graphical models, and support vector machines. A few recent studies explored using deep learning …
-
Towards Conversational Search and Recommendation
2018
Conversational search and recommendation based on user-system dialogs exhibit major differences from conventional search and recommendation tasks in that 1) the user and system can interact for multiple semantically coherent rounds on a task through …
-
Hierarchical Transformers for Multi-Document Summarization
2019
In this paper, we develop a neural summarization model which can effectively process multiple input documents and distill Transformer architecture with the ability to encode documents in a hierarchical manner. We represent cross-document relationships via …
-
Reducing Word Omission Errors in Neural Machine Translation: A Contrastive Learning Approach
2019
While neural machine translation (NMT) has achieved remarkable success, NMT systems are prone to make word omission errors. In this work, we propose a contrastive learning approach to reducing word omission errors in NMT. The …
-
Improving the Transformer Translation Model with Document-Level Context
2018
Although the Transformer translation model In this work, we extend the Transformer model with a new context encoder to represent document-level context, which is then incorporated into the original encoder and decoder. As large-scale document-level …
-
A Dependency-Based Neural Network for Relation Classification
2015
Yang Liu, Furu Wei, Sujian Li, Heng Ji, Ming Zhou, Houfeng Wang. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume …
-
Minimum Risk Training for Neural Machine Translation
2016
We propose minimum risk training for end-to-end neural machine translation. Unlike conventional maximum likelihood estimation, minimum risk training is capable of optimizing model parameters directly with respect to arbitrary evaluation metrics, which are not necessarily …
-
Learning to Remember Translation History with a Continuous Cache
2018 · Transactions of the Association for Computational Linguistics
Existing neural machine translation (NMT) models generally translate sentences in isolation, missing the opportunity to take advantage of document-level information. In this work, we propose to augment NMT models with a very light-weight cache-like memory …
-
Text Summarization with Pretrained Encoders
2019 · arXiv (Cornell University)
Bidirectional Encoder Representations from Transformers (BERT) represents the latest incarnation of pretrained language models which have recently advanced a wide range of natural language processing tasks. In this paper, we showcase how BERT can be …
-
Iterative Dual Domain Adaptation for Neural Machine Translation
2019
Jiali Zeng, Yang Liu, Jinsong Su, Yubing Ge, Yaojie Lu, Yongjing Yin, Jiebo Luo. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language …
-
MediaSum: A Large-scale Media Interview Dataset for Dialogue Summarization
2021
This paper introduces MEDIASUM 1 , a largescale media interview dataset consisting of 463.6K transcripts with abstractive summaries. To create this dataset, we collect interview transcripts from NPR and CNN and employ the overview and …
-
QMSum: A New Benchmark for Query-based Multi-domain Meeting Summarization
2021
Ming Zhong, Da Yin, Tao Yu, Ahmad Zaidi, Mutethia Mutuma, Rahul Jha, Ahmed Hassan Awadallah, Asli Celikyilmaz, Yang Liu, Xipeng Qiu, Dragomir Radev. Proceedings of the 2021 Conference of the North American Chapter of the …
-
DialogSum: A Real-Life Scenario Dialogue Summarization Dataset
2021
Proposal of large-scale datasets has facilitated research on deep neural models for news summarization. Deep learning can also be potentially useful for spoken dialogue summarization, which can benefit a range of reallife scenarios including customer …
-
Parameter-efficient fine-tuning of large-scale pre-trained language models
2023 · Nature Machine Intelligence
Abstract With the prevalence of pre-trained language models (PLMs) and the pre-training–fine-tuning paradigm, it has been continuously shown that larger models tend to yield better performance. However, as PLMs scale up, fine-tuning and storing all …