Researcher profile

Richong Zhang

7 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. DPS: A DSM-based Parameter Server for Machine Learning

    2017

    To solve the problem of efficient storing and updating of model parameters in the learning process, the parameter server is concerned as a high-throughput distributed machine learning (ML) architecture with the emergence of big models …

  2. MixUp as Locally Linear Out-of-Manifold Regularization

    2019 · Proceedings of the AAAI Conference on Artificial Intelligence

    MixUp (Zhang et al. 2017) is a recently proposed dataaugmentation scheme, which linearly interpolates a random pair of training examples and correspondingly the one-hot representations of their labels. Training deep neural networks with such additional …

  3. Parallel Interactive Networks for Multi-Domain Dialogue State Generation

    2020 · arXiv (Cornell University)

    The dependencies between system and user utterances in the same turn and across different turns are not fully considered in existing multidomain dialogue state tracking (MDST) models. In this study, we argue that the incorporation …

  4. On Scalar Embedding of Relative Positions in Attention Models

    2021 · Proceedings of the AAAI Conference on Artificial Intelligence

    Attention with positional encoding has been demonstrated as a powerful component in modern neural network models, such as transformers. However, why positional encoding works well in attention models remains largely unanswered. In this paper, we …

  5. Hierarchical Modeling of Label Dependency and Label Noise in Fine-grained Entity Typing

    2021

    Fine-grained entity typing (FET) aims to annotate the entity mentions in a sentence with fine-grained type labels. It brings plentiful semantic information for many natural language processing tasks. Existing FET approaches apply hard attention to …

  6. LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

    2024 · arXiv (Cornell University)

    Efficient fine-tuning is vital for adapting large language models (LLMs) to downstream tasks. However, it requires non-trivial efforts to implement these methods on different models. We present LlamaFactory, a unified framework that integrates a suite …

  7. LH-Mix: Local Hierarchy Correlation Guided Mixup over Hierarchical Prompt Tuning

    2024 · arXiv (Cornell University)

    Hierarchical text classification (HTC) aims to assign one or more labels in the hierarchy for each text. Many methods represent this structure as a global hierarchy, leading to redundant graph structures. To address this, incorporating …