Bowen Zhou
22 papers in the PaperMetrix corpus
Papers by this author
-
Improved Neural Relation Detection for Knowledge Base Question Answering
2017 · arXiv (Cornell University)
Relation detection is a core component for many NLP applications including Knowledge Base Question Answering (KBQA). In this paper, we propose a hierarchical recurrent neural network enhanced by residual learning that detects KB relations given …
-
On the Convergence and Robustness of Adversarial Training
2021 · arXiv (Cornell University)
Improving the robustness of deep neural networks (DNNs) to adversarial examples is an important yet challenging problem for secure deep learning. Across existing defense techniques, adversarial training with Projected Gradient Decent (PGD) is amongst the …
-
A Structured Self-Attentive Sentence Embedding.
2017 · International Conference on Learning Representations
This paper proposes a new model for extracting an interpretable sentence embedding by introducing self-attention. Instead of using a vector, we use a 2-D matrix to represent the embedding, with each row of the matrix …
-
Diverse Few-Shot Text Classification with Multiple Metrics
2018
Mo Yu, Xiaoxiao Guo, Jinfeng Yi, Shiyu Chang, Saloni Potdar, Yu Cheng, Gerald Tesauro, Haoyu Wang, Bowen Zhou. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human …
-
Improving Prosody Modelling with Cross-Utterance BERT Embeddings for End-to-end Speech Synthesis
2020 · arXiv (Cornell University)
Despite prosody is related to the linguistic information up to the discourse structure, most text-to-speech (TTS) systems only take into account that within each sentence, which makes it challenging when converting a paragraph of texts …
-
Selective Attention Based Graph Convolutional Networks for Aspect-Level Sentiment Classification
2021
Recent work on aspect-level sentiment classification has employed Graph Convolutional Networks (GCN) over dependency trees to learn interactions between aspect terms and opinion words. In some cases, the corresponding opinion words for an aspect term …
-
Enhancing Chat Language Models by Scaling High-quality Instructional Conversations
2023 · arXiv (Cornell University)
Fine-tuning on instruction data has been widely validated as an effective practice for implementing chat language models like ChatGPT. Scaling the diversity and quality of such data, although straightforward, stands a great chance of leading …
-
CRaSh: Clustering, Removing, and Sharing Enhance Fine-tuning without Full Large Language Model
2023
Instruction tuning has recently been recognized as an effective way of aligning Large Language Models (LLMs) to enhance their generalization ability across various tasks. However, when tuning publicly accessible, centralized LLMs with private instruction data, …
-
Interactive Continual Learning: Fast and Slow Thinking
2024
Advanced life forms, sustained by the synergistic interaction of neural cognitive mechanisms, continually acquire and transfer knowledge throughout their lifespan. In contrast, contemporary machine learning paradigms exhibit limitations in emulating the facets of continual learning …
-
Less is More: Efficient Model Merging with Binary Task Switch
2024 · arXiv (Cornell University)
As an effective approach to equip models with multi-task capabilities without additional training, model merging has garnered significant attention. However, existing methods face challenges of redundant parameter conflicts and the excessive storage burden of parameters. …
-
Classifying Relations by Ranking with Convolutional Neural Networks
2015 · arXiv (Cornell University)
Cícero dos Santos, Bing Xiang, Bowen Zhou. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
-
LSTM-based Deep Learning Models for Non-factoid Answer Selection
2015 · arXiv (Cornell University)
In this paper, we apply a general deep learning (DL) framework for the answer selection task, which does not depend on manually defined features or linguistic tools. The basic framework is to build the embeddings …
-
ABCNN: Attention-Based Convolutional Neural Network for Modeling Sentence Pairs
2016 · Transactions of the Association for Computational Linguistics
How to model a pair of sentences is a critical issue in many NLP tasks such as answer selection (AS), paraphrase identification (PI) and textual entailment (TE). Most prior work (i) deals with one individual …
-
Attentive Pooling Networks
2016 · arXiv (Cornell University)
In this work, we propose Attentive Pooling (AP), a two-way attention mechanism for discriminative model training. In the context of pair-wise ranking or classification with neural networks, AP enables the pooling layer to be aware …
-
Pointing the Unknown Words
2016 · arXiv (Cornell University)
The problem of rare and unknown words is an important issue that can potentially influence the performance of many NLP systems, including both the traditional count-based and the deep learning models. We propose a novel …
-
Multiresolution Recurrent Neural Networks: An Application to Dialogue Response Generation
2017 · Proceedings of the AAAI Conference on Artificial Intelligence
We introduce a new class of models called multiresolution recurrent neural networks, which explicitly model natural language generation at multiple levels of abstraction. The models extend the sequence-to-sequence framework to generate two parallel stochastic processes: …
-
Improved Representation Learning for Question Answer Matching
2016
Passage-level question answer matching is a challenging task since it requires effective representations that capture the complex semantic relations between questions and answers. In this work, we propose a series of deep learning models to …
-
SummaRuNNer: A Recurrent Neural Network based Sequence Model for Extractive Summarization of Documents
2016 · arXiv (Cornell University)
We present SummaRuNNer, a Recurrent Neural Network (RNN) based sequence model for extractive summarization of documents and show that it achieves performance better than or comparable to state-of-the-art. Our model has the additional advantage of …
-
R$^3$: Reinforced Reader-Ranker for Open-Domain Question Answering
2017 · arXiv (Cornell University)
In recent years researchers have achieved considerable success applying neural network methods to question answering (QA). These approaches have achieved state of the art results in simplified closed-domain settings such as the SQuAD (Rajpurkar et …
-
Multi-hop Reading Comprehension across Multiple Documents by Reasoning over Heterogeneous Graphs
2019
Multi-hop reading comprehension (RC) across documents poses new challenge over single-document RC because it requires reasoning over multiple documents to reach the final answer. In this paper, we propose a new model to tackle the …
-
SummaRuNNer: A Recurrent Neural Network Based Sequence Model for Extractive Summarization of Documents
2017 · Proceedings of the AAAI Conference on Artificial Intelligence
We present SummaRuNNer, a Recurrent Neural Network (RNN) based sequence model for extractive summarization of documents and show that it achieves performance better than or comparable to state-of-the-art. Our model has the additional advantage of …
-
Applying deep learning to answer selection: A study and an open task
2015
We apply a general deep learning framework to address the non-factoid question answering task. Our approach does not rely on any linguistic tools and can be applied to different languages or domains. Various architectures are …