ملف الباحث

Issei Sato

6 أوراق في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Infinite Plaid Models for Infinite Bi-Clustering

    2016 · Proceedings of the AAAI Conference on Artificial Intelligence

    We propose a probabilistic model for non-exhaustive and overlapping (NEO) bi-clustering. Our goal is to extract a few sub-matrices from the given data matrix, where entries of a sub-matrix are characterized by a specific distribution …

  2. Normalized Flat Minima: Exploring Scale Invariant Definition of Flat Minima for Neural Networks using PAC-Bayesian Analysis

    2019 · arXiv (Cornell University)

    The notion of flat minima has played a key role in the generalization studies of deep learning models. However, existing definitions of the flatness are known to be sensitive to the rescaling of parameters. The …

  3. Solving NP-Hard Problems on Graphs with Extended AlphaGo Zero

    2019 · arXiv (Cornell University)

    There have been increasing challenges to solve combinatorial optimization problems by machine learning. Khalil et al. proposed an end-to-end reinforcement learning framework, S2V-DQN, which automatically learns graph embeddings to construct solutions to a wide range …

  4. Few-shot Domain Adaptation by Causal Mechanism Transfer

    2020 · International Conference on Machine Learning

    We study few-shot supervised domain adaptation (DA) for regression problems, where only a few labeled target domain data and many labeled source domain data are available. Many of the current DA methods base their transfer …

  5. Are Transformers with One Layer Self-Attention Using Low-Rank Weight Matrices Universal Approximators?

    2023 · arXiv (Cornell University)

    Existing analyses of the expressive capacity of Transformer models have required excessively deep layers for data memorization, leading to a discrepancy with the Transformers actually used in practice. This is primarily due to the interpretation …

  6. Understanding Linear Probing then Fine-tuning Language Models from NTK Perspective

    2024 · arXiv (Cornell University)

    The two-stage fine-tuning (FT) method, linear probing (LP) then fine-tuning (LP-FT), outperforms linear probing and FT alone. This holds true for both in-distribution (ID) and out-of-distribution (OOD) data. One key reason for its success is …