Issei Sato
6 papers in the PaperMetrix corpus
Papers by this author
-
Infinite Plaid Models for Infinite Bi-Clustering
2016 · Proceedings of the AAAI Conference on Artificial Intelligence
We propose a probabilistic model for non-exhaustive and overlapping (NEO) bi-clustering. Our goal is to extract a few sub-matrices from the given data matrix, where entries of a sub-matrix are characterized by a specific distribution …
-
Normalized Flat Minima: Exploring Scale Invariant Definition of Flat Minima for Neural Networks using PAC-Bayesian Analysis
2019 · arXiv (Cornell University)
The notion of flat minima has played a key role in the generalization studies of deep learning models. However, existing definitions of the flatness are known to be sensitive to the rescaling of parameters. The …
-
Solving NP-Hard Problems on Graphs with Extended AlphaGo Zero
2019 · arXiv (Cornell University)
There have been increasing challenges to solve combinatorial optimization problems by machine learning. Khalil et al. proposed an end-to-end reinforcement learning framework, S2V-DQN, which automatically learns graph embeddings to construct solutions to a wide range …
-
Few-shot Domain Adaptation by Causal Mechanism Transfer
2020 · International Conference on Machine Learning
We study few-shot supervised domain adaptation (DA) for regression problems, where only a few labeled target domain data and many labeled source domain data are available. Many of the current DA methods base their transfer …
-
Are Transformers with One Layer Self-Attention Using Low-Rank Weight Matrices Universal Approximators?
2023 · arXiv (Cornell University)
Existing analyses of the expressive capacity of Transformer models have required excessively deep layers for data memorization, leading to a discrepancy with the Transformers actually used in practice. This is primarily due to the interpretation …
-
Understanding Linear Probing then Fine-tuning Language Models from NTK Perspective
2024 · arXiv (Cornell University)
The two-stage fine-tuning (FT) method, linear probing (LP) then fine-tuning (LP-FT), outperforms linear probing and FT alone. This holds true for both in-distribution (ID) and out-of-distribution (OOD) data. One key reason for its success is …