Researcher profile

Jia Li

11 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Determining Gains Acquired from Word Embedding Quantitatively Using Discrete Distribution Clustering

    2017

    Word embeddings have become widelyused in document analysis. While a large number of models for mapping words to vector spaces have been developed, it remains undetermined how much net gain can be achieved over traditional …

  2. Cooperative Bi-path Metric for Few-shot Learning

    2020 · arXiv (Cornell University)

    Given base classes with sufficient labeled samples, the target of few-shot classification is to recognize unlabeled samples of novel classes with only a few labeled samples. Most existing methods only pay attention to the relationship …

  3. Multimodal Data Fusion Using Canonical Variates Analysis Confusion Matrix Fusion

    2021

    Data fusion from a variety of sources requires alignment, association, and analysis. One method to determine the relationship between two variables measuring the same information is a correlation analysis. The canonical variates analysis (CVA) supports …

  4. Max-Min Diversification with Fairness Constraints: Exact and Approximation Algorithms

    2023 · arXiv (Cornell University)

    Diversity maximization aims to select a diverse and representative subset of items from a large dataset. It is a fundamental optimization task that finds applications in data summarization, feature selection, web search, recommender systems, and …

  5. Deconfounded Causal Collaborative Filtering

    2021 · arXiv (Cornell University)

    Recommender systems may be confounded by various types of confounding factors (also called confounders) that may lead to inaccurate recommendations and sacrificed recommendation performance. Current approaches to solving the problem usually design each specific model …

  6. Theoretical Proof that Auto-regressive Language Models Collapse when Real-world Data is a Finite Set

    2024 · arXiv (Cornell University)

    Auto-regressive language models (LMs) have been widely used to generate data in data-scarce domains to train new LMs, compensating for the scarcity of real-world data. Previous work experimentally found that LMs collapse when trained on …

  7. Large language models show fragile cognitive reasoning about human emotions

    2025 · ArXiv.org

    Affective computing seeks to support the holistic development of artificial intelligence by enabling machines to engage with human emotion. Recent foundation models, particularly large language models (LLMs), have been trained and evaluated on emotion-related tasks, …

  8. LONGCODEU: Benchmarking Long-Context Language Models on Long Code Understanding

    2025 · arXiv (Cornell University)

    Current advanced long-context language models offer great potential for real-world software engineering applications. However, progress in this critical domain remains hampered by a fundamental limitation: the absence of a rigorous evaluation framework for long code …

  9. Agent4S: The Transformation of Research Paradigms from the Perspective of Large Language Models

    2025 · arXiv (Cornell University)

    While AI for Science (AI4S) serves as an analytical tool in the current research paradigm, it doesn't solve its core inefficiency. We propose "Agent for Science" (Agent4S)-the use of LLM-driven agents to automate the entire …

  10. Latent Cross

    2018

    The success of recommender systems often depends on their ability to understand and make use of the context of the recommendation request. Significant research has focused on how time, location, interfaces, and a plethora of …

  11. Intent Contrastive Learning for Sequential Recommendation

    2022 · Proceedings of the ACM Web Conference 2022

    Users’ interactions with items are driven by various intents (e.g., preparing for holiday gifts, shopping for fishing equipment, etc.). However, users’ underlying intents are often unobserved/latent, making it challenging to leverage such latent intents for …