Researcher profile

Andy T. Liu

5 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Representation Learning of Structured Data for Medical Foundation Models

    2024 · arXiv (Cornell University)

    Large Language Models (LLMs) have demonstrated remarkable performance across various domains, including healthcare. However, their ability to effectively represent structured non-textual data, such as the alphanumeric medical codes used in records like ICD-10 or SNOMED-CT, …

  2. Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks

    2024 · arXiv (Cornell University)

    Multimodal foundation models, such as Gemini and ChatGPT, have revolutionized human-machine interactions by seamlessly integrating various forms of data. Developing a universal spoken language model that comprehends a wide range of natural language instructions is …

  3. Process-Supervised Reinforcement Learning for Interactive Multimodal Tool-Use Agents

    2025 · arXiv (Cornell University)

    Effective interactive tool use requires agents to master Tool Integrated Reasoning (TIR): a complex process involving multi-turn planning and long-context dialogue management. To train agents for this dynamic process, particularly in multi-modal contexts, we introduce …

  4. TERA: Self-Supervised Learning of Transformer Encoder Representation for Speech

    2021 · IEEE/ACM Transactions on Audio Speech and Language Processing

    We introduce a self-supervised speech pre-training method called TERA, which stands for Transformer Encoder Representations from Alteration. Recent approaches often learn by using a single auxiliary task like contrastive prediction, autoregressive prediction, or masked reconstruction. …

  5. SUPERB: Speech Processing Universal PERformance Benchmark

    2021

    Self-supervised learning (SSL) has proven vital for advancing research in natural language processing (NLP) and computer vision (CV).The paradigm pretrains a shared model on large volumes of unlabeled data and achieves state-of-the-art (SOTA) for various …