ملف الباحث

Yun Yang

7 أوراق في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. An Algorithm for Finding the Minimum Cost of Storing and Regenerating Datasets in Multiple Clouds

    2015 · IEEE Transactions on Cloud Computing

    The proliferation of cloud computing allows users to flexibly store, re-compute or transfer large generated datasets with multiple cloud service providers. However, due to the pay-as-you-go model, the total cost of using cloud services depends …

  2. Bayesian fractional posteriors

    2016 · arXiv (Cornell University)

    We consider the fractional posterior distribution that is obtained by updating a prior distribution via Bayes theorem with a fractional likelihood function, a usual likelihood function raised to a fractional power. First, we analyze the …

  3. A Novel Data Set Importance Based Cost-Effective and Computation-Efficient Storage Strategy in the Cloud

    2017

    The rapid development of cloud computing service allows data and computation intensive applications to be easily moved into cloud. Users pay for computing and storing resources to deal with their data in the cloud, therefore …

  4. An Agent-Based Decentralised Service Monitoring Approach in Multi-Tenant Service-Based Systems

    2017

    Service monitoring is an important research problem in service-based systems (SBSs), which aims to monitor the failure of services in a timely manner while using resources as few as possible. Most of the existing service …

  5. A Deep Context-wise Method for Coreference Detection in Natural Language Requirements

    2020

    Requirements are usually written by different stakeholders with diverse backgrounds and skills and evolve continuously. Therefore inconsistency caused by specialized jargons and different domains, is inevitable. In particular, entity coreference in Requirement Engineering (RE) is …

  6. Sketch-and-Lift: Scalable Subsampled Semidefinite Program for $K$-means Clustering

    2022 · arXiv (Cornell University)

    Semidefinite programming (SDP) is a powerful tool for tackling a wide range of computationally hard problems such as clustering. Despite the high accuracy, semidefinite programs are often too slow in practice with poor scalability on …

  7. Distributed Database Design and Performance Tuing

    2023

    it is necessary to design a distributed database for real-time data storage to meet the needs of rapid data collection, for the operation and maintenance of large and complex systems. The system adopts the distributed …