ملف الباحث

Jonathan Will

3 أوراق في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Towards Collaborative Optimization of Cluster Configurations for Distributed Dataflow Jobs

    2020

    Analyzing large datasets with distributed dataflow systems requires the use of clusters. Public cloud providers offer a large variety and quantity of resources that can be used for such clusters. However, picking the appropriate resources …

  2. Training Data Reduction for Performance Models of Data Analytics Jobs in the Cloud

    2021 · 2021 IEEE International Conference on Big Data (Big Data)

    Distributed dataflow systems like Apache Flink and Apache Spark simplify processing large amounts of data on clusters in a data-parallel manner. However, choosing suitable cluster resources for distributed dataflow jobs in both type and number …

  3. Selecting Efficient Cluster Resources for Data Analytics: When and How to Allocate for In-Memory Processing?

    2023

    Distributed dataflow systems such as Apache Spark or Apache Flink enable parallel, in-memory data processing on large clusters of commodity hardware. Consequently, the appropriate amount of memory to allocate to the cluster is a crucial …