Jonathan Will
3 papers in the PaperMetrix corpus
Papers by this author
-
Towards Collaborative Optimization of Cluster Configurations for Distributed Dataflow Jobs
2020
Analyzing large datasets with distributed dataflow systems requires the use of clusters. Public cloud providers offer a large variety and quantity of resources that can be used for such clusters. However, picking the appropriate resources …
-
Training Data Reduction for Performance Models of Data Analytics Jobs in the Cloud
2021 · 2021 IEEE International Conference on Big Data (Big Data)
Distributed dataflow systems like Apache Flink and Apache Spark simplify processing large amounts of data on clusters in a data-parallel manner. However, choosing suitable cluster resources for distributed dataflow jobs in both type and number …
-
Selecting Efficient Cluster Resources for Data Analytics: When and How to Allocate for In-Memory Processing?
2023
Distributed dataflow systems such as Apache Spark or Apache Flink enable parallel, in-memory data processing on large clusters of commodity hardware. Consequently, the appropriate amount of memory to allocate to the cluster is a crucial …