Dominik Scheinert
3 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
Training Data Reduction for Performance Models of Data Analytics Jobs in the Cloud
2021 · 2021 IEEE International Conference on Big Data (Big Data)
Distributed dataflow systems like Apache Flink and Apache Spark simplify processing large amounts of data on clusters in a data-parallel manner. However, choosing suitable cluster resources for distributed dataflow jobs in both type and number …
-
Efficient Runtime Profiling for Black-box Machine Learning Services on Sensor Streams
2022 · arXiv (Cornell University)
In highly distributed environments such as cloud, edge and fog computing, the application of machine learning for automating and optimizing processes is on the rise. Machine learning jobs are frequently applied in streaming conditions, where …
-
Selecting Efficient Cluster Resources for Data Analytics: When and How to Allocate for In-Memory Processing?
2023
Distributed dataflow systems such as Apache Spark or Apache Flink enable parallel, in-memory data processing on large clusters of commodity hardware. Consequently, the appropriate amount of memory to allocate to the cluster is a crucial …