Lauritz Thamsen
6 papers in the PaperMetrix corpus
Papers by this author
-
Implicit Parallelism through Deep Language Embedding
2015
The appeal of MapReduce has spawned a family of systems that implement or extend it. In order to enable parallel collection processing with User-Defined Functions (UDFs), these systems expose extensions of the MapReduce programming model …
-
SMiPE: Estimating the Progress of Recurring Iterative Distributed Dataflows
2017
Distributed dataflow systems such as Apache Spark allow the execution of iterative programs at large scale on clusters. In production use, programs are often recurring and have strict latency requirements. Yet, choosing appropriate resource allocations …
-
Towards Collaborative Optimization of Cluster Configurations for Distributed Dataflow Jobs
2020
Analyzing large datasets with distributed dataflow systems requires the use of clusters. Public cloud providers offer a large variety and quantity of resources that can be used for such clusters. However, picking the appropriate resources …
-
Chiron: Optimizing Fault Tolerance in QoS-aware Distributed Stream Processing Jobs
2021 · arXiv (Cornell University)
Fault tolerance is a property which needs deeper consideration when dealing with streaming jobs requiring high levels of availability and low-latency processing even in case of failures where Quality-of-Service constraints must be adhered to. Typically, …
-
Training Data Reduction for Performance Models of Data Analytics Jobs in the Cloud
2021 · 2021 IEEE International Conference on Big Data (Big Data)
Distributed dataflow systems like Apache Flink and Apache Spark simplify processing large amounts of data on clusters in a data-parallel manner. However, choosing suitable cluster resources for distributed dataflow jobs in both type and number …
-
Selecting Efficient Cluster Resources for Data Analytics: When and How to Allocate for In-Memory Processing?
2023
Distributed dataflow systems such as Apache Spark or Apache Flink enable parallel, in-memory data processing on large clusters of commodity hardware. Consequently, the appropriate amount of memory to allocate to the cluster is a crucial …