Odej Kao
5 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
Implicit Parallelism through Deep Language Embedding
2015
The appeal of MapReduce has spawned a family of systems that implement or extend it. In order to enable parallel collection processing with User-Defined Functions (UDFs), these systems expose extensions of the MapReduce programming model …
-
SMiPE: Estimating the Progress of Recurring Iterative Distributed Dataflows
2017
Distributed dataflow systems such as Apache Spark allow the execution of iterative programs at large scale on clusters. In production use, programs are often recurring and have strict latency requirements. Yet, choosing appropriate resource allocations …
-
Chiron: Optimizing Fault Tolerance in QoS-aware Distributed Stream Processing Jobs
2021 · arXiv (Cornell University)
Fault tolerance is a property which needs deeper consideration when dealing with streaming jobs requiring high levels of availability and low-latency processing even in case of failures where Quality-of-Service constraints must be adhered to. Typically, …
-
Efficient Runtime Profiling for Black-box Machine Learning Services on Sensor Streams
2022 · arXiv (Cornell University)
In highly distributed environments such as cloud, edge and fog computing, the application of machine learning for automating and optimizing processes is on the rise. Machine learning jobs are frequently applied in streaming conditions, where …
-
Selecting Efficient Cluster Resources for Data Analytics: When and How to Allocate for In-Memory Processing?
2023
Distributed dataflow systems such as Apache Spark or Apache Flink enable parallel, in-memory data processing on large clusters of commodity hardware. Consequently, the appropriate amount of memory to allocate to the cluster is a crucial …