ملف الباحث

Odej Kao

5 أوراق في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Implicit Parallelism through Deep Language Embedding

    2015

    The appeal of MapReduce has spawned a family of systems that implement or extend it. In order to enable parallel collection processing with User-Defined Functions (UDFs), these systems expose extensions of the MapReduce programming model …

  2. SMiPE: Estimating the Progress of Recurring Iterative Distributed Dataflows

    2017

    Distributed dataflow systems such as Apache Spark allow the execution of iterative programs at large scale on clusters. In production use, programs are often recurring and have strict latency requirements. Yet, choosing appropriate resource allocations …

  3. Chiron: Optimizing Fault Tolerance in QoS-aware Distributed Stream Processing Jobs

    2021 · arXiv (Cornell University)

    Fault tolerance is a property which needs deeper consideration when dealing with streaming jobs requiring high levels of availability and low-latency processing even in case of failures where Quality-of-Service constraints must be adhered to. Typically, …

  4. Efficient Runtime Profiling for Black-box Machine Learning Services on Sensor Streams

    2022 · arXiv (Cornell University)

    In highly distributed environments such as cloud, edge and fog computing, the application of machine learning for automating and optimizing processes is on the rise. Machine learning jobs are frequently applied in streaming conditions, where …

  5. Selecting Efficient Cluster Resources for Data Analytics: When and How to Allocate for In-Memory Processing?

    2023

    Distributed dataflow systems such as Apache Spark or Apache Flink enable parallel, in-memory data processing on large clusters of commodity hardware. Consequently, the appropriate amount of memory to allocate to the cluster is a crucial …