Representations and Optimizations for Embedded Parallel Dataflow Languages-Reference-Cited by-同舟云学术

Representations and Optimizations for Embedded Parallel Dataflow Languages

Published:2019-01-29 Issue:1 Volume:44 Page:1-44
ISSN:0362-5915
Container-title:ACM Transactions on Database Systems
language:en
Short-container-title:ACM Trans. Database Syst.

Author:

Alexandrov Alexander¹,Krastev Georgi¹,Markl Volker¹

Affiliation:

1. TU Berlin, Germany

Abstract

Parallel dataflow engines such as Apache Hadoop, Apache Spark, and Apache Flink are an established alternative to relational databases for modern data analysis applications. A characteristic of these systems is a scalable programming model based on distributed collections and parallel transformations expressed by means of second-order functions such as map and reduce. Notable examples are Flink’s DataSet and Spark’s RDD programming abstractions. These programming models are realized as EDSLs—domain specific languages embedded in a general-purpose host language such as Java, Scala, or Python. This approach has several advantages over traditional external DSLs such as SQL or XQuery. First, syntactic constructs from the host language (e.g., anonymous functions syntax, value definitions, and fluent syntax via method chaining) can be reused in the EDSL. This eases the learning curve for developers already familiar with the host language. Second, it allows for seamless integration of library methods written in the host language via the function parameters passed to the parallel dataflow operators. This reduces the effort for developing analytics dataflows that go beyond pure SQL and require domain-specific logic. At the same time, however, state-of-the-art parallel dataflow EDSLs exhibit a number of shortcomings. First, one of the main advantages of an external DSL such as SQL—the high-level, declarative Select-From-Where syntax—is either lost completely or mimicked in a non-standard way. Second, execution aspects such as caching, join order, and partial aggregation have to be decided by the programmer. Optimizing them automatically is very difficult due to the limited program context available in the intermediate representation of the DSL. In this article, we argue that the limitations listed above are a side effect of the adopted type-based embedding approach. As a solution, we propose an alternative EDSL design based on quotations. We present a DSL embedded in Scala and discuss its compiler pipeline, intermediate representation, and some of the enabled optimizations. We promote the algebraic type of bags in union representation as a model for distributed collections and its associated structural recursion scheme and monad as a model for parallel collection processing. At the source code level, Scala’s comprehension syntax over a bag monad can be used to encode Select-From-Where expressions in a standard way. At the intermediate representation level, maintaining comprehensions as a first-class citizen can be used to simplify the design and implementation of holistic dataflow optimizations that accommodate for nesting and control-flow. The proposed DSL design therefore reconciles the benefits of embedded parallel dataflow DSLs with the declarativity and optimization potential of external DSLs like SQL.

Funder

Deutsche Forschungsgemeinschaft

European Commission

Bundesministerium für Wissenschaft und Forschung

Oracle Labs

Publisher

Association for Computing Machinery (ACM)

Subject

Information Systems

Link

https://dl.acm.org/doi/pdf/10.1145/3281629

Reference67 articles.

1. The dataflow model

2. Implicit Parallelism through Deep Language Embedding

3. Implicit Parallelism through Deep Language Embedding

4. SSA is functional programming

Cited by 10 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Handling Iterations in Distributed Dataflow Systems;ACM Computing Surveys;2022-12-31

2. TranSQL: A Transformer-based Model for Classifying SQL Queries;2022 21st IEEE International Conference on Machine Learning and Applications (ICMLA);2022-12

3. Imperative or Functional Control Flow Handling;ACM SIGMOD Record;2022-05-31

4. A two-level formal model for Big Data processing programs;Science of Computer Programming;2022-03

5. Risk management of data flow under cross-cultural english language understanding;International Journal of System Assurance Engineering and Management;2022-01-30