Skip to main navigation Skip to search Skip to main content

Automated Planning for Optimal Data Pipeline Instantiation

  • Leonardo Rosa Amado
  • , Adriano Vogel
  • , Dalvan Griebler
  • , Gabriel Paludo Licks
  • , Eric Simon
  • , Felipe Meneguzzi
  • Pontifícia Universidade Católica do Rio Grande do Sul
  • University of Rome La Sapienza

Research output: Working paperPreprint

1 Downloads (Pure)

Abstract

Data pipeline frameworks provide abstractions for implementing sequences of data-intensive transformation operators, automating the deployment and execution of such transformations in a cluster. Deploying a data pipeline, however, requires computing resources to be allocated in a data center, ideally minimizing the overhead for communicating data and executing operators in the pipeline while considering each operator's execution requirements. In this paper, we model the problem of optimal data pipeline deployment as planning with action costs, where we propose heuristics aiming to minimize total execution time. Experimental results indicate that the heuristics can outperform the baseline deployment and that a heuristic based on connections outperforms other strategies.
Original languageEnglish
PublisherArXiv
Number of pages8
DOIs
Publication statusPublished - 16 Mar 2025

Bibliographical note

We are grateful for the financial support and computing resources from SAP Labs.

Funding

We are grateful for the financial support and computing resources from SAP Labs.

Keywords

  • cs.AI
  • cs.DC

Fingerprint

Dive into the research topics of 'Automated Planning for Optimal Data Pipeline Instantiation'. Together they form a unique fingerprint.

Cite this