SIM-PIPE DryRunner: An approach for testing container-based big data pipelines and generating simulation data

Abstract

Big data pipelines are becoming increasingly vital in a wide range of data intensive application domains such as digital healthcare, telecommunication, and manufacturing for efficiently processing data. Data pipelines in such domains are complex and dynamic and involve a number of data processing steps that are deployed on heterogeneous computing resources under the realm of the Edge-Cloud paradigm. The processes of testing and simulating big data pipelines on heterogeneous resources need to be able to accurately represent this complexity. However, since big data processing is heavily resource-intensive, it makes testing and simulation based on historical execution data impractical. In this paper, we introduce the SIM - PIPE Dry Runner approach - a dry run approach that deploys a big data pipeline step by step in an isolated environment and executes it with sample data; this approach could be used for testing big data pipelines and realising practical simulations using existing simulators.

Read publication

Client

Research Council of Norway (RCN) / 323325
Research Council of Norway (RCN) / 309691
EC/H2020 / 101016835

Language

English

Author(s)

Affiliation

SINTEF Digital / Sustainable Communication Technologies
OsloMet - Oslo Metropolitan University

Year

2022

Published in

Computer Software and Applications Conference

ISSN

0730-3157

Publisher

IEEE conference proceedings

Page(s)

1159 - 1164

External resources

https://github.com/datacloud-project/sim-pipe

DOI

https://doi.org/10.1109/compsac54236.2022.00182

Read fulltext

https://hdl.handle.net/11250/3055097

View this publication at Cristin

Contact us

Our services

Career

Sustainability

Management and board

Institutes

Other units

About us

Follow us