Search Swinburne Research Bank
Home
List of Titles
A cost-effective strategy for intermediate data storage in scientific cloud workflow systems
List of Titles
A cost-effective strategy for intermediate data storage in scientific cloud workflow systems
Please use this identifier to cite or link to this item: http://hdl.handle.net/1959.3/88596
- Title
- A cost-effective strategy for intermediate data storage in scientific cloud workflow systems
- Author(s)
- Yuan, Dong; Yang, Yun; Liu, Xiao; Chen, Jinjun
- Abstract
- Many scientific workflows are data intensive where a large volume of intermediate data is generated during their execution. Some valuable intermediate data need to be stored for sharing or reuse. Traditionally, they are selectively stored according to the system storage capacity, determined manually. As doing science on cloud has become popular nowadays, more intermediate data can be stored in scientific cloud workflows based on a pay-for-use model. In this paper, we build an Intermediate data Dependency Graph (IDG) from the data provenances in scientific workflows. Based on the IDG, we develop a novel intermediate data storage strategy that can reduce the cost of the scientific cloud workflow system by automatically storing the most appropriate intermediate datasets in the cloud storage. We utilise Amazon's cost model and apply the strategy to an astrophysics pulsar searching scientific workflow for evaluation. The results show that our strategy can reduce the overall cost of scientific cloud workflow execution significantly.
- Publication type
- Conference paper
- Research centre
- Swinburne University of Technology. Faculty of Information and Communication Technologies
- Source
- Proceedings of the 2010 IEEE International Symposium on Parallel & Distributed Processing (IPDPS 2010), Atlanta, Georgia, United States, 19-23 April 2010
- Publication year
- 2010
- FOR Code(s)
- 0805 Distributed Computing
- Keyword(s)
- Cloud computing; Cost; Data storage; Scientific workflow
- Publisher
- IEEE
- ISSN
- 1530-2075
- ISBN
- 9781424464432, 1424464439
- Publisher URL
- http://dx.doi.org/10.1109/IPDPS.2010.5470453
- Copyright
- Copyright © 2010 IEEE. Published version of the paper reproduced here in accordance with the copyright policy of the publisher. Personal use of this material is permitted. However, permission to reprint/republish this material for advertising or promotional purposes or for creating new collective works for resale or redistribution to servers or lists, or to reuse any copyrighted component of this work in other works must be obtained from the IEEE.
- Research Projects
-
Novel cloud computing based workflow technology for managing large numbers of process instances, Australian Research Council grant number LP0990393
- Full text

- Peer reviewed


