NDLI: Distilling structure in Taverna scientific workflows: a refactoring approach

Content Provider	Springer Nature : BioMed Central
Author	Cohen-Boulakia, Sarah Chen, Jiuqiang Missier, Paolo Goble, Carole Williams, Alan R Froidevaux, Christine
Abstract	Background Scientific workflows management systems are increasingly used to specify and manage bioinformatics experiments. Their programming model appeals to bioinformaticians, who can use them to easily specify complex data processing pipelines. Such a model is underpinned by a graph structure, where nodes represent bioinformatics tasks and links represent the dataflow. The complexity of such graph structures is increasing over time, with possible impacts on scientific workflows reuse. In this work, we propose effective methods for workflow design, with a focus on the Taverna model. We argue that one of the contributing factors for the difficulties in reuse is the presence of \"anti-patterns\", a term broadly used in program design, to indicate the use of idiomatic forms that lead to over-complicated design. The main contribution of this work is a method for automatically detecting such anti-patterns, and replacing them with different patterns which result in a reduction in the workflow's overall structural complexity. Rewriting workflows in this way will be beneficial both in terms of user experience (easier design and maintenance), and in terms of operational efficiency (easier to manage, and sometimes to exploit the latent parallelism amongst the tasks). Results We have conducted a thorough study of the workflows structures available in Taverna, with the aim of finding out workflow fragments whose structure could be made simpler without altering the workflow semantics. We provide four contributions. Firstly, we identify a set of anti-patterns that contribute to the structural workflow complexity. Secondly, we design a series of refactoring transformations to replace each anti-pattern by a new semantically-equivalent pattern with less redundancy and simplified structure. Thirdly, we introduce a distilling algorithm that takes in a workflow and produces a distilled semantically-equivalent workflow. Lastly, we provide an implementation of our refactoring approach that we evaluate on both the public Taverna workflows and on a private collection of workflows from the BioVel project. Conclusion We have designed and implemented an approach to improving workflow structure by way of rewriting preserving workflow semantics. Future work includes considering our refactoring approach during the phase of workflow design and proposing guidelines for designing distilled workflows.
Related Links	https://bmcbioinformatics.biomedcentral.com/counter/pdf/10.1186/1471-2105-15-S1-S12.pdf
Ending Page	14
Page Count	14
Starting Page	1
File Format	HTM / HTML
ISSN	14712105
DOI	10.1186/1471-2105-15-S1-S12
Journal	BMC Bioinformatics
Issue Number	1
Volume Number	15
Language	English
Publisher	BioMed Central
Publisher Date	2014-01-10
Access Restriction	Open
Subject Keyword	Bioinformatics Microarrays Computational Biology Computer Appl. in Life Sciences Algorithms Output Port Input Port Recursive Call Trace Link Dataflow Model Computational Biology/Bioinformatics
Content Type	Text
Resource Type	Article
Subject	Molecular Biology Biochemistry Computer Science Applications Applied Mathematics Structural Biology
Journal Impact Factor	2.9/2023
5-Year Journal Impact Factor	3.6/2023

Sl.	Authority	Responsibilities	Communication Details
1	Ministry of Education (GoI), Department of Higher Education	Sanctioning Authority	https://www.education.gov.in/ict-initiatives
2	Indian Institute of Technology Kharagpur	Host Institute of the Project: The host institute of the project is responsible for providing infrastructure support and hosting the project	https://www.iitkgp.ac.in
3	National Digital Library of India Office, Indian Institute of Technology Kharagpur	The administrative and infrastructural headquarters of the project	Dr. B. Sutradhar bsutra@ndl.gov.in
4	Project PI / Joint PI	Principal Investigator and Joint Principal Investigators of the project	Dr. B. Sutradhar bsutra@ndl.gov.in Prof. Saswat Chakrabarti will be added soon
5	Website/Portal (Helpdesk)	Queries regarding NDLI and its services	support@ndl.gov.in
6	Contents and Copyright Issues	Queries related to content curation and copyright issues	content@ndl.gov.in
7	National Digital Library of India Club (NDLI Club)	Queries related to NDLI Club formation, support, user awareness program, seminar/symposium, collaboration, social media, promotion, and outreach	clubsupport@ndl.gov.in
8	Digital Preservation Centre (DPC)	Assistance with digitizing and archiving copyright-free printed books	dpc@ndl.gov.in
9	IDR Setup or Support	Queries related to establishment and support of Institutional Digital Repository (IDR) and IDR workshops	idr@ndl.gov.in

Distilling structure in Taverna scientific workflows: a refactoring approach.

Distilling structure in Taverna scientific workflows: a refactoring approach

Distilling structure in Taverna scientific workflows: a refactoring approach

SegMine workflows for semantic microarray data analysis in Orange4WS

Bioinformatics recipes: creating, executing and distributing reproducible data analysis workflows

Classification of bioinformatics workflows using weighted versions of partitioning and hierarchical clustering algorithms

Tavaxy: Integrating Taverna and Galaxy workflows with cloud computing support

Scientific workflow optimization for improved peptide and protein identification

Performing statistical analyses on quantitative data in Taverna workflows: An example using R and maxdBrowse to identify differentially-expressed genes from microarray data

Distilling structure in Taverna scientific workflows: a refactoring approach

Similar Documents

Distilling structure in Taverna scientific workflows: a refactoring approach.

Distilling structure in Taverna scientific workflows: a refactoring approach

Distilling structure in Taverna scientific workflows: a refactoring approach

SegMine workflows for semantic microarray data analysis in Orange4WS

Bioinformatics recipes: creating, executing and distributing reproducible data analysis workflows

Classification of bioinformatics workflows using weighted versions of partitioning and hierarchical clustering algorithms

Tavaxy: Integrating Taverna and Galaxy workflows with cloud computing support

Scientific workflow optimization for improved peptide and protein identification

Performing statistical analyses on quantitative data in Taverna workflows: An example using R and maxdBrowse to identify differentially-expressed genes from microarray data

Distilling structure in Taverna scientific workflows: a refactoring approach