Showing posts with label BRANE. Show all posts
Showing posts with label BRANE. Show all posts

September 27, 2018

BRANE power: Gene network inference with graph optimization

And the IPFEN 2018 Yves Chauvin PhD prize is awarded to Aurélie Pirayre, for IFPEN first thesis on bioinformatics with graph optimization for gene networks. In French, you can now read ‘‘BRANE Power’’ : gènes et algorithmes, une alliance pour la chimie verte.

Aurélie Pirayre, Grégoire Allaire, Didier Houssin, Pierre-Henri Bigeard, Eric Heintzé

PhD manuscript and slides for her thesis:
  • Reconstruction and clustering with graph optimization and priors on gene networks and images (manuscript)
    • Abstract: The discovery of novel gene regulatory processes improves the understanding of cell phenotypic responses to external stimuli for many biological applications, such as medicine, environment or biotechnologies. To this purpose, transcriptomic data are generated and analyzed from DNA microarrays or more recently RNAseq experiments. They consist in genetic expression level sequences obtained for all genes of a studied organism placed in different living conditions. From these data, gene regulation mechanisms can be recovered by revealing topological links encoded in graphs. In regulatory graphs, nodes correspond to genes. A link between two nodes is identified if a regulation relationship exists between the two corresponding genes. Such networks are called Gene Regulatory Networks (GRNs). Their construction as well as their analysis remain challenging despite the large number of available inference methods. In this thesis, we propose to address this network inference problem with recently developed techniques pertaining to graph optimization. Given all the pairwise gene regulation information available, we propose to determine the presence of edges in the final GRN by adopting an energy optimization formulation integrating additional constraints. Either biological (information about gene interactions) or structural (information about node connectivity) a priori have been considered to restrict the space of possible solutions. Different priors lead to different properties of the global cost function, for which various optimization strategies, either discrete and continuous, can be applied. The post-processing network refinements we designed led to computational approaches named BRANE for \Biologically-Related A priori for Network Enhancement". For each of the proposed methods --- BRANE Cut, BRANE Relax and BRANE Clust --- our contributions are threefold: a priori-based formulation, design of the optimization strategy and validation (numerical and/or biological) on benchmark datasets from DREAM4 and DREAM5 challenges showing numerical improvement reaching 20%. In a ramification of this thesis, we slide from graph inference to more generic data processing such as inverse problems. We notably invest in HOGMep, a Bayesian-based approach using a Variational Bayesian Approximation framework for its resolution. This approach allows to jointly perform reconstruction and clustering/segmentation tasks on multi-component data (for instance signals or images). Its performance in a color image deconvolution context demonstrates both quality of reconstruction and segmentation. A preliminary study in a medical data classification context linking genotype and phenotype yields promising results for forthcoming bioinformatics adaptations.
  • Slides (PhD defense on July 3rd, 2017)
are finally online (check the EURASIP Library of Ph.D. Theses). The work was ruled by the concept of BRANE power; a methodology for gene regulatory network inference and clustering based on graph optimization and biological priors. BRANE stands for Biologically Related Apriori Network Enhancement. It rhymes with cell membrane (and brain, for who it's worth).

Gene regulatory network inference with BRANE Cut

State-of-the-art results are obtained on synthetic and real transcriptomic data (DREAM-4, DREAM-5 for DREAM consortium challenges, Escherichia coli dataset). Derived methods are BRANE Cut (with graph cuts), BRANE Relax (with proximal optimization) and BRANE Clust (with graph Laplacian). 

Gene network joint inference and clustering with BRANE Clust



Used concepts include:
  • data science, optimization on graphs: maximal flow, minimum cut, random walker algorithm, variational and Bayes variational formalism, convex relaxation, alternating optimization, combinatorial Dirichlet problem, hard-clustering and soft-clustering
  • biology, biotechnology, bioinformatics: transcription factors (TFs) as regulators and non-transcription factors (TFs) as targets, modular networks, biological priors, in-silico data, second generation bio-fuel production, DREAM4 challenge, DREAM5 challenge
  • use to biofuels and green chemistry production (with fungus Trichoderma reesei)
Supervising team:
PhD Thesis reporters
PhD Thesis Examiners
More links:

April 11, 2017

BRANE Clust: cluster-assisted gene regulatory network inference refinement

The joined Gene Regulatory Network (GRN)  inference and clustering tool BRANE Clust has just been published in BRANE Clust: cluster-assisted gene regulatory network inference refinement in IEEE/ACM Transactions on Computational Biology and Bioinformatics, 2017 (doi:10.1109/TCBB.2017.2688355).

It is also featured on RNA-seq blog and OMIC tools.

Alternative versions are available as a preprint, on biorxiv, with a page and software and in HAL. Another brick in the BRANE series wall, a series of bioinformatics tools based on graphs and optimization, dedicated to -omics gene expression data for GRN (Gene Regulatory Network) inference.

While traditional Next-generation sequencing (NGS) pipelines often combine motley assumptions (correlation, normalization, clustering, inference), this work is an first step toward gracefully combining network inference and clustering. 

BRANE Clust works as a post-processing tool upon classical network thresholding refinement. From a complete weighted network (obtained from any network inference method) BRANE Clust favors edges both having higher weights (as in standard thresholding) and linking nodes belonging to a same cluster. It  relies on an optimization procedure. It  computes an optimal gene clustering (random walker algorithm) and an optimal edge selection jointly. The introduction of a clustering step in the edge selection process improves gene regulatory network inference. This is demonstrated on both synthetic (five networks of  DREAM4 and network 1 of DREAM5) and real (network 3 of DREAM5) data. These conclusions are drawn after comparing classical thresholding on CLR and GENIE3 networks to our proposed post-processing. Significant improvements in terms of Area Under Precision-Recall curve are obtained. The  predictive power on real data yields promising results: predicted links specific to BRANE Clust reveal plausible biological interpretation. GRN approaches that produce a complete weighted network to prune could benefit from BRANE Clust post-processing.

Escherichia coli network built using BRANE Clust on GENIE3 weights and containing 236 edges. Large dark gray nodes refers to transcription factors (TFs). Inferred edges also reported in the ground truth are colored in black while predictive edges are light gray. Dashed edges correspond to a link inferred by both BRANE Clust and GENIE3 while solid links refer to edges specifically inferred by BRANE Clust.
Abstract:
Discovering meaningful gene interactions is crucial for the identification of novel regulatory processes in cells.
Building accurately the related graphs remains challenging due to the large number of possible solutions from available data. Nonetheless, enforcing a priori on the graph structure, such as modularity, may reduce network indeterminacy issues. BRANE Clust (Biologically-Related A priori Network Enhancement with Clustering) refines gene regulatory network (GRN) inference thanks to cluster information. It works as a post-processing tool for inference methods (i.e. CLR, GENIE3). In BRANE Clust, the clustering is based on the inversion of a system of linear equations involving a graph-Laplacian matrix promoting a modular structure. Our approach is validated on DREAM4 and DREAM5 datasets with objective measures, showing significant comparative improvements. We provide additional insights on the discovery of novel regulatory or co-expressed links in the inferred Escherichia coli network evaluated using the STRING database. The comparative pertinence of clustering is discussed computationally (SIMoNe, WGCNA, X-means) and biologically (RegulonDB). BRANE Clust software is available at:
http://www-syscom.univ-mlv.fr/~pirayre/Codes-GRN-BRANE-clust.html

Film-opéra-concert Ariodante

  #Ariodante de #Händel par les Arts Florissants  en opéra-concert-film est #amazing ; trois raisons, deux futiles.  c'est 16 euros, ...