System Biology and Systems Medicine




The mechanisms behind life are inherently complex. In the last years, increasing research activity has shown that our understaning of life, and human diseases, is not complete if only one source of information (proteomics, genomics, transcriptomics, metabolomics, connectomics, etc) is used.

Systems biology is an approach which accounts for different sources of information to build an integrate model of life and disease. In this context, methodologies based con complex systems science and network science are required.



Multilayer networks approach to molecular biology



A fairly standard approach to the analysis of human diseases is based on the enrichment of topological and functional connectivity representing protein-protein interaction networks.

Figure: PPI network and disease. The concept of disease modules exemplified using a sample PPIN. One or more topological modules (highlighted red) contain proteins involved in similar biological processes forming functional modules (highlighted blue). A disease module (highlighted green) is a sub-network of proteins enriched with disease-relevant proteins, e.g., known disease associated proteins. (Figure and caption from this source)


Multilayer networks provide one of the most promising for modeling systems biology and for their analysis.

I have verified this claim quantitatively, by building multilayer networks where nodes are genes and proteins, links represent their interactions, and each layer encodes a different relation (physical chemical, genetic, including regulatory, inhibitory, etc.) and quantify the overall structural reducibility of the resulting system.

Figure: mesoscale organization of the human multilayer PPI network. Changes in the meso-scale organization of human multiplex proteome for varying relax rate. Clusters with at least 100 proteins are considered for clarity. This alluvial plot shows how partitions split and merge for increasing rate: larger clusters are quite stable, highlighting that differences in meso-scale are mostly due to smaller sets of proteins. (Figure and caption from our paper.)


References:

More recently, I have applied this framework to map complex disease-gene and disease-symptom interactions to a multiplex networks with two layers. Layers encode phenotypic and genotypic relations, allowing a better characterization and classification of human diseases.


Figure: Multiplex disease network. (A) Two bipartite networks of disease-gene and disease-symptom interactions are projected onto diseases, (B) where diseases are connected in the genotype layer (blue) if they share a common gene and connected in the phenotype layer (green) if they share a symptom. (C) The two networks are considered as layers of a multiplex system, where nodes are the diseases and colored links encode their interactions. Disease-disease interactions that are present in both layers are named overlapping links.


Data »

In parallel, I was part of an extraordinary international team effort to assess network module identification across complex diseases.
Many bioinformatics methods have been proposed for reducing the complexity of large gene or protein networks into relevant subnetworks or modules. Yet, how such methods compare to each other in terms of their ability to identify disease-relevant modules in different types of network remains poorly understood. We launched the ‘Disease Module Identification DREAM Challenge’, an open competition to comprehensively assess module identification methods across diverse protein–protein interaction, signaling, gene co-expression, homology and cancer-gene networks. Predicted network modules were tested for association with complex traits and diseases using a unique collection of 180 genome-wide association studies. Our robust assessment of 75 module identification methods reveals top-performing algorithms, which recover complementary trait-associated modules. We find that most of these modules correspond to core disease-relevant pathways, which often comprise therapeutic targets. This community challenge establishes biologically interpretable benchmarks, tools and guidelines for molecular network analysis to study human disease biology.

Figure: Assessment of module identification methods. a, Main types of module identification approach used in the challenge. b, Final scores of the 42 module identification methods applied in Sub-challenge 1 for each of the six networks, as well as the overall score summarizing performance across networks (evaluated using the holdout GWAS set at 5% FDR; method IDs are defined in Supplementary Table 2). Ranks are indicated for the top ten methods. The last row shows the mean performance of 17 random modularizations of the networks (error bars show the standard deviation). c, Robustness of the overall ranking was evaluated by subsampling the GWAS set used for evaluation 1,000 times. For each method, the resulting distribution of ranks is shown as a boxplot. d, Number of trait-associated modules per network. Boxplots show the number of trait-associated modules across the 42 methods, normalized by the size of the respective network.
Learn more »

References:

Application to COVID-19

This approach is promising for applications in both network and systems medicine. Recently, we have built a multilayer network for SARS-CoV-2, the virus of COVID-19, by integrating data sets for Proteins, Diseases, Drugs, and Symptoms, focusing on the part of the human interactome targeted by viral proteins. We recapitulate many of the known symptoms of the disease and we find the most similar diseases to COVID-19 reflect conditions that are risk factors in patients. In particular, the comparison between CovMulNet19 and randomized networks recovers many of the known associated comorbidities that are important risk factors for COVID-19 patients, through identified similarities with intestinal, hepatic, and neurological diseases as well as with respiratory conditions, in line with reported comorbidities. CovMulNet19 can be suitably used for network medicine analysis, as a valuable tool for exploring drug repurposing while accounting for the intervening multidimensional factors, from molecular interactions to symptoms.

Figure: CovMulNet19 COVID-19 genotype–phenotype–drug interaction network. Result of the data integration and processing procedures illustrated schematically in Figure 1. (A) Nodes and schematic map of interdependencies among different layers encoding diseases, symptoms, drugs, GO terms, human proteins, and viral proteins. (B) Map of the reconstructed structural interactions (e.g., protein–protein) and functional interdependencies (e.g., protein–disease, protein–GO term, or disease–symptom). Overall, the network consists of 1999 protein–protein, 19,755 protein–disease, 10,152 protein–symptom, 13,018 drug–target, 9210 protein–GO, and 3056 disease–symptom relationships.
Interactive Viz » Data »

Multiscale statistical physics of the pan-viral interactome unravels the systemic nature of SARS-CoV-2 infections


Figure: Impact of virus interactions with the human interactome at micro, meso and macroscopic scales. Schematic illustration of virus-host interactions across scales, where viral proteins attack human protein targets and the corresponding effects are investigated with distinct techniques from statistical physics of complex networks. Addition of viral components dG to the human Protein-Protein Interaction (PPI) network G generates the virus-human interactome G'.
Micro: percolation analysis evaluates static structural properties (S(G') size of the giant connected component) and robustness of the interactome under removal of proteins. Here, the underlying biological hypothesis is that viral proteins might inhibit the usual function of human targets, and we map this activity into the removal of protein from the system. We also test another less invasive hypothesis: the viral proteins interact with the human targets while altering, and not just inhibiting, their functions: the resulting perturbations are propagated (dashed lines mimicking the propagation) and we analyze the system response.
Meso: in this case, the underlying hypothesis is that viral proteins alter the function of the human interactome at the mesoscopic level, i.e., interfering with the functional organization in modules (green shaded areas) typical of biomolecular systems. This interference is mapped into the isolation of the target proteins, and the modular and hierarchical re-organization of the interactome is detected according to two popular methods for community and hierarchy detection.
Macro: viral interactions dG perturb macroscopic properties of the interactome which are captured by the analysis of the network density matrix von Neumann entropy, Massieu function and energy functions at temporal scale beta.



Figure: Virus-host interactome as an interdependent network. BIOSTR Human PPI (Protein-Protein Interactions) used in this study, is obtained from data fusion of two comprehensive public repositories, namely STRING and BIOGRID (see the text for details). The network consists of N=19,945 proteins linked by |E|=737,668 edges, and the largest connected component (99.8% nodes, 99.6% edges) is shown (a). Proteins targeted by viruses are highlighted in two ways. On the one hand, markers of distinct size identify targeted proteins: bigger the marker larger the number of times a protein is targeted by viruses in our data set. On the other hand, distinct colored markers of constant size encode distinct viruses (93 in total, including SARS-CoV-2) and the same color scheme is used to show the contribution of each virus to one of the most frequently targeted proteins (b), TP53, as an example.
Multiscale statistical physics \rev{of the pan-viral interactome unravels the systemic nature of SARS-CoV-2 infections} }

References:

Multilayer networks approach to computational neuroscience



After proposing a novel multilayer representation of fuctional connectivity in human brain, I have shown that multilayer analysis allows to identify crucial regions (ie, hubs) that can be used to increase accuracy in distinguishing between healthy and schizophrenic patients from resting-state fMRI measurements. If you are interesting in knowing more about multilayer approaches to model and to analyze the human brain, you can read my very recent review on the topic.


Figure: Multilayer functional brain. Brain activity is measured in different regions and signals are decomposed in the frequency domain. The frequency domain consists of (possibly overlapping) frequency bands and, for each band, coherence -- or other similarity descriptors -- is measured between all pairs of regions. A similarity matrix is built for each frequency domain and statistical analysis of significance is used to map each matrix into a functional network, constituting a functional layer of the overall multilayer system.



Figure: Visualizing the multilayer functional brain. Three-dimensional representations of the multilayer functional brain of a schizophrenic subject, based on frequency decomposition (11 layers, non-overlapping frequency bands between 0.01~Hz and 0.23~Hz). Only links with at least 6 standard deviations from the mean are shown. Top panels: edge-colored representation, where connections are colored according to the frequency band and node size is proportional to their functional versatility. Bottom panel: multi-slice representation, where each layer encodes information about a specific frequency bandand inter-layer connectivity is not shown explicitly for sake of simplicity. The color scheme is the same in the two representations.




References: