Sunday, September 8, 2013
Abstract for Talk at Georgia Tech (i0gatech)
Plummeting sequencing costs have led to a great increase in the number
of personal genomes. Interpreting the large number of variants in
them, particularly in non-coding regions, is a central challenge for
genomics. We investigate patterns of selection in DNA elements from
the ENCODE project using the full spectrum of sequence variants from
1,092 individuals in the 1000 Genomes Project Phase 1, including
single-nucleotide variants (SNVs), short insertions and deletions
(indels) and structural variants (SVs). We analyze both coding and
non-coding regions, with the former corroborating the latter. We
identify a specific sub-group of non-coding categories that exhibit
very strong selection constraint, comparable to coding genes:
"ultra-sensitive" regions. We also find variants that are disruptive
due to mechanistic effects on transcription-factor binding (i.e.
"motif-breakers").
We make great use of networks -- contrasting them with linear
annotation -- and describe how we construct a practical instantiation
of the human regulatory network. Using connectivity information
between elements from protein-protein interaction and regulatory
networks, we find that variants in regions with higher network
centrality tend to be deleterious. Indels and SVs follow a similar
pattern as SNVs, with some notable exceptions (e.g. certain deletions
and enhancers).
Using these results, we develop a scheme and a practical tool to
prioritize non-coding variants based on their potential deleterious
impact. As a proof of principle, we experimentally validate and
characterize a small number of candidate variants prioritized by the
tool. Application of the tool to ~90 cancer genomes (breast, prostate
and medulloblastoma) reveals ~100 candidate non-coding cancer drivers.
This approach can be readily used in precision medicine to prioritize
variants.
References
Architecture of the human regulatory network derived from ENCODE data.
Gerstein et al. Nature 489: 91
Interpretation of genomic variants using a unified biological network approach.
E Khurana, Y Fu, J Chen, M Gerstein (2013). PLoS Comput Biol 9: e1002886.
Integrative annotation of variants from 1,092 humans: application to
cancer genomics
E Khurana et al. Science (in press)
Thursday, May 16, 2013
Abstract for Talk at Harvard CCSB (i0farb)
networks as a "next-generation annotation" for interpreting personal
genomes. I will initially describe current approaches to genome
annotation in terms of one-dimensional browser tracks. Here I will discuss
approaches for annotating pseudogenes and also
for developing predictive models for gene expression.
Then I will describe various aspects of networks. In particular, I will touch on
the following topics: (1) I will show how analyzing the structure of
the regulatory network indicates that it has a hierarchical layout
with the "middle-managers" acting as information-flow bottlenecks and
with more "influential" TFs on top. (2) I will show that most human
variation occurs at the periphery of the network. (3) I will compare
the topology and variation of the regulatory network to the call graph
of a computer operating system, showing that they have different
patterns of variation. (4) I will talk about web-based tools for the
analysis of networks (TopNet and tYNA).
http://networks.gersteinlab.org
http://tyna.gersteinlab.org
Architecture of the human regulatory network derived from ENCODE data.
Gerstein et al. Nature 489: 91
Classification of human genomic regions based on experimentally
determined binding sites of more than 100 transcription-related
factors.
KY Yip et al. (2012). Genome Biol 13: R48.
Understanding transcriptional regulation by integrative analysis of
transcription factor binding data.
C Cheng et al. (2012). Genome Res 22: 1658-67.
The GENCODE pseudogene resource.
B Pei et al. (2012). Genome Biol 13: R51.
Comparing genomes to computer operating systems in terms of the
topology and evolution of their regulatory control networks.
KK Yan et al. (2010). Proc Natl Acad Sci U S A 107:9186-91.
Monday, March 25, 2013
Abstract for talk i0cmg - Comp_Meth_Prioritizing_Var_Exome_Seq_Pgenes_trueLoF_nethubs--20130319-i0cmg
-Certain genes have lots of similar pseudogenes
which could confound variant calling
* Finding True LoF Mutations
-Not just stop codon finding: tricky if one takes into
account splicing, NMD, indels, &c
* Using High Network Connectivity
-More connected genes in many networks have a
greater chance of being disease causing
Abstract for Talk at Chicago (i0chi12)
networks as a "next-generation annotation" for interpreting personal
genomes. I will initially describe current approaches to genome
annotation in terms of one-dimensional browser tracks. Here I will discuss
approaches for annotating pseudogenes and also
for developing predictive models for gene expression.
Then I will describe various aspects of networks. In particular, I will touch on
the following topics: (1) I will show how analyzing the structure of
the regulatory network indicates that it has a hierarchical layout
with the "middle-managers" acting as information-flow bottlenecks and
with more "influential" TFs on top. (2) I will show that most human
variation occurs at the periphery of the network. (3) I will compare
the topology and variation of the regulatory network to the call graph
of a computer operating system, showing that they have different
patterns of variation. (4) I will talk about web-based tools for the
analysis of networks (TopNet and tYNA).
http://networks.gersteinlab.org
http://tyna.gersteinlab.org
Architecture of the human regulatory network derived from ENCODE data.
Gerstein et al. Nature 489: 91
Classification of human genomic regions based on experimentally
determined binding sites of more than 100 transcription-related
factors.
KY Yip et al. (2012). Genome Biol 13: R48.
Understanding transcriptional regulation by integrative analysis of
transcription factor binding data.
C Cheng et al. (2012). Genome Res 22: 1658-67.
The GENCODE pseudogene resource.
B Pei et al. (2012). Genome Biol 13: R51.
Comparing genomes to computer operating systems in terms of the
topology and evolution of their regulatory control networks.
KK Yan et al. (2010). Proc Natl Acad Sci U S A 107:9186-91.
Abstract for Genome_Annotation_Compare_n_Func_Description_SVs_n_Nets--20130322-i0simons
the Population
Methods
Read-depth: MSB+CNVnator
Breakpoints & Split Read: SRiC, AGE & BreakSeq
Applications : 1000G & Somatic Variation
2 ## A Networks View on Large-scale Organization of Genomic Elements
Understanding the human regulatory network as a hierarchy with
information flow bottlenecks
Understanding the impact of variation and constraint on the network
Particularly with network analogies
Thursday, October 11, 2012
Abstract for Talk at Harvard (i0hsph)
My talk will be concerned with the analysis of networks and the use of
networks as a "next-generation annotation" for interpreting personal
genomes. I will initially describe current approaches to genome
annotation in terms of one-dimensional browser tracks. Here I will discuss
approaches for annotating pseudogenes and also
for developing predictive models for gene expression.
Then I will describe various aspects of networks. In particular, I will touch on
the following topics: (1) I will show how analyzing the structure of
the regulatory network indicates that it has a hierarchical layout
with the "middle-managers" acting as information-flow bottlenecks and
with more "influential" TFs on top. (2) I will show that most human
variation occurs at the periphery of the network. (3) I will compare
the topology and variation of the regulatory network to the call graph
of a computer operating system, showing that they have different
patterns of variation. (4) I will talk about web-based tools for the
analysis of networks (TopNet and tYNA).
http://networks.gersteinlab.org
http://tyna.gersteinlab.org
Architecture of the human regulatory network derived from ENCODE data.
Gerstein et al. Nature 489: 91
Classification of human genomic regions based on experimentally
determined binding sites of more than 100 transcription-related
factors.
KY Yip et al. (2012). Genome Biol 13: R48.
Understanding transcriptional regulation by integrative analysis of
transcription factor binding data.
C Cheng et al. (2012). Genome Res 22: 1658-67.
The GENCODE pseudogene resource.
B Pei et al. (2012). Genome Biol 13: R51.
Comparing genomes to computer operating systems in terms of the
topology and evolution of their regulatory control networks.
KK Yan et al. (2010). Proc Natl Acad Sci U S A 107:9186-91.
Thursday, October 4, 2012
2012 HUPO Abstract Entry Confirmation 294
Your abstract for the HUPO2012 conference was submitted on 5/3/2012.The log number for your abstract is 294. | ||||||||||||||||||||
| Analysis of Protein Networks | ||||||||||||||||||||
| Mark Gerstein | ||||||||||||||||||||
| Yale Comp. Bio., New Haven, CT | ||||||||||||||||||||
| Abstract My talk will be concerned with understanding protein function on a genomic scale. My lab approaches this through the prediction and analysis of biological networks, focusing on protein-protein interaction and transcription-factor-target ones. I will describe how these networks can be determined through integration of many genomic features and how they can be analyzed in terms of various topological statistics. In particular, I will discuss a number of recent analyses: (1) Improving the prediction of molecular networks through systematic training-set expansion; (2) Showing how the analysis of biochemical pathways across environments potentially allows them to act as biosensors; (3) Analyzing the structure of the regulatory network indicates that it has a hierarchical layout with the "middle-managers" acting as information bottlenecks; (4) Integrating the protein-interaction network with molecular structures and motions; (5) Showing the some motions are conflicting with protein-protein interactions and (6) Creating practical web-based tools for the analysis of these networks (DynaSIN and tYNA).
| ||||||||||||||||||||
| ||||||||||||||||||||