Gerstein Lab Abstracts

Saturday, May 17, 2014

Fwd: CSHL Meeting Abstract Submission

Thank you for submitting your abstract for Systems Biology: Global
Regulation of Gene Expression 2014 .

Comparative network analysis of ENCODE and modENCODE data

Mark Gerstein1,2, Koon-Kiu Yan1,2, modENCODE/ENCODE transcriptome group1

1Yale University, Program of Computational Biology and Bioinformatics,
New Haven, CT, 2Yale University, Molecular Biophysics and
Biochemistry, New Haven, CT

The ENCODE and modENCODE consortia have generated a resource
containing large amounts of transcriptomic data, extensive mapping of
chromatin states, as well as the binding locations of over 300
transcription-regulatory factors for human, worm and fly. We performed
extensive data integration by constructing genome-wide co-expression
networks and transcriptional regulatory networks, revealing
fundamental principles of transcription and network organization that
are conserved across the three highly divergent animals. First, we
developed a novel cross-species clustering algorithm to integrate the
co-expression networks of the three species, resulted at conserved
modules shared between the organisms. These modules are enriched in
developmental genes and exhibited hourglass behavior. They were then
used to align the stages in worm and fly development, finding the
normal embryo-to-embryo and larvae-to-larvae pairings in addition to a
novel pairing between worm embryo and fly pupae. Second, we defined a
new score to quantify the degree of hierarchy (preponderance of
downward information flow) of a network, and developed a global
optimization algorithm to compare the hierarchical organization of the
three species. We found that, despite extensive rewiring of binding
targets, high-level organization principles like the hierarchical
structures are conserved across three species. Finally, we found the
gene expression levels in the organisms, both coding and non-coding,
can be predicted consistently based on their upstream histone marks.
In fact, a "universal model" with a single set of cross-organism
parameters can predict expression level for both protein coding genes
and ncRNAs. The algorithms introduced by this study can be easily
applied to other model datasets such as those from yeast.

i0cshsb

Thursday, December 5, 2013

Abstract for Talk at University of Connecticut Health Center (i0ccam)

Human Genome Analysis

Plummeting sequencing costs have led to a great increase in the number
of personal genomes. Interpreting the large number of variants in
them, particularly in non-coding regions, is a central challenge for
genomics. We investigate patterns of selection in DNA elements from
the ENCODE project using the full spectrum of sequence variants from
1,092 individuals in the 1000 Genomes Project Phase 1, including
single-nucleotide variants (SNVs), short insertions and deletions
(indels) and structural variants (SVs). We analyze both coding and
non-coding regions, with the former corroborating the latter. We
identify a specific sub-group of non-coding categories that exhibit
very strong selection constraint, comparable to coding genes:
"ultra-sensitive" regions. We also find variants that are disruptive
due to mechanistic effects on transcription-factor binding (i.e.
"motif-breakers").

We make great use of networks -- contrasting them with linear
annotation -- and describe how we construct a practical instantiation
of the human regulatory network. Using connectivity information
between elements from protein-protein interaction and regulatory
networks, we find that variants in regions with higher network
centrality tend to be deleterious. Indels and SVs follow a similar
pattern as SNVs, with some notable exceptions (e.g. certain deletions
and enhancers).

Using these results, we develop a scheme and a practical tool to
prioritize non-coding variants based on their potential deleterious
impact. As a proof of principle, we experimentally validate and
characterize a small number of candidate variants prioritized by the
tool. Application of the tool to ~90 cancer genomes (breast, prostate
and medulloblastoma) reveals ~100 candidate non-coding cancer drivers.
This approach can be readily used in precision medicine to prioritize
variants.

References

Architecture of the human regulatory network derived from ENCODE data.
Gerstein et al. Nature 489: 91
[encodenets.gersteinlab.org]

Interpretation of genomic variants using a unified biological network approach.
E Khurana, Y Fu, J Chen, M Gerstein (2013). PLoS Comput Biol 9: e1002886.

Integrative annotation of variants from 1,092 humans: application to
cancer genomics
E Khurana et al. (2013) Science 342:1235587.
[funseq.gersteinlab.org]

Monday, December 2, 2013

RE: Abstract for Talk at Keystone Big Data Symposium (i0keybdata)

Dear Dr. Gerstein,

Thank you for your abstract. Co you please provide me with an abstract title and if there are more authors then just you, please send me their names and institutes.

Thank you and have a good day!

Jenny Hindorff
Keystone Symposia
Program Implementation Associate
www.keystonesymposia.org
jennyh@keystonesymposia.org
970-262-2661

-----Original Message-----
From: Mark Gerstein [mailto:mark@gersteinlab.org]
Sent: Saturday, November 30, 2013 11:45 AM
To: Programs
Cc: glabstracts.mbglab@blogger.com
Subject: Abstract for Talk at Keystone Big Data Symposium (i0keybdata)

My talk will discuss Human Genome Analysis from a data science perspective.

Plummeting sequencing costs have led to a great increase in the number of personal genomes. Interpreting the large number of variants in them, particularly in non-coding regions, is a central challenge for genomics.

One data science construct that is particularly useful for genome interpretation is networks. My talk will be concerned with the analysis of networks and the use of networks as a "next-generation annotation" for interpreting personal genomes. I will initially describe current approaches to genome annotation in terms of one-dimensional browser tracks. Here I will discuss approaches for annotating pseudogenes and also for developing predictive models for gene expression.
Then I will describe various aspects of networks. In particular, I will touch on the following topics: (1) I will show how analyzing the structure of the regulatory network indicates that it has a hierarchical layout with the "middle-managers" acting as information-flow bottlenecks and with more "influential" TFs on top. (2) I will show that most human variation occurs at the periphery of the network. (3) I will compare the topology and variation of the regulatory network to the call graph of a computer operating system, showing that they have different patterns of variation. (4) I will talk about web-based tools for the analysis of networks (TopNet and tYNA).

http://networks.gersteinlab.org
http://tyna.gersteinlab.org

Architecture of the human regulatory network derived from ENCODE data.
Gerstein et al. Nature 489: 91

Classification of human genomic regions based on experimentally determined binding sites of more than 100 transcription-related factors.
KY Yip et al. (2012). Genome Biol 13: R48.

Understanding transcriptional regulation by integrative analysis of transcription factor binding data.
C Cheng et al. (2012). Genome Res 22: 1658-67.

The GENCODE pseudogene resource.
B Pei et al. (2012). Genome Biol 13: R51.

Comparing genomes to computer operating systems in terms of the topology and evolution of their regulatory control networks.
KK Yan et al. (2010). Proc Natl Acad Sci U S A 107:9186-91.

Saturday, November 30, 2013

Abstract for Talk at Keystone Big Data Symposium (i0keybdata)

My talk will discuss Human Genome Analysis from a data science perspective.

Plummeting sequencing costs have led to a great increase in the number
of personal genomes. Interpreting the large number of variants in
them, particularly in non-coding regions, is a central challenge for
genomics.

One data science construct that is particularly useful for genome
interpretation is networks. My talk will be concerned with the
analysis of networks and the use of
networks as a "next-generation annotation" for interpreting personal
genomes. I will initially describe current approaches to genome
annotation in terms of one-dimensional browser tracks. Here I will discuss
approaches for annotating pseudogenes and also
for developing predictive models for gene expression.
Then I will describe various aspects of networks. In particular, I will touch on
the following topics: (1) I will show how analyzing the structure of
the regulatory network indicates that it has a hierarchical layout
with the "middle-managers" acting as information-flow bottlenecks and
with more "influential" TFs on top. (2) I will show that most human
variation occurs at the periphery of the network. (3) I will compare
the topology and variation of the regulatory network to the call graph
of a computer operating system, showing that they have different
patterns of variation. (4) I will talk about web-based tools for the
analysis of networks (TopNet and tYNA).

http://networks.gersteinlab.org
http://tyna.gersteinlab.org

Architecture of the human regulatory network derived from ENCODE data.
Gerstein et al. Nature 489: 91

Classification of human genomic regions based on experimentally
determined binding sites of more than 100 transcription-related
factors.
KY Yip et al. (2012). Genome Biol 13: R48.

Understanding transcriptional regulation by integrative analysis of
transcription factor binding data.
C Cheng et al. (2012). Genome Res 22: 1658-67.

The GENCODE pseudogene resource.
B Pei et al. (2012). Genome Biol 13: R51.

Comparing genomes to computer operating systems in terms of the
topology and evolution of their regulatory control networks.
KK Yan et al. (2010). Proc Natl Acad Sci U S A 107:9186-91.

Friday, November 29, 2013

Fwd: Pot. Abstract for Talk at ASHG '14 (i0ashg14)

Network Analysis for Human Genomics

Plummeting sequencing costs have led to a great increase in the number
of personal genomes. Interpreting the large number of variants in
them, particularly in non-coding regions, is a central challenge for
genomics. We investigate patterns of selection in DNA elements from
the ENCODE project using the full spectrum of sequence variants from
1,092 individuals in the 1000 Genomes Project Phase 1, including
single-nucleotide variants (SNVs), short insertions and deletions
(indels) and structural variants (SVs). We analyze both coding and
non-coding regions, with the former corroborating the latter. We
identify a specific sub-group of non-coding categories that exhibit
very strong selection constraint, comparable to coding genes:
"ultra-sensitive" regions. We also find variants that are disruptive
due to mechanistic effects on transcription-factor binding (i.e.
"motif-breakers").

We make great use of networks -- contrasting them with linear
annotation -- and describe how we construct a practical instantiation
of the human regulatory network. Using connectivity information
between elements from protein-protein interaction and regulatory
networks, we find that variants in regions with higher network
centrality tend to be deleterious. Indels and SVs follow a similar
pattern as SNVs, with some notable exceptions (e.g. certain deletions
and enhancers).

Using these results, we develop a scheme and a practical tool to
prioritize non-coding variants based on their potential deleterious
impact. As a proof of principle, we experimentally validate and
characterize a small number of candidate variants prioritized by the
tool. Application of the tool to ~90 cancer genomes (breast, prostate
and medulloblastoma) reveals ~100 candidate non-coding cancer drivers.
This approach can be readily used in precision medicine to prioritize
variants.

References

Architecture of the human regulatory network derived from ENCODE data.
Gerstein et al. Nature 489: 91
[encodenets.gersteinlab.org]

Interpretation of genomic variants using a unified biological network approach.
E Khurana, Y Fu, J Chen, M Gerstein (2013). PLoS Comput Biol 9: e1002886.

Integrative annotation of variants from 1,092 humans: application to
cancer genomics
E Khurana et al. (2013) Science 342:1235587.
[funseq.gersteinlab.org]

Saturday, October 19, 2013

Fwd: Abstract for Talk at Duke (i0duke)

Sunday, September 8, 2013

Abstract for Talk at Georgia Tech (i0gatech)

Human Genome Analysis: Application to Cancer

Plummeting sequencing costs have led to a great increase in the number
of personal genomes. Interpreting the large number of variants in
them, particularly in non-coding regions, is a central challenge for
genomics. We investigate patterns of selection in DNA elements from
the ENCODE project using the full spectrum of sequence variants from
1,092 individuals in the 1000 Genomes Project Phase 1, including
single-nucleotide variants (SNVs), short insertions and deletions
(indels) and structural variants (SVs). We analyze both coding and
non-coding regions, with the former corroborating the latter. We
identify a specific sub-group of non-coding categories that exhibit
very strong selection constraint, comparable to coding genes:
"ultra-sensitive" regions. We also find variants that are disruptive
due to mechanistic effects on transcription-factor binding (i.e.
"motif-breakers").

We make great use of networks -- contrasting them with linear
annotation -- and describe how we construct a practical instantiation
of the human regulatory network. Using connectivity information
between elements from protein-protein interaction and regulatory
networks, we find that variants in regions with higher network
centrality tend to be deleterious. Indels and SVs follow a similar
pattern as SNVs, with some notable exceptions (e.g. certain deletions
and enhancers).

Using these results, we develop a scheme and a practical tool to
prioritize non-coding variants based on their potential deleterious
impact. As a proof of principle, we experimentally validate and
characterize a small number of candidate variants prioritized by the
tool. Application of the tool to ~90 cancer genomes (breast, prostate
and medulloblastoma) reveals ~100 candidate non-coding cancer drivers.
This approach can be readily used in precision medicine to prioritize
variants.

References

Architecture of the human regulatory network derived from ENCODE data.
Gerstein et al. Nature 489: 91

Interpretation of genomic variants using a unified biological network approach.
E Khurana, Y Fu, J Chen, M Gerstein (2013). PLoS Comput Biol 9: e1002886.

Integrative annotation of variants from 1,092 humans: application to
cancer genomics
E Khurana et al. Science (in press)