Hello, everyone. Thank you for joining us, and welcome to the Global Neurodegenerative Disease Summit. Today's presentation will feature a talk by Dr. Ryan Corces focused on the use of single-cell epigenomics to reveal causal non-coding variants in neurodegenerative disease. My name is Kelly Miller, I'm with 10x, and I'll be moderating today's session. We have just one housekeeping slide before getting started. We do want to make our presentation as interactive as possible. This particular presentation will not include a live Q&A session, but we encourage you to submit your questions in the Q&A box adjacent to the slide window, and we'll follow up with you at a later date. Also, you can find a list of resources related to this week's topics in the resource list link on the right of your screen. Please note that all attendees are on mute. Also, this webinar is being recorded. We'll send you an email when the on-demand webinar recording is available for viewing. Just a reminder that today's presentation is part of a series. We have an impressive agenda this week that features discussions of some of the most impactful recent work and cutting-edge applications in the neurodegenerative disease field. I hope you'll make it to hear all the important talks that will be presented. I thank all of our amazing speakers in advance. I'd like to introduce our speaker for today. Dr. Corces started his scientific career at Princeton University, where he graduated with a focus in molecular biology and computer science. Ryan's thesis work at Stanford University in Dr. Ravi Majeti's lab centered on the genetic evolution of acute myeloid leukemia, where he showed that the earliest mutations that occur in AML affect genes that regulate the epigenome. Following this line of study, Ryan began his postdoc work in the laboratories of Dr. Howard Chang and Thomas Montine, studying epigenetics in human diseases, including neurodegeneration and cancer. Ryan joined the Gladstone Institute of Neurological Disease in July 2020 to study the contributions of genetic and non-genetic factors to neurodegenerative diseases. Using computational biology, large-scale screens, and single-cell technologies, Ryan's lab probes the epigenome of patient-derived cells with the aim of understanding the impact on disease risk and developing novel avenues for therapeutic interventions. Thank you so much for being with us, Ryan. I'll go ahead and turn it over so you can get started. Thanks so much for that introduction, and thank you all for being here and listening in to some of the work that my lab has done and has ongoing in the space of single-cell epigenomics, with a particular focus on using single-cell epigenomics to annotate the function of non-coding variants in neurodegenerative disease. My name's Ryan Corces, and I'm an assistant investigator at the Gladstone Institutes, and I just started my lab in the summer of 2020. It's really a pleasure to tell you about our pursuit to understand this puzzle of the nucleus and how the epigenome affects disease. My lab is broadly focused on this question of why do some people develop neurodegenerative disease while others do not? We focus on this through a few overarching themes. The one that I'll focus on today is the contribution of inherited genetics to the disease. We're also very focused on epigenomic aspects of the disease, which we view as cognitive resilience. The propensity of some individuals to remain resilient to cognitive decline or selective vulnerability, which would be the propensity of certain neurons to decline with these diseases. We're also very focused on developing new technologies or modifying existing technologies and developing new software and analytical paradigms. Really the overall goal of all of these studies is to create better prognostication strategies or new therapeutic targets. With that as the background for what my lab is interested in, we'll dive into this inherited genetics. How much do we really understand about the inherited genetics of late-onset Alzheimer's disease? Well, if you look at twin studies, there is approximately 60% of Alzheimer's disease is heritable and genetic. The vast majority of Alzheimer's can be explained by inherited genetics. Of course, this is somewhat confounded in that twin studies, the twins often share early environments, so that does obviously play a role as well. As an upper bound, we could consider about 60% genetic heritability in Alzheimer's. If we just look at common SNPs, those above some threshold in the population, about 33% of the phenotypic variance in Alzheimer's disease can be explained. A large portion of this is explained by the sole effect of APOE, which is a gene that harbors two coding polymorphisms and is a major risk factor for Alzheimer's disease. There's a large portion of phenotypic variance that we can't explain by the known genetics of Alzheimer's. The question becomes: where is this missing heritability? It exists in a couple of different places. It could be that there's less common SNPs that are affecting this. There could be structural variations, which are a very interesting topic these days. The vast majority of this is going to reside in the non-coding genome. I'll give you a little bit of background on why I make that argument. A genome-wide association study typically results in a plot like this. This is a Manhattan plot where the significance of association of a given variant with the disease is shown on the Y-axis, and the position along the linear genome on the X-axis. Any regions that fall above this red dotted line are significant. Of course, we see regions such as the APOE gene or BIN1, which are very well now studied, but there's many, many genes down here which are annotated as being associated with Alzheimer's, but we don't really understand their function. In large part, that's because those associations are driven by variants in the non-coding genome. Just to really drive this point home, this table represents what you might get out of a genome-wide association study where you have a list of SNPs, and they're annotated because they have the highest p-value, but there are often nearby SNPs which also have very significant p-values. With each of those SNPs is an annotated gene, and that's often annotated because it's the nearest gene. In the case of non-coding polymorphisms, it's very well understood that non-coding gene regulatory interactions can occur over very large distances, and so the nearest gene may or may not be the correct functionally relevant gene. To summarize this all in words, genome-wide association studies, they're very good at identifying large regions of the genome where genetic variation is associated with a particular disease. What they're not good at doing is determining which cell type is affected. They're not good at predicting which genes will be affected, especially in the case of non-coding variants, and they're really, really not good at pinpointing which SNP is functional. This is largely due to linkage disequilibrium, which if you remember back to your undergrad genetics class, is really the propensity of SNPs that are located nearby to be co-inherited together. I'll show you some data throughout the next 30 minutes or so where we kind of tick off these three points using single-cell ATAC-seq, using HiChIP and single-cell ATAC-seq, so three-dimensional chromosome confirmation capture techniques. Then also using machine learning to kind of tie all of this together. Within Alzheimer's and Parkinson's disease, and this largely extends to every disease that has been studied, the vast majority of genome-wide association study polymorphisms reside within the non-coding genome. Most loci have no plausible coding alteration that would explain their association with the disease. We are charged then with understanding how these non-coding polymorphisms can possibly be functional. The underlying hypothesis is that to be associated with the disease, a SNP has to exert an effect, otherwise it couldn't be associated. For a non-coding SNP, that means that it has to affect gene expression or splicing, because it's not going to affect the protein sequence. How does a sequence change in the non-coding genome affect gene expression? Let's take this toy example where you have a T allele and a C allele. Here, I'm highlighting a GATA transcription factor motif, so you can see that motif here. Wouldn't it be great if we could identify the regions in the genome where a transcription factor was bound or where a gene regulatory element was present so that we could determine whether or not this sequence change affected transcription factor binding and gene regulation? What we would want is to find those sites. We could map them to nearby genes. Of course, the way that we do this is using chromatin accessibility profiling. In In my lab, we use ATAC-seq, and so we find peaks of chromatin accessibility where transcription factors are bound and gene regulatory elements exist. Then we look under those peaks for sequence changes that may affect canonical transcription factor binding sites. This is the type of SNP that we're looking for in the context of a functional non-coding SNP. One of the ways that we identify these is through this concept of allelic accessibility. In this toy example, this GATA motif is only present on the T allele. When you have the C allele present, it disrupts that motif, the GATA transcription factor only binds to this particular allele. You see that allele overrepresented in your sequencing data, and this will come back into play later. We started this whole journey quite a few years ago now by doing bulk ATAC-seq, and we took samples from controls and Alzheimer's disease and Parkinson's disease individuals and profiled bulk ATAC-seq in these seven different regions. Of course, when we do dimensionality reduction on those samples, we can see that they generally group by the brain region of origin. This is largely, it turns out, due to different cell types present in those different regions. For example, in the striatal regions, you see the dopamine D2 receptor. In the nigral regions, you see markers of that part of the brain, in particular IRX3 transcription factor, et cetera. What we ended up finding was that these bulk assays showed very little significant difference between cases and controls. Of course, if we compare different regions, we can see significant differences. When we compare, for example, controls that have very low pathology to cognitively healthy individuals that have very high pathology, we see no significant differences. I'm only highlighting this to create a foil for why single-cell data is really important. Just to drive this point home, in the bulk ATAC-seq data, what we're missing is cell type specific signal. We know that chromatin accessibility is extremely cell type specific because gene regulation is highly cell type specific. To illustrate that point, here are ATAC-seq tracks of various different cell types. This is just around a random gene in the genome that I chose, IGF-1. I hope that you can appreciate that effectively every single cell type here, even though IGF-1 is expressed in most of these cells, they have very different ways in which they regulate the IGF-1 gene. You have excitatory neuron specific peaks, inhibitory neuron specific peaks, microglia, oligodendrocyte, astrocyte, et cetera. Hopefully this shows you that cell type specificity is really important for gene regulation. The question becomes: how do we obtain these cell type specific regulatory landscapes in the brain? Unlike in the blood, we don't have paradigms to FACS sort all of these very intricately defined cell types. One of the most effective ways that we have found to do this is to use single-cell profiling. This essentially needs no introduction in this webinar, but the way that this works with the 10x Genomics platform is that we transpose in bulk, we use the Chromium system to encapsulate individual nuclei with bar-coded gel beads. We do our amplification, split the GEMs, and then sequence all of the fragments that result, and map each of those fragments back to the cell of origin based on the barcode that they have. In the context of the brain, rather than flow sorting up front and doing bulk ATAC-seq of different populations, we can now do this all in one pot reaction and then identify the neurons, the glia, et cetera, based on their cell type specific signals. What does this look like? Here's a dimensionality reduction of about 70,000 single cells. We can call clusters and try to annotate those clusters. We do that in the chromatin accessibility space by making what we call gene activity scores, which are inferences of how highly expressed a gene might be based on its patterns of chromatin accessibility. This works quite well. You can identify excitatory neurons, inhibitory neurons, microglia, oligodendrocytes, OPCs, et cetera. To highlight how different this data is compared to the bulk ATAC-seq data, if we take all of the peaks identified in hundreds of bulk ATAC-seq samples from the brain and compare those to peaks identified in just 10 single-cell ATAC-seq reactions, we find almost twice as many peaks in the single-cell ATAC-seq data. We do capture the vast majority of peaks from the bulk ATAC-seq data with our single-cell ATAC-seq, but a very large portion of the single-cell ATAC-seq peaks are not captured by the bulk ATAC-seq peaks. You can see this in this heat map where of the 350,000 or so total peaks, approximately 220,000 are specific to one cell type or a pair of cell types. For example, these are specific to neurons, excitatory neurons, inhibitory neurons, microglia, et cetera. This is more than half of the peaks are cell type specific. When you think back to the bulk data, and we look at which of these cell type specific peaks are captured by the bulk, you can see that there's a strong underrepresentation of peaks from microglia, astrocytes, and OPCs, which happen to be the least abundant cell types. What this ends up showing is that cell types that are less than about 20% of your total sample are just completely missed by bulk profiling. Doing this sort of single-cell profiling really illuminates a lot of the cell type specific biology. To show you how far this can be taken, here is a dimensionality reduction just of the neurons in our data, and we can identify really fine-grained subclasses of neurons, including different subclasses of interneurons like somatostatin, parvalbumin, or VIP. We can identify even multiple subtypes of medium spiny neurons in the basal ganglia and similar phenomena. To kind of wrap up this section on why single-cell compared to bulk, we took that single-cell data and used it to deconvolve the bulk ATAC-seq data. We do this using CIBERSORT, which is a program developed at Stanford by Aaron Newman and Ash Alizadeh. What this basically does is takes a bulk ATAC-seq profile and splits it up into profiles that represent contributions from different cell types. Again, we do this using CIBERSORT. We get these sorts of signature matrices which define the cell types of interest. We can see that when we use this in the bulk data and compare to the known ground truth in the single-cell ATAC-seq data, it performs extremely well. When I run this sort of analysis across all of the bulk data that we obtained previously, you can see that there is a massive amount of heterogeneity in the cell type composition of these individual tissues. If you imagine trying to identify statistically significant differences across macro-dissected frozen brain, hopefully you can appreciate how much variability they would have across different samples and how problematic that would be for identifying statistically significant signals. Back to our plan to understand genome-wide association studies. One of the first things that we did was try to see if there's a specific cell type that's enriched for polymorphisms from Alzheimer's or Parkinson's disease or other cell types. It's very well known, at this point in time, that microglia, shown here in light blue, are enriched for these polymorphisms in Alzheimer's disease. This doesn't really tell us anything about a specific polymorphism, but about the disease in general. If we look at different neuronal subtypes, none of them are enriched for polymorphisms that would be associated with Alzheimer's or Parkinson's disease. This gets a little bit at the cell type level, but we still want to dig a little bit further and try to annotate which precise genes are being affected by each of the individual polymorphisms. For this, we've done HiChIP, which is a chromosome confirmation capture technique. The only one real difference between Hi-C and HiChIP, which Hi-C would capture all interactions, HiChIP uses an antibody enrichment to capture specific regions. In this case, we're using the active chromatin mark H3K27 acetylation, which will capture interactions between enhancers and promoters and other regions of the genome. The other way that we're going to map regulatory elements to the genes that they interact with is by using what's called co-accessibility. Imagine you have a promoter and many different enhancers, and you wanted to predict which of these enhancers might be affecting that gene's expression. You could look for situations in which the accessibility of the promoter was correlated with the accessibility of the enhancer. Hopefully, you can appreciate that the E3 enhancer is relatively well correlated with the accessibility at the promoter. You could plot this in multiple different ways, but the end story is that one of these enhancers has a highly correlated accessibility with the accessibility of the promoter. Using these two different orthogonal techniques, we can start to try to map these SNPs to potential genes. If we just take all of the lead SNPs from GWAS studies and we ask how many genes do we map them to, well, we map them to the nearest gene. It's one gene per SNP, and that's 51 at the time when the study was done. If we use the HiChIP and co-accessibility data, we can expand the realm of genes where these polymorphisms are mapping, and notably, about half of those predictions from the lead SNPs are actually incorrect. You see a very similar picture for Parkinson's disease. This hopefully gives you the feel that we can now map which genes might be affected by which SNPs. The last piece of this puzzle is really predicting which SNP is functional. Once we know that, we can map it to genes and cell types and all of that. The way that we're going to do this is through machine learning. Machine learning, what we're going to use it for is to predict this functionality. Here's the paradigm that we'll go through. First, we start with a lead SNP. We expand this in linkage disequilibrium to identify all of the SNPs that might be important. We start to whittle that list down. We overlap those with the peaks from our single-cell chromatin accessibility profiling because our hypothesis is that for a SNP to be functional in the non-coding genome, it has to affect a regulatory element which would be highlighted by one of these peaks. For the subset of those SNPs that affect peaks, we're going to try to predict the effect of that sequence change on transcription factor binding. In this toy example, you have a C to an A change here, and that might affect the binding. Once we do that, we'll use our HiChIP or co-accessibility data to map that particular SNP to the genes that it might be regulating. Really the crux of this is to predict the SNP effect. We take all of our clusters and we use a gapped k-mer support vector machine, which essentially just is learning the underlying grammars of chromatin accessibility so that you can feed it a wild type and variant sequence and look for differences in how those sequences may be bound by a transcription factor. Here you have the effect allele with predicted higher accessibility indicated by a higher height of these logos, the non-effect allele with a lower, and then you have this delta track, which should in theory highlight the motif that is present in that binding. Okay, I'll walk you through a few of these examples. First, the PICALM locus. This has classically been associated with microglia, but I'll try to provide some evidence that it may actually be affecting oligodendrocytes. Across a couple of GWAS studies, we have a few lead SNPs. When we expand those in LD, we see about 165 SNPs and 24 of those overlap peak regions, which I'll show you in a second. These are all in the vicinity of the PICALM gene, and that is why these SNPs have been annotated as affecting PICALM in the past. I'm going to convince you that it's this one SNP right here that has an effect. That SNP does overlap this prominent oligodendrocyte specific peak. It does interact both based on HiChIP and co-accessibility with the PICALM gene. We also see some evidence for it interacting downstream with this gene EED, which is part of a Polycomb group, which is another provocative hypothesis. When we look at the machine learning, we see a pretty strong prediction that this G to A change disrupts a Fos enhancer, which is shown here. Again, when you have the A allele, you have very little predicted accessibility. When you have the G allele, you have higher predicted accessibility. When you subtract those two tracks, you essentially pick up this beautiful Fos motif. Taking this a step further and going back to our bulk data and looking for allelic accessibility, you can see that in a large number of cases, all of which are heterozygotes here, and the ones shown in color are heterozygotes, the reference allele is more accessible than the variant allele. Here, the non-effect allele is the reference allele, and so the G is more accessible, more strongly bound by Fos than the variant allele. That is a very strong indication that this SNP is functional. How about a few other examples? I'll try to breeze through these rather quickly. In this particular locus, which is annotated as the KCNIP3 locus, we expand an LD, we cover a pretty large region of this locus with about 100 SNPs, and we can whittle those down to 22 that affect peak regions. I'm going to convince you or show you evidence that it could be either this SNP shown here in red or this one over here. They have two very different stories. This one on the left affects an oligodendrocyte specific peak. This one on the right affects a neuronal peak. The oligodendrocyte specific peak seems to interact with this MAL gene, which is a key gene for oligodendrocyte function. This SNP interacts with the KCNIP3 gene, which is known to be involved in neuronal function. On the oligodendrocyte side, we have support by a machine learning prediction where the effect allele is more accessible than the non-effect allele, and this maps very strongly to a SOX6 motif, where SOX6 is a known regulator of oligodendrocyte function. On the neuronal side, we have enough individuals where we can find evidence of allelic accessibility. Unfortunately, for this SNP, we did not have enough individuals that were heterozygous. These two SNPs give two very different interpretations of what's going on. I'll give you one last example, and that's in a much better understood locus, which is BIN1. Here, these two SNPs could potentially have function. They both affect microglial specific peaks, and they both, with some evidence, interact with the promoter of BIN1. However, we believe that this SNP, rs13025717, is the causative SNP because in our machine learning prediction, we find a very strong loss of accessibility with the effect allele, and that maps very strongly to a KLF4 motif, which is a known regulator of microglial identity. This has all been published, and we're actively working on validating some of these findings. How do we go about validating that one of these predictions from the machine learning side of things is actually important and functional in the disease? The first way that we're attempting to do this is with what we call scarless single-base editing. The idea here is that you take a wild-type allele, and you use some gene editing approach, in our case, we're using prime editing, to change that in an isogenic fashion from a G to an A, and then you differentiate these cells and test allelic differences in gene expression. We've been able to do this for a variety of loci. I'm just showing you some Sanger traces here, showing an A to a G conversion if it's the G allele, or a G to an A conversion. I will say that prime editing has been a little bit finicky and is quite locus dependent, so we're still actively working on a lot of this. The other way that we're using functional genomics to validate some of these findings is through massively parallel reporter assays. What these assays do is you have a wild-type version and a variant version of a particular regulatory element, and you clone that upstream of a minimal promoter and an open reading frame, and you use sequencing to determine the differential activity of these two alleles in a reporter assay. You do this across thousands of different allelic transcripts in this assay. This, of course, has the advantages of being very high throughput. You can do this in any cell type that you can query and grow in cell culture. This can really be applied to any disease, and it does give you a quantitative readout. The disadvantage here is that it's not in situ in the correct location in the genome, so the correct in situ context isn't maintained. These are things that we're actively pursuing and hoping to have more to share in the future. I'll share one more vignette with you about how we can use this sort of epigenomic data to understand pretty complex associations, and in this case, with Parkinson's disease. MAPT, which is the gene that encodes the tau protein, is actually one of the strongest GWAS loci associated with Parkinson's disease, even though we canonically think of tau as an Alzheimer's disease protein. The MAPT locus is very interesting from a genetic and evolutionary standpoint. Many, many years ago, there was an inversion that occurred in this locus, which occurs between this location and this location, which swaps the orientation of everything inside the inversion with respect to those things outside of the inversion. Along with that inversion, there are a few thousand SNPs that are also different between these two haplotypes. If you inherit a copy of the H2 haplotype, this region is flipped, and you also have thousands of SNPs within that region. This creates a very complicated problem to understand is it the inversion that's causing the disease association? Is it one of these SNPs? How do we understand the epigenetics underlying this particular complex association? We can use publicly available transcriptome data to look at how the expression of the MAPT gene, for example, changes with these different haplotypes. There is a significant loss of MAPT expression in the H2 haplotype, but it's relatively mild. What we sought to explain was how does this gene expression change occur from an epigenetic standpoint? We're going to do an allelic comparison of these two haplotypes. We have an H1 haplotype homozygotes, H2 homozygotes, and heterozygotes. For the H1 and H2, we can just handle them as is. We don't have to split their reads. For the H1 H2 heterozygotes, we're actually going to split the reads based on the SNPs that occur here in this haplotype region into the H1 and H2 reads. In both cases, we're going to do some differential testing. This case in particular is interesting because it's exquisitely well-controlled. The H1 and H2 reads are coming from the same individuals from the same cells. Here is that same locus. I'm highlighting for you the promoter of the MAPT gene. I'm going to animate for you on top of this thing. Here I'm showing you bulk ATAC-seq data so you can see high accessibility at the MAPT promoter, and you can see some changes in accessibility that are haplotype specific. This peak is H1 specific. These peaks are significantly higher in H2 haplotype, we'll focus on those two regions and we'll call these other regions A and B. When we look at this B region down here, there's a very strong increase in interaction frequency based on HiChIP with this A region upstream. Now we'll follow these throughout. This dotted line represents our focal point for the HiChIP assay, where everything on this Y-axis is relative to its interaction with that point. Now we can move upstream to the MAPT promoter, and you can see a slight increase in interaction with this H1 specific peak. If we shift to that H1 specific peak, you can see highly increased interaction with the promoter and with this downstream MAPT enhancer. Of course, if we shift all the way over to this A region, you see the reciprocal high interaction strength with the B region downstream. All of this goes to say that there are big changes both in chromatin accessibility and in 3D enhancer-promoter interactions that are changing in this locus. Those do have changes in the corresponding gene expression. Here I'm showing you each gene as a bar. Some of them you don't see because they're just at zero because they have no difference in H1 or H2. These genes down here are upregulated in H2 individuals, and these are upregulated in H1 individuals. Hopefully you can appreciate that whatever is happening with these sorts of interactions between the A and the B regions is definitely changing gene expression. We don't particularly think that this gene expression is driving the association because these are largely pseudogenes and antisense transcripts, but it is a relatively provocative mechanism. This is all happening inside of the breakpoints, but what happens outside of the breakpoints? That's interesting because the region inside these breakpoints is actually being inverted in the H2 haplotype. Here I'm showing you the focal point on the MAPT gene, and if we look upstream, both in homozygous individuals where we haven't done any read splitting and in heterozygotes with the allelic read splitting, there's an increased interaction with this region here, which could be a potential long-range enhancer. When we look at that region, we see multiple neuron-specific peaks that are likely enhancers. We believe that these regions are interacting over a long distance with the MAPT promoter to increase its expression, specifically in the H1 individuals compared to the H2 individuals. To put this schematically in the H1 individuals, you have the inversion in this direction and a long-range interaction between this enhancer and the promoter and a production of more MAPT transcript. In the H2 haplotype, this A and B region, they're inverted, and they're also interacting with much higher frequency, which seems to insulate the MAPT gene from this distal enhancer. Hopefully this gives you a flavor for how this sort of integrative epigenomic analysis can really provide some interesting and informative interpretations of different genetic associations with disease. Why should you care if you're not in this field? I do think it's important that we continue to identify new genetic targets, and this, of course, has translational implications. Those new genes that we implicate in the disease will nominate new molecular mechanisms. They'll provide new insights into this multifactorial genetic interaction that's occurring between different variants that we inherit. Hopefully long-term, this leads to better prognostication. Just in the last few minutes, I want to highlight some of the work that we've been doing in this kind of large-scale multi-omics and in particular on the software development side. ATAC-seq, in particular single-cell ATAC-seq, there's not a lot of analytical approaches that really enable these sorts of massive scale analyses. Some of that is due to challenges that are present in ATAC-seq in particular. If you were to compare ATAC-seq and RNA-seq, the number of features is very different. Most people end up with a few thousand to 10,000 genes per cell. For ATAC-seq, this is much, much higher. There's hundreds of thousands of regulatory elements across the genome. This becomes really complicated because the feature set in ATAC-seq is also dynamic. Different cell types might express the same genes. There's only a limited number of genes in the genome, but the number of regulatory elements is 10 or morefold higher and is extremely cell type specific. As you add new cell types to your analysis, the feature set changes, and the problem gets larger and larger with ATAC-seq, whereas for RNA-seq, there's a certain number of features, and the feature set doesn't really change. This creates computational challenges, where the matrices and the analyses that are being performed require some pretty fine-tuned solutions to really enable the scale of data generation that's now possible with the 10x Genomics system. Our solution to this has been to develop this software package called ArchR, which is highly robust and scales very well to what we consider to be massive scale data sets in the millions of cells. On a standard MacBook laptop now, you can analyze about a million cells through dimensionality reduction and clustering in under eight hours. This really enables the scale of data generation that's now possible. We think that the underlying infrastructure of ArchR is highly intuitive. It's super well annotated with a manual. At the risk of continuing this infomercial, I'll just leave it at that and say you should visit archrproject.com to check it out if you haven't already. To summarize what I've told you today, the tissues and cell types, especially in the brain, are extremely complicated and multifactorial, and using single-cell assays can really help understand what's important about that different cell type-specific biology. Hopefully, I've shown you that the non-coding genome matters, the work that we're doing really heavily relies on single-cell chromatin accessibility profiling to identify these functional non-coding mutations. I showed you one example where this vertical integration of multi-omic data can really provide key insights into disease predisposition. This is not an isolated example. There's many examples where integrating these types of data are hugely important. I think that the field of machine learning really only adds to this. The field is really developing quite quickly, and the insights that we'll gain by combining assays and layering on machine learning are really impressive in my opinion. Lastly, I told you about our efforts to support the new age of single-cell genomics by developing new tools and analytical paradigms. With that, I'll just thank the people involved. My lab, as I mentioned, is quite new, and we're actively growing. We're very thankful for the generous funding that we received from multiple sources. Thank you for your attention, and I look forward to the rest of the seminars. Thanks so much for a great talk, Ryan. If any of you have questions for Dr. Corces or for 10x, you'll have five more minutes to submit those questions in the Q&A box, and we'll get back to you at a later time with the reply. You can also contact us with any follow-up questions you have at the email address support@10xgenomics.com. Thank you so much for joining us, and have a great rest of your day.
Loading workspace