One Gene at a Time: How Perturb-seq Maps Cause and Effect
Editing One Gene, Reading the Whole Cell: How Perturb-seq Maps Cause and Effect
Imagine walking into a control room filled with thousands of unlabeled switches. Several of them are flipped on at the same time, but nothing tells you which switch controls what. Maybe one switch turns on all the others. Maybe something behind the wall is flipping all of them together. You could stare at that panel for years and still be guessing.
Biologists run into almost exactly this problem when they study genes, and for a long time they had no good way to start flipping switches on purpose.
Scientists can already observe which genes are active inside a cell. They can see that a certain set of genes turns on during an immune response, during a disease, or at a specific stage of development. The catch is that watching two genes turn on together does not prove that one of them controls the other. Correlation is not causation, and inside a human cell with roughly 20,000 genes there are a lot of ways to fool yourself.
Perturb-seq was built for this exact problem. It combines CRISPR gene editing with single-cell RNA sequencing, which means researchers can change, or “perturb,” one gene and then read how thousands of other genes respond inside individual cells. Basically, scientists edit one part of the system and then listen to the entire system react.
From Watching Cells to Testing Them
Traditional gene-expression studies are mostly observational. Researchers collect cells, measure their RNA, and compare the patterns they find. That work is genuinely useful, because RNA tells you which genes a cell is actually using at a given moment rather than which genes it happens to carry. But observation has a ceiling, and the ceiling is low.
Say Gene A and Gene B are both highly active in cancer cells. Gene A might activate Gene B, or Gene B might activate Gene A. A third gene might control both of them. They might also be active at the same time for no connected reason at all. An observational study can rank these possibilities, but it cannot settle the question.
Perturb-seq lets researchers actually test the relationship. Scientists can shut down Gene A, then look at what happens to Gene B. If Gene B changes, that is much stronger evidence that Gene A influences it. The word “perturb” just means to disturb or change something, so a Perturb-seq experiment is exactly what it sounds like: disturb a gene, then measure the damage.
I think this shift is the part worth sitting with. The technology is impressive, but the real change is philosophical. Biology spent decades as a watching science in this area, and Perturb-seq turns a piece of it into a doing science.
How Perturb-seq Works
A Perturb-seq experiment starts with CRISPR. CRISPR uses a molecule called a single-guide RNA, or sgRNA, to steer a CRISPR protein toward a specific stretch of DNA. From there, scientists have options. They can disable a gene completely with a CRISPR knockout, they can lower a gene’s activity without cutting the DNA using CRISPR interference (CRISPRi), or they can push a gene’s activity up using CRISPR activation.
The trick that makes the method scalable is pooling. Researchers build thousands of different guide RNAs, mix them into a single library, and deliver that library into a large population of cells at once. The concentration is usually tuned so that each cell picks up one guide, which means each cell ends up with one gene targeted. Instead of running thousands of separate experiments, you run one experiment containing thousands of tiny experiments.
Of course, that creates a bookkeeping nightmare. If every cell in the dish got a different guide, how do you know which cell got which one?
Perturb-seq solves this by giving each guide an identifiable sequence that works like a barcode. After the CRISPR system has had time to take effect, researchers run single-cell RNA sequencing, which separates individual cells and measures the RNA inside each one. The barcode gets read alongside everything else, so the final dataset links two pieces of information for every single cell: which gene was targeted, and the expression level of thousands of genes in that same cell.
Researchers then compare edited cells against control cells that received non-targeting guides, which should not affect any gene. If the cells carrying a guide against Gene A consistently shift the same pathway, a map of Gene A’s effects starts to appear.
Why Single-Cell Data Matters
Before single-cell sequencing became common, researchers usually studied thousands or millions of cells together in what is called bulk sequencing. Bulk sequencing works, but it produces an average, and averages hide things.
Cells are not identical, even when they come from the same tissue. Some are dividing, some are not. Some are responding to a signal, some are ignoring it. A perturbation might hit one group of cells hard and do almost nothing to another group sitting right next to it. In a bulk experiment, those two effects blur into one middling number that describes neither population.
Perturb-seq measures cells one at a time, so the differences survive. This makes it possible to find out that the same gene plays different roles in different cell types. A developmental gene might be essential in an immature neuron and close to irrelevant in a mature one, and only single-cell data can show you that.
This is one of the biggest strengths of the method, and honestly it is the part I underestimated when I first read about it. Perturb-seq does not only ask, “What does this gene do?” It asks, “What does this gene do in this cell type, under these conditions, at this moment?” That is a much harder question, and it is also much closer to how biology actually behaves.
From Small Experiments to Millions of Cells
Perturb-seq arrived in 2016 from several groups working at the same time on gene regulation, cellular stress, and immune cells (Dixit et al., Cell, 2016; Adamson et al., Cell, 2016; Jaitin et al., Science, 2016; and the closely related CROP-seq from Datlinger et al., Nature Methods, 2017). By modern standards these first experiments were small. Dixit and colleagues targeted about two dozen transcription factors, and Adamson and colleagues studied around ten perturbations in the unfolded protein response using roughly 30,000 cells. What they proved was the concept: you could run a pooled CRISPR screen and still walk away with detailed information from every individual cell.
The scale grew fast after that. In 2022, Replogle and colleagues published a genome-scale Perturb-seq study covering more than 2.5 million human cells (Cell, 2022). In K562 leukemia cells they used CRISPRi to target 9,866 expressed genes, sampled eight days after the guides were delivered, plus a second screen against 2,057 common essential genes at day six. They then repeated an essential-gene screen in RPE1 cells, a non-cancerous line, to check whether the patterns held up outside a cancer background.
Numbers like that are hard to picture. Every one of those 2.5 million cells contributes its own full transcriptome, so the dataset holds billions of measurements. The useful thing is what you can do once the data exists. Researchers can group genes into pathways and predict what poorly understood genes are doing, and the Replogle screen used exactly this approach to identify a previously unrecognized subunit of the Integrator complex, a protein machine involved in RNA processing.
The logic behind that kind of discovery is simple, which is part of why I like it. Suppose you knock down two different genes and the resulting RNA changes look nearly identical. That similarity suggests the two genes work in the same pathway or the same protein complex, even if nobody knew anything about one of them beforehand.
In my opinion this is the coolest thing about Perturb-seq. Genetics has always had a popularity problem, where the famous genes get studied over and over while thousands of others sit in the genome with names like C7orf26 and almost no literature attached. Perturb-seq gives those genes a way in. You do not need a hypothesis about a mystery gene to learn something about it. You just need it to behave like something you already understand.
New Versions of Perturb-seq
The original method has real limitations, and researchers have been building variations to work around them.
Compressed Perturb-seq attacks cost. A standard experiment needs many cells per perturbation, which gets expensive quickly. Yao and colleagues borrowed ideas from compressed sensing and designed experiments that measure multiple random perturbations per cell, or multiple cells per droplet, then computationally separate the effects afterward using the fact that regulatory circuits are sparse (Nature Biotechnology, 2023). They reported cutting costs by roughly an order of magnitude compared with existing strategies. It is a bit like recording several instruments playing at once and using software to work out what each one contributed.
In vivo Perturb-seq moves the whole experiment into a living animal instead of a dish. This matters because cells behave differently inside real tissue, where they receive signals from blood vessels, immune cells, neighbors, and the physical structure around them. Jin and colleagues used CRISPR-Cas9 to introduce frameshift mutations in 35 autism and neurodevelopmental delay risk genes, in pools, inside the developing mouse neocortex before birth, then sequenced the cells in the postnatal brain (Science, 2020). They found effects in both neurons and glia, and several different risk genes converged on shared gene modules. More recent work has pushed the same idea further into cortical development at larger scale (Zheng et al., Cell, 2024).
Multiome versions measure RNA and chromatin accessibility in the same cell. Chromatin is the DNA-protein structure that packages the genome, and some regions of it are open and readable while others are closed off. Reading both layers at once, as in Perturb-ATAC (Rubin et al., Cell, 2019), lets researchers watch a change in DNA accessibility and the change in gene expression that follows from it.
Spatial approaches add the piece that standard single-cell sequencing throws away, which is location. To sequence cells individually you normally have to dissociate the tissue, and at that point you know what each cell was but not where it sat or who its neighbors were. Perturb-map kept that information by tagging perturbed cells with protein barcodes and imaging intact tissue in a mouse lung cancer model, which allowed the researchers to see how each knockout changed the tumor and the immune cells around it (Dhainaut, Rose et al., Cell, 2022).
That last one is the direction I find most interesting, because a lot of the genes that matter in disease are not doing their work inside a single cell. They are changing how cells talk to each other, and you cannot see a conversation if you have already separated everyone in the room.
Does Perturb-seq Prove Causation?
Not on its own, no. It gives much stronger causal evidence than an observational study, because the scientists actually intervene rather than waiting for nature to vary. But it does not hand anyone a finished map.
Plenty can go wrong. A guide RNA might fail to change its target gene, or it might hit an unintended stretch of DNA instead. Single-cell sequencing captures only a fraction of the RNA molecules in each cell, so missing data is normal rather than exceptional, and this is one reason the Replogle team filtered out perturbations backed by too few cells before analyzing them. Timing matters as well, since a measurement taken early may capture the direct response while a later one shows stress, adaptation, or cells dying.
The deepest limitation is more conceptual. If changing Gene A changes Gene B, that still does not prove Gene A controls Gene B directly. Gene A might change Gene C, and Gene C changes Gene B. Perturb-seq reads the endpoint of a chain without necessarily showing the links in between, so follow-up experiments are still required to pin down the mechanism.
I think it is worth being honest about this, because the scale of the datasets can make the results feel more final than they are. Millions of cells sounds like proof. It is really just a very good starting point, produced very quickly.
A New Way to Understand Cells
Perturb-seq does not reduce a cell to a solved machine. Cells are far messier than that, and their behavior depends on their genes, their environment, their history, their location, and everything happening in the cells around them.
What Perturb-seq changes is the kind of question a scientist can realistically ask. Instead of only watching which genes light up together, researchers can switch one off and record how the rest of the cell responds, then repeat that thousands of times in a single pooled experiment and assemble the results into a map of pathways.
It comes back to that control room. Someone is finally testing the switches one at a time instead of standing there taking notes on the blinking lights. The map that comes out of it will not be complete, and some of the wiring will stay hidden behind the wall. But with every perturbation, the panel makes a little more sense than it did before.
Sources
Adamson, B. et al. (2016). A multiplexed single-cell CRISPR screening platform enables systematic dissection of the unfolded protein response. Cell.
Datlinger, P. et al. (2017). Pooled CRISPR screening with single-cell transcriptome readout. Nature Methods.
Dhainaut, M., Rose, S. A. et al. (2022). Spatial CRISPR genomics identifies regulators of the tumor microenvironment. Cell, 185(7), 1223-1239. doi:10.1016/j.cell.2022.02.015
Dixit, A. et al. (2016). Perturb-Seq: Dissecting molecular circuits with scalable single-cell RNA profiling of pooled genetic screens. Cell.
Jaitin, D. A. et al. (2016). Dissecting immune circuits by linking CRISPR-pooled screens with single-cell RNA-seq. Science.
Jin, X. et al. (2020). In vivo Perturb-Seq reveals neuronal and glial abnormalities associated with autism risk genes. Science, 370(6520). doi:10.1126/science.aaz6063
Replogle, J. M. et al. (2022). Mapping information-rich genotype-phenotype landscapes with genome-scale Perturb-seq. Cell, 185, 2559-2575. doi:10.1016/j.cell.2022.05.013
Rubin, A. J. et al. (2019). Coupled single-cell CRISPR screening and epigenomic profiling reveals causal gene regulatory networks. Cell.
Yao, D. et al. (2023). Scalable genetic screening for regulatory circuits using compressed Perturb-seq. Nature Biotechnology. doi:10.1038/s41587-023-01964-9
Zheng, X. et al. (2024). Massively parallel in vivo Perturb-seq reveals cell-type-specific transcriptional networks in cortical development. Cell. doi:10.1016/j.cell.2024.04.050

