Say you’re a technician and you find yourself standing in a control room (a cell) with 20,000 unlabeled switches (genes). These switches all seem to operate on their own (via transcription factors and regulatory proteins), and there’s a constant clicking sound of switches turning off and on as you stand there. You notice fans turning off and on as the switches flick.
You could investigate the wires (DNA) connected to each switch, and with a clamp meter, measure the current output (RNA). You might start picking up on patterns, like how the current spikes in different wires at different times, but you’d still just be observing rather than directing anything yourself.
With the invention of CRISPR, you can finally grab one specific switch and flip it fully off, hold it partway down, or jam it to full-on (a knockout, CRISPRi, or CRISPR activation). Instead of passively measuring the currents, you can finally force something to happen.
It’d take a long time to map out what every single switch does with CRISPR, but if you layer Perturb-seq on top of CRISPR, it’s as if your boss hired a whole fleet of technicians. With Perturb-seq activated, instead of one person testing one switch in one room, technicians like you now go to thousands of rooms at once, randomly assigned to force one specific switch (pooling). Afterward, inspectors go room to room and record which switch got forced, and what every fan in that room is doing now.
Over thousands of iterations, you’ll finally start being able to map out which switch(es) influence which outcomes.
The Limitations of Observational Studies
Traditionally, studying gene expression has been mostly observational. It relies on collecting cells, measuring the transcriptional output of their genes (RNA), and observing patterns. This work is valuable, but it cannot determine causation since you can’t directly manipulate anything.
For example, let’s say you observe that Gene X71 and Gene Y918 are both active in cancer cells. You might believe that Gene X71 activates Gene Y918, but it’s just as likely that Gene Y918 is actually activating Gene X71, or that there’s a third gene (a confounding variable) you’re unaware of that’s activating both of them.
How Perturb-seq Works
The word “perturb” means to disturb or change something, which is an apt description of what Perturb-seq does. It disturbs a gene and then measures the outcome.
Let’s say you’re studying Gene X71 and Gene Y918 (the genes you noticed were active in cancer cells). To test their relationship using Perturb-seq, researchers could shut down Gene X71 and see what happens to Gene Y918. If Gene Y918 changes, it’s much stronger evidence that Gene X71 influences it.
Here’s how that experiment works in more detail:
All Perturb-seq experiments start with CRISPR, which uses a molecule called a single-guide RNA (sgRNA). Single-guide RNA is like a key that physically guides the CRISPR protein towards the target section of DNA. From there, scientists can:
Disable the gene completely (CRISPR knockout).
Lower the gene’s activity without cutting the DNA (CRISPR interference (CRISPRi)).
Increase the gene’s activity (CRISPR activation).
CRISPR lets you manipulate one gene at a time, but a human cell has roughly 20,000 genes, and testing them one by one would take ages. To speed up the process, researchers use a technique called pooling.
With pooling, scientists create thousands of different guide RNAs and mix them into a single library. That library gets delivered to a large population of cells, tuned so each cell picks up just one guide. Across the full population, though, each guide ends up represented in many different cells, so instead of testing one gene once, you’re testing thousands of genes at once, with each one tested many times over.
To manage all this data, Perturb-seq gives each guide RNA something like a barcode, so results can be sorted by cell, noting which gene was targeted and how much every gene in that cell was expressed.
Researchers then compare cells that received a real guide against control cells that got a blank guide, evaluating how RNA differs from that baseline. If cells that received a Gene X71 guide consistently show the same RNA changes, that pattern starts to map the actual effects of Gene X71.
The Value of Single-Cell Data
Before single-cell sequencing, there was bulk sequencing. Bulk sequencing is the process of studying thousands or millions of cells together and then averaging the results. It still has modern research applications, but it doesn’t account for variations across cells.
Cells, even those from the same tissue, are not identical. At any given moment, some are dividing while others are at rest, and some are responding to a signal while others are ignoring it entirely. A perturbation might hit one group of cells hard and do almost nothing to another group sitting right next to it. In a bulk experiment, those two effects merge into one middling number that describes neither population.
Like bulk sequencing, Perturb-seq still involves large numbers of cells (often millions across an experiment). But it measures each cell individually rather than averaging them, which makes it possible to explore whether the same gene has different roles in different cell types.
This part is important and bears repeating. Perturb-seq doesn’t only ask, “What does this gene do?” It asks, “What does this gene do in this cell type, under these conditions, at this moment?” That’s a much more precise question to work from.
The Rapid Growth of Perturb-seq
The first Perturb-seq experiments happened in 2016, developed by several groups working independently (Dixit et al., Cell, 2016; Adamson et al., Cell, 2016; Jaitin et al., Science, 2016; Datlinger et al., Nature Methods, 2017). Dixit and colleagues, for example, targeted about two dozen transcription factors, while Adamson and colleagues used roughly 30,000 cells to study ten perturbations. These studies were small by today’s standards, but they proved that pooled CRISPR methods could produce highly detailed, single-cell data.
Since then, the scale of studies using pooled CRISPR has exploded. By 2022, Replogle and colleagues published an entire genome-scale Perturb-seq study covering more than 2.5 million human cells (Cell, 2022). Using CRISPRi, they studied nearly 10,000 genes in K562 leukemia cells and repeated the experiment in a non-cancerous cell line to compare results.
Think about the sheer scale of this work for a moment. Every one of those 2.5 million cells contributes its own full transcriptome, meaning the dataset holds billions of individual measurements. And with that much data, researchers can group genes into shared pathways and infer what poorly understood genes might be doing. Using this approach, Replogle and colleagues identified a previously unrecognized subunit of a protein complex involved in RNA processing.
The scale of these experiments can feel abstract, but the logic behind a discovery like that is pretty straightforward when broken down into an example.
Say Gene P22 has been studied for years and is already known to play a role in cancer. Gene Q83, by contrast, is barely studied, and nobody’s sure what it does. But you suspect it might be connected to Gene P22, so you knock down each gene separately using CRISPR and compare the results. If the RNA changes look nearly identical in both cases, that’s a strong clue the two genes are doing similar work, perhaps as part of the same pathway, or even the same protein complex. By tracking how Gene Q83 behaves relative to Gene P22, you developed compelling evidence for what the mysterious Gene Q83 does, too.
This is one of the most valuable things about Perturb-seq. Genetics has long over-emphasized a small set of famous genes while leaving countless others understudied. Perturb-seq gives researchers a way into those overlooked genes, simply by tracing their associations with genes we already understand.
How Perturb-seq Innovations Address Its Limitations
While the discovery of Perturb-seq was a great scientific advancement, it has its own set of challenges.
Compressed Perturb-seq
A standard Perturb-seq experiment requires many cells per perturbation, which gets expensive quickly as you scale it.
Yao and colleagues designed a way around this. Normally, each cell gets exactly one guide RNA, so any change in that cell’s RNA can be traced back to one clear cause. Yao’s approach breaks that rule by letting a single cell receive multiple guides at once, or spreading guides across cells sharing the same droplet (Nature Biotechnology, 2023).
Since most genes only affect a small, specific set of other genes (meaning RNA changes from different guides rarely overlap completely), there is a strong enough pattern for software to work backwards and determine which guide most likely caused which change. It’s a bit like recording a band together and using software to pull out each instrument. Using this method, Yao and colleagues were able to cut the cost of a Perturb-seq experiment by an order of magnitude.
In Vivo Perturb-Seq
Standard Perturb-seq experiments take place in a petri dish, which is useful for reasons like cost and the ability to control the environment. But cells behave differently in a petri dish than they do inside living tissue, where they’re receiving signals from neighboring cells, immune cells, blood vessels, and more. In vivo Perturb-seq solves for this by running the experiment inside a living host.
In one of these studies, Jin and colleagues used CRISPR to disrupt 35 genes linked to autism and developmental delay inside the brain of a developing mouse before birth. Then, they sequenced the cells after the mouse was born (Science, 2020). They found effects in neurons and glia, the brain’s other major supporting cell type. Interestingly, several of these 35 genes produced similar RNA changes when disrupted, suggesting those genes may have a shared pathway.
More recent work has pushed the same idea further into brain development at a larger scale (Zheng et al., Cell, 2024).
Multiome Perturb-seq
Standard Perturb-seq only measures RNA, which means it only accounts for what is transcribed from DNA. But some genes never get transcribed because they sit in a region of DNA that stays tightly packed, in a DNA-protein structure called chromatin, and can’t be read by the cell’s machinery. Standard Perturb-seq has no way of accounting for this.
Multiome versions of Perturb-seq solve for this by measuring RNA and chromatin accessibility together. By reading both layers at once, in a process called Perturb-ATAC (Rubin et al., Cell, 2019), researchers can see how DNA accessibility changes, and how gene expression changes with it.
Spatial Perturb-seq
To run standard single-cell sequencing, researchers first have to dissociate the tissue, meaning they physically break it apart so individual cells can be pulled out and measured one at a time. That process tells researchers exactly what each cell was, but once cells are separated, they’ve lost their location relative to other cells.
Spatial approaches solve for this by keeping the tissue intact. Instead of separating cells to read them, researchers tag perturbed cells with a visible label and image the tissue as a whole to preserve location data.
In one study, Dhainaut, Rose, and colleagues used this method, called Perturb-map, in a mouse lung cancer model, allowing them to see how disrupting a single gene changed not just the tumor cell itself, but the immune cells around it (Cell, 2022).
Spatial Perturb-seq has enormous potential for studying how genes shape the relationships between different cell types in disease, such as how a single gene inside a tumor cell can change the immune cells surrounding it.
Does Perturb-seq Prove Causation?
No, Perturb-seq gives strong evidence of a real relationship between genes, since scientists are directly intervening rather than just observing. But it isn’t perfect. Several things can undermine any individual result, for example:
A guide RNA might fail to change its target gene, or it might accidentally hit the wrong stretch of DNA instead.
Single-cell sequencing only captures a fraction of the RNA in each cell, so missing data is probable (this is one reason the Replogle team filtered out results backed by too few cells before making conclusions).
A measurement taken early might capture a gene’s direct response, while a later measurement might instead reflect stress, adaptation, or cells dying off.
Perhaps the deepest limitation is that Perturb-seq can’t control for all potential indirect or mediated effects. If changing Gene A17 changes Gene D5, that still does not prove Gene A17 controls Gene D5 directly. Gene A17 might act through a mediator (Gene Z3), and Gene Z3 in turn changes Gene D5. Perturb-seq reads the endpoint of a chain without necessarily showing the links in between, so follow-up experiments are still required to validate the mechanism.
Millions of cells can make results feel more conclusive than they really are. Large-scale data is valuable, and it goes much further to establish a causal influence between variables than an observational study can, but even Perturb-seq cannot tell you whether an effect is direct.
Perturb-seq Offers A New Way Forward
With observational studies, you can observe patterns, asking questions like, “Are Gene P22 and Gene Q83 both active at the same time?” With CRISPR, you can start asking narrow causal questions, one at a time, such as “If I shut down Gene P22, does Gene Q83 change?”
While Perturb-seq isn’t perfect, it can change the kind of questions a scientist can ask. It allows you to combine CRISPR with pooling and single-cell readout to ask questions about a gene you know nothing about. For example, you could take the mysterious Gene Q83 and ask, “What does forcing Gene Q83 do across different cells, and does its behavior mimic that of any known genes?”
That’s what’s exciting about Perturb-seq. Rather than just testing a hypothesis you have with CRISPR, you’re generating hypotheses just by virtue of noticing patterns across perturbations. That’s a novel place for gene expression research to find itself, and it means researchers are starting to find things they didn’t even know to look for.
Sources
Adamson, B. et al. (2016). A multiplexed single-cell CRISPR screening platform enables systematic dissection of the unfolded protein response. Cell.
Datlinger, P. et al. (2017). Pooled CRISPR screening with single-cell transcriptome readout. Nature Methods.
Dhainaut, M., Rose, S. A. et al. (2022). Spatial CRISPR genomics identifies regulators of the tumor microenvironment. Cell, 185(7), 1223-1239. doi:10.1016/j.cell.2022.02.015
Dixit, A. et al. (2016). Perturb-Seq: Dissecting molecular circuits with scalable single-cell RNA profiling of pooled genetic screens. Cell.
Jaitin, D. A. et al. (2016). Dissecting immune circuits by linking CRISPR-pooled screens with single-cell RNA-seq. Science.
Jin, X. et al. (2020). In vivo Perturb-Seq reveals neuronal and glial abnormalities associated with autism risk genes. Science, 370(6520). doi:10.1126/science.aaz6063
Replogle, J. M. et al. (2022). Mapping information-rich genotype-phenotype landscapes with genome-scale Perturb-seq. Cell, 185, 2559-2575. doi:10.1016/j.cell.2022.05.013
Rubin, A. J. et al. (2019). Coupled single-cell CRISPR screening and epigenomic profiling reveals causal gene regulatory networks. Cell.
Yao, D. et al. (2023). Scalable genetic screening for regulatory circuits using compressed Perturb-seq. Nature Biotechnology. doi:10.1038/s41587-023-01964-9
Zheng, X. et al. (2024). Massively parallel in vivo Perturb-seq reveals cell-type-specific transcriptional networks in cortical development. Cell. doi:10.1016/j.cell.2024.04.050

