The Non-Biologists Guide to CRISPR

CRISPR

If you've paid any attention to the news, you know that CRISPR is a gene-editing tool. In December 2023, the world got its first CRISPR-based therapy with Casgevy. Last year, researchers used base editing to save an treat an infant with severe CPS1 deficiency (a rare liver disorder that causes ammonia to accumulate in the blood). But what does all this really mean? How does it work?

CRISPR is the address; Cas9 is the machine

The phrase “CRISPR–Cas9” tends to collapse several jobs into one. CRISPR began as a microbial defense system: pieces of genetic material from previous invaders are stored and later expressed as RNAs that help recognize matching sequences. Cas9 is the protein effector. Give it a guide RNA and it becomes a programmable search complex.

CRISPR stands for Clustered Regularly Interspaced Short Palindromic Repeats. In it's early days before the hype, all that was known was that CRISPR was a bacterial adaptive immune system found in around 40% of bacteria. Experiments showed that the CRISPR system allow bacteria to recognize an infection with a virus and ultimately kill the virus. But it wasn't known how the molecules that participate in this destruction worked. But it was known that these CAS (CRISPR Associated Sequences) Proteins were the ones that did the work. Cas9 is one of those proteins, but at the time it was known as Csn1.

In nature it was seen that CRISPR systems have CrRNA, just little copies of the virus that the bacteria had seen before. And it was also seen that there was a tracrRNA, which is a little RNA that helps the Cas9 protein find the CrRNA.

Martin Jinek a scientist from Doudna Lab and Krzysztof Chylinski from Charpentier lab realized that together, the crRNA and tracrRNA could be fused into a single guide RNA. This was a huge breakthrough because it meant that Cas9 could be programmed to target any DNA sequence by simply changing the guide RNA sequence.

The example here is Streptococcus pyogenes Cas9, usually shortened to SpyCas9. It is the workhorse most people mean when they say “Cas9”: a 1,368-amino-acid protein whose two lobes wrap around RNA and DNA.

The protein is a moving checkpoint

SpyCas9 has a broad recognition lobe (REC) and a nuclease lobe (NUC). The guide RNA threads between them. This architecture matters because Cas9 is not simply a pair of scissors waiting in an open position; guide binding reshapes the protein into a surveillance complex, and target binding drives more conformational checks before catalysis.

The recognition lobe helps sense the RNA–DNA hybrid. The nuclease lobe contains the catalytic machinery, the bridge helix that helps communicate between lobes, and the PAM-interacting region that grants access to a potential target.

It reads a short motif first. For SpyCas9, the canonical protospacer-adjacent motif is 5′-NGG-3′: any base followed by two guanines. The PAM is not part of the guide-matching sequence. It sits immediately beside it, acting more like a license to inspect.

Two arginines in the PAM-interacting domain—R1333 and R1335—contact those guanines. If either contact is lost, SpyCas9 cannot reliably recognize the PAM and the rest of the targeting sequence never gets a proper audition.

From PAM to R-loop

Once the PAM is recognized, the nearby DNA duplex begins to open. The spacer portion of the guide tests the exposed target strand by base pairing. A good match allows the RNA–DNA hybrid to extend; the non-target DNA strand is displaced. This three-stranded arrangement is called an R-loop.

Matching is therefore kinetic and sequential, not a single yes/no comparison. PAM recognition starts the process, pairing propagates away from the PAM, and mismatches—especially close to the PAM—can stop or destabilize the transition toward a cleavage-ready state. That layered checking is useful, but not perfect, which is why guide design and off-target validation matter in real experiments.

One break, two active sites

SpyCas9 cleaves the two DNA strands with different nuclease domains.

The catalytic and PAM-reading residues preserved in the study

Molecular job SpyCas9 residues What they do
RuvC active site D10, E762, H983, D986 Cleaves the displaced, non-target DNA strand
HNH active site D839, H840, N863 Cleaves the guide-complementary target strand
PAM recognition R1333, R1335 Contacts the two guanines in the NGG PAM

HNH is positioned against the target strand; RuvC receives the displaced strand. Cleavage typically occurs about three base pairs upstream of the PAM, producing a mostly blunt double-strand break. A D10A mutation disables RuvC and an H840A mutation disables HNH, so either change turns SpyCas9 into a single-strand nickase. Combining them yields the binding-competent but catalytically inactive dCas9 used for gene regulation and molecular labeling.

Cutting DNA is not the same as editing it

Cas9’s direct product is a DNA break. The cell’s repair machinery creates the lasting edit.

  • End joining reconnects the broken ends and can introduce small insertions or deletions. In a coding region, those changes can disrupt a gene.
  • Template-directed repair can copy information from a supplied donor template into the break, enabling a planned sequence change under the right cellular conditions.
  • Newer CRISPR systems can avoid a full double-strand break altogether, but base editors and prime editors are redesigned machines built on the targeting logic explained above.

This distinction is important: the guide determines where the complex concentrates, the Cas9 variant determines what molecular action is possible, and the cell type and repair context strongly influence the final outcome.

A molecular walkthrough

A search engine with molecular scissors

CRISPR provides the address; Cas9 performs the search and the cut. In this walkthrough, we will assemble one editing complex piece by piece before following it from DNA recognition to repair.

01 · Meet

Meet SpyCas9

SpyCas9 is a 1,368-amino-acid protein built around a central channel. Its REC domains inspect nucleic acids; HNH and RuvC provide two cutting centers; the PI region reads the PAM.

Recognition lobe + nuclease lobe + PAM reader

02 · Load

Load the single-guide RNA

The orange sgRNA enters the central channel. Its folded scaffold is the handle Cas9 grips, while the exposed 20-nucleotide spacer carries the sequence that will interrogate DNA.

Scaffold = handle · spacer = programmable address

03 · Approach

Bring in the DNA duplex

Double-stranded DNA now passes the loaded complex. The target and non-target strands remain paired at first; Cas9 can sample many such segments without opening them completely.

Loaded Cas9 meets intact double-stranded DNA

04 · Search

Read the PAM first

The complex samples DNA, but it does not unzip every sequence it meets. SpyCas9 first looks for an NGG protospacer-adjacent motif. R1333 and R1335 make the decisive contacts with its two guanines.

NGG is permission to inspect the neighboring DNA

05 · Pair

Let the RNA interrogate the DNA

PAM recognition destabilizes the nearby duplex. The guide RNA can now pair with the complementary target strand, building an RNA–DNA hybrid while the other DNA strand is displaced.

A matching sequence grows into an R-loop

06 · Cut

Two nuclease domains, two cuts

The HNH domain cleaves the guide-complementary target strand. RuvC cleaves the displaced non-target strand. Together they usually leave a blunt double-strand break about three bases upstream of the PAM.

HNH → target strand · RuvC → non-target strand

07 · Repair

The cell writes the ending

Cas9 makes the break; the cell performs the edit. End joining can create small insertions or deletions, while template-directed repair can copy in a designed sequence when a repair template is available.

Break → disruption or template-guided change

What the animation leaves out

The walkthrough is intentionally schematic. Real Cas9 moves through many conformational states; DNA sampling is stochastic; magnesium ions participate in catalysis; guide sequence and chromatin context affect access; and an impressive predicted structure is not evidence of cleavage in a cell. Structural models narrow the experimental search. They do not replace biochemical assays.

Primary references

  • https://www.youtube.com/watch?v=cuHD7jCY8X4
  • https://www.youtube.com/watch?v=MZG1KmWMP2Q
  • Jinek et al., “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity,” Science (2012). DOI
  • Gasiunas et al., “Cas9–crRNA ribonucleoprotein complex mediates specific DNA cleavage for adaptive immunity in bacteria,” PNAS (2012). DOI
  • Nishimasu et al., “Crystal structure of Cas9 in complex with guide RNA and target DNA,” Cell (2014). DOI
  • Anders et al., “Structural basis of PAM-dependent target DNA recognition by the Cas9 endonuclease,” Nature (2014). DOI