Home | Projects | Blog | Personal

Blog

Paper notes. Informal, sometimes wrong, mostly for me. If any of it is useful to you, great.

Recent readings

Read Aug 25, 2026 · Tang et al., Nature Methods, 2009

mRNA-Seq whole-transcriptome analysis of a single cell

  • Before this, you needed thousands of cells to read gene expression, so you only ever saw an average across a crowd. They figured out how to read the whole transcriptome from one mouse embryo cell.
  • The trick they used was to turn the tiny RNA into cDNA, then use PCR to copy it millions of times. They added a poly-A tail to every fragment so one universal primer could kick off amplification for all of them.
  • They found way more genes than microarrays could catch, because microarrays only see what they were built to look for, while sequencing reads whatever is actually there.
  • Real world: this is the origin of single-cell RNA-seq. It's how we now know tumors are mixed cell types, map immune cells, and trace embryo development step by step.
Read Aug 24, 2026 · Sanger, Nicklen & Coulson, PNAS, 1977

DNA sequencing with chain-terminating inhibitors

  • They figured out how to read DNA letter by letter. The trick is a broken nucleotide (ddNTP) that stops the copying enzyme cold, so you get fragments ending at every position of one letter.
  • Run four tubes, one for each letter, then line them up on a gel and read the sequence from bottom to top.
  • Real world: this was the method behind basically every genome sequenced for the next 30 years, including the Human Genome Project.
Read Aug 23, 2026 · Watson & Crick, Nature, 1953

Molecular Structure of Nucleic Acids: A Structure for Deoxyribose Nucleic Acid

  • One page. They propose DNA is two chains twisted around each other in a double helix, with the bases paired in the middle: A with T, G with C.
  • Because each base only pairs with its partner, one strand tells you exactly what the other strand says. They drop a one line hint that this suggests how DNA copies itself.
  • Real world: this is the shape everything else in genetics is built on.
Read Aug 22, 2026 · Jinek et al., Science, 2012

A programmable dual RNA-guided DNA endonuclease in adaptive bacterial immunity

  • The CRISPR paper. Bacteria use Cas9 plus a guide RNA to cut up viral DNA.
  • They showed you can swap in your own guide RNA and Cas9 will cut wherever you point it. Also merged the two natural RNAs into one single guide.
  • Real world: this is where gene editing as a tool comes from.
Read Aug 21, 2026 · Hanahan & Weinberg, Cell, 2011

Hallmarks of Cancer: The Next Generation

  • They boiled all of cancer down to 6 core tricks a cell picks up.
  • 1. Grow on its own. Normal cells only divide when they get a signal. Cancer cells figure out how to make that signal themselves or ignore the stop signs.
  • 2. Ignore the brakes. The body has built-in off switches that stop cells from dividing when they shouldn't. Cancer disables them.
  • 3. Refuse to die. Damaged cells are supposed to self-destruct. Cancer flips that switch off so damaged cells just keep living.
  • 4. Keep dividing forever. Normal cells hit a limit and stop. Cancer turns on a workaround so they can divide basically forever.
  • 5. Grow its own blood supply. Tumors need food and oxygen. They send out signals that make new blood vessels grow right to them.
  • 6. Spread. Cancer cells learn to break loose, travel, and start growing somewhere else. That's what makes it deadly.
  • Real world: this framework is how basically everyone now thinks about cancer. If you can find which trick a tumor is using, you can try to block that specific one.
Read Aug 20, 2026 · Beadle & Tatum, PNAS, 1941

Genetic Control of Biochemical Reactions in Neurospora

  • The one gene, one enzyme paper. They zapped bread mold (Neurospora) with X-rays to cause mutations, then showed that a single mutated gene knocked out a single specific enzyme, breaking one specific step in the mold's metabolism.
  • First solid experimental link between a gene and a specific biochemical function. Genes work by encoding enzymes.
  • Real world: this is the foundation of molecular biology. Every gene-enzyme-disease connection traces back here.
Read Aug 19, 2026 · Baek et al., Nature Methods, 2024

Accurate prediction of protein-nucleic acid complexes using RoseTTAFoldNA

  • RoseTTAFold was a protein structure predictor that used a three-track architecture: the 1D sequence, the 2D distance map, and the 3D shape all talk to each other at the same time and update each other. RoseTTAFoldNA extends that same idea to DNA and RNA, so now you can predict how a protein binds to DNA or RNA, not just other proteins.
  • One trained network handles both protein-DNA and protein-RNA complexes. It also gives you a confidence estimate per region so you know which parts to trust.
  • This is the same gap AlphaFold3 later filled (modeling proteins touching DNA and RNA), but RoseTTAFoldNA got there first and is fast.
  • Real world: could help design proteins that bind specific DNA or RNA sequences, which is useful for gene editing and for understanding how cells read their own genome.
Read Aug 18, 2026 · Hershey & Chase, Journal of General Physiology, 1952

Independent Functions of Viral Protein and Nucleic Acid in Growth of Bacteriophage

  • For today's update I wanted to read a paper that is a little more foundational. I wanted to take a step back.
  • The Hershey-Chase experiment was foundational for the discovery that DNA was the genetic material. Since DNA contains phosphorus and not sulfur, and proteins contain sulfur and not phosphorus, they were able to label phosphorus for DNA and sulfur for proteins.
  • They had two batches of phages (one with radioactive sulfur and the other with radioactive phosphorus), let them infect the bacteria, and literally shook it in a kitchen blender. They found that all the sulfur (proteins) stayed outside the bacteria and all the phosphorus stayed inside, showing that DNA was the thing being injected. Set the stage for further evidence in DNA being the genetic material.
  • Real world: set the stage for the whole field of molecular biology. No DNA as the genetic material = no double helix, no genetic code, no genome sequencing, no CRISPR.
Read Aug 17, 2026 · Watson et al., Nature, 2023

De novo design of protein structure and function with RFdiffusion

  • AlphaFold: existing protein → shape. RFdiffusion runs it in reverse: you give it the shape/job you want → it designs a brand new protein for it. Doesn't exist in nature.
  • Same diffusion trick as AF3, but instead of placing atoms of a known molecule it's building a protein backbone from scratch.
  • Key part: when they actually built these designs in the lab, a lot of them worked. That's the step most AI ideas die at.
  • Real world: custom proteins for vaccines, enzymes, therapies that evolution never got around to making.
Read Aug 16, 2026 · Abramson et al., Nature, 2024

Accurate structure prediction of biomolecular interactions with AlphaFold3

  • AF2 only did proteins alone. In real cells proteins are always touching something else (DNA, RNA, drug molecules, ions). AF3 does the whole complex together = much closer to reality.
  • Final 3D step swapped for a diffusion model. Starts from random atom positions, nudges them into place over many steps.
  • Still shaky on antibodies. Current fix = run it a bunch of times, pick the best.
  • Real world: drug companies can model how a candidate binds its target on a computer before burning months in the lab.
Read Aug 15, 2026 · Jumper et al., Nature, 2021

Highly accurate protein structure prediction with AlphaFold

  • Protein = chain of amino acids that folds into a 3D shape. The shape decides what it does. Reading the chain has been easy forever. Figuring out the shape used to take years of lab work per protein.
  • AlphaFold = neural net that predicts the shape in minutes from the chain. Also spits out a confidence score per region so you know which parts to trust.
  • Real world: shape is step 1 for designing basically any drug. Turned a years-long bottleneck into minutes.
Read Aug 14, 2026 · Schaffer, Hu et al. (Ideker Lab), Nature, 2025

Multimodal cell maps as a foundation for structural and functional genomics

  • Proteins do most of the work in a cell. To understand the cell you need to know where each protein sits + what it interacts with. This paper builds that map for human cells.
  • Two methods combined: affinity purification (pull a protein out, see what's bound to it → interaction partners) + immunofluorescence imaging (tag proteins, image under microscope → subcellular location). Each alone misses stuff, but they get a much sharper picture when used together.
  • Found hundreds of protein communities. Used an LLM to sort through the data and name them. A lot were completely new.
  • Real world: diseases like cancer often = one of these communities malfunctioning. Knowing which one is broken means drugs can target just that instead of the whole cell.
Read Aug 13, 2026 · Puniya et al., Bioinformatics Advances, 2024

Perspectives on computational modeling of biological systems and the significance of the SysMod community

  • This is an overview of how people mathematically model biological systems.
  • There are different frameworks for different questions: reaction rates → ODEs. gene on/off → Boolean networks. populations of cells → agent based.
  • There's not a single method that covers everything. Each subfield mostly works in isolation. SysMod exists to get them talking + combining models.
  • Real world: hybrid models could simulate a whole cell or tissue. Useful for testing drugs in silico before touching a lab.
← Back home