# 06. Grouping Similar Cells into Clusters

> Calculate expression similarity among cells in a quality-controlled cell-by-gene matrix and find groups of cells that are densely connected to one another.

Clustering is not drawing boundaries around islands on a two-dimensional plot. It adjusts measurement-depth differences among cells, connects cells with similar expression patterns, and finds densely connected groups.

> **Question for this lesson**: How are cells with tens of thousands of gene values compared to find recurring expression patterns?

## Make rows comparable

Cells differ in total **unique molecular identifier counts** (UMI counts), the numbers of original RNA molecules detected after deduplication. If cell A has 10,000 and cell B has 2,000, untransformed raw measurement depth can dominate their distance.

A common exploratory transformation scales total counts per cell and then applies `log1p`, which replaces each value $x$ with $\log(1+x)$.

```python
counts = adata.X.copy()       # preserve raw counts for later tests
normalize_total(adata)        # adjust per-cell count depth
log1p(adata)                  # reduce the influence of large values
```

This is conceptual pseudocode modelled on AnnData, the data object used by the Python package Scanpy. Normalized values support visualization and distance calculations. They do not reconstruct absolute RNA molecules. Raw counts and transformed values should be kept in separate storage areas called layers.

## Select informative gene columns

Many genes are nearly constant or rarely detected. Including all of them adds noise and computation.

**Highly variable genes, or HVGs**, vary more across cells than expected at their average expression. This is feature selection for the clustering representation. An HVG is not automatically a disease-related gene, and the HVG list is not the list of genes tested for a condition effect.

## Find neighbours after reducing the number of gene dimensions

**Principal component analysis** (PCA) compresses thousands of selected gene columns into tens of coordinates before the data are placed in two dimensions. This reduces correlated dimensions and weak noise that can dominate distances.

[PCA](/en/reference/pca/) explains scores and loadings.

Connecting every cell to its `k` nearest neighbours in PCA space produces a **k-nearest-neighbour graph**. Cells are now graph nodes, with edges linking similar expression profiles.



## Partition the neighbour graph and display it in two dimensions

Leiden clustering finds graph regions with dense internal connections. A higher `resolution` tends to split the graph more finely, while a lower value tends to merge larger regions.

**Uniform Manifold Approximation and Projection** (UMAP) places high-dimensional neighbourhoods in two dimensions for viewing. Cluster IDs can be coloured on the UMAP, but they do not come directly from its x and y coordinates. [Distance and dimensionality reduction](/en/reference/distance-and-dimension-reduction/) compares what PCA and UMAP preserve.

- Inspect whether nearby cells are connected consistently.
- Do not assign physical meaning to global orientation, empty space, or absolute UMAP distance.
- Test whether major structure survives changes in seed, neighbours, and preprocessing.
- Do not call every island a cell type automatically.

## Experimental differences can look like biological differences

A **batch effect** is a technical difference caused by experiment date, instrument, or processing group. Cells from different experiments, donors, or dates may separate by batch. Harmony and Seurat integration can reduce technical separation for clustering and visualization.

If condition and batch overlap completely, no algorithm can learn which difference is technical. Preserve sample identity and raw counts. Integrated coordinates may guide clustering, but condition-level expression tests still need real donors and the appropriate count layer.

## Keep each derived output

| stage | representative output |
| --- | --- |
| normalization | normalized or log-transformed layer |
| feature selection | HVG list and parameters |
| PCA | cell by PC coordinates and gene loadings |
| neighbour calculation | cell-cell graph |
| Leiden | cluster ID per cell |
| UMAP | two-dimensional viewing coordinates |

## Summary

Clustering selects HVGs from a normalized matrix, finds neighbours in PCA space, and uses Leiden to identify densely connected graph regions. UMAP is a view of those relationships, not an automatic cell-type labeler.

Next, [Naming Cell Types with Gene Patterns](/en/lessons/single-cell-annotation/) develops evidence for biological labels on clusters 0, 1, and 2.

---

### Sources

- Scanpy preprocessing and clustering: [official workflow](https://scanpy.readthedocs.io/en/stable/tutorials/basics/clustering.html)
- Seurat PBMC workflow: [guided clustering tutorial](https://satijalab.org/seurat/articles/pbmc3k_tutorial)
- Single-cell analysis best practices: [Heumos et al., *Nature Reviews Genetics* (2023)](https://www.nature.com/articles/s41576-023-00586-w)
- Harmony integration: [Korsunsky et al., *Nature Methods* (2019)](https://www.nature.com/articles/s41592-019-0619-0)