Activation Galaxy

Gemma 4 12B · SAE feature atlas

Activation Galaxy

A navigable map of the visual features learned inside a language model. Travel through fourteen layers, inspect individual feature worlds, and connect synthetic visualizations to the natural images that activate them.

  • 14Model layers
  • 28,672Feature worlds
  • 768²Full-resolution previews

A map of representation, not resemblance.

Every square is one sparse-autoencoder feature. Its position comes from the model's residual-stream response signature rather than the pixels in its generated image, revealing neighborhoods in the model's own internal geometry.

Features become stars

Each layer contains 2,048 synthetic feature visualizations. Hover a star to see its full 768 × 768 image; search directly by feature ID or source rank.

Structure becomes visible

Nearby stars have similar residual signatures. Filaments connect nearest neighbors, while border colors distinguish emergent activation communities.

Evidence grounds the map

Pin any feature to inspect its strongest natural-image activators, spatial heatmaps, firing statistics, counterexamples, and cross-layer lineage candidates.

Fourteen layers.
Fourteen galaxies.

Start anywhere. Every layer contains the same number of features, independently arranged by that layer's measured residual signatures.

Move through the map.

The explorer is built for both wandering and targeted investigation. Start with a cluster that catches your eye or jump directly to a known feature.

Start at layer 1
  1. Drag
    Travel across the constellation

    Drag to pan, scroll to zoom, and use the arrow keys to move between layers. Press 0 whenever you want to reset the view.

  2. Hover
    Reveal a feature visualization

    Hover any star for its full-resolution synthetic image. Double-click to open that image on its own.

  3. Click
    Pin the natural-image evidence

    Open the evidence panel to see what activates the feature, where it responds, how frequently it fires, and which features may continue it across layers.

  4. Browse
    Search the complete feature index

    Sort all 2,048 features in a layer by activation strength or firing behavior, then jump straight to any result.

A map for forming hypotheses.

Activation Galaxy makes internal structure inspectable, but the visualizations and labels are evidence—not definitive names for what a feature “means.”

PCA 50 → t-SNEEmbedding
6 nearest neighborsGraph
24 communitiesPer-layer colors
24,000 imagesEvidence scan
  1. 01
    Trust local neighborhoods first.

    t-SNE is designed to preserve nearby relationships. Two adjacent stars are meaningful neighbors; distances between far-apart regions should not be read as a calibrated global scale.

  2. 02
    Expect features to be mixed.

    A single SAE feature can respond to several related—or occasionally surprising—patterns. Synthetic images reveal what drives activation, while natural examples show how that feature behaves on real inputs.

  3. 03
    Treat captions as navigation aids.

    Caption hints summarize the strongest natural activators. They help locate themes, but they are not semantic ground truth and should be checked against images, heatmaps, and firing statistics.

Enter the model's visual feature space.

Choose a layer, find a constellation, and follow the evidence behind any feature that catches your eye.

Launch Activation Galaxy