LLM explainability
Experiments and notes on figuring out what language models are actually doing under the hood.
-
LLM Content Overview
Why this section exists, and the explainers and resources that inspired it.
-
The Illustrated Anatomy of a Model
An interactive explainer for what parameter, neuron, activation, concept, and token actually mean — five words people use interchangeably, made concrete on a real toy network.
-
The Superposition Problem
An interactive explainer for why capturing every activation in a neural network doesn't let you read its thoughts.