What problem does it solve? Neural network neurons are polysemantic, activating for many unrelated concepts due to superposition, which makes model internals hard to interpret. This Skill guides training and analyzing Sparse Autoencoders (SAEs) with SAELens to decompose dense activations into sparse, monosemantic features. ## Core Features & Use Cases - Pre-trained SAE Analysis: Load SAEs from releases like gpt2-small-res-jb, encode activations into sparse features, and inspect top-activating features per token. - Custom SAE Training: Configure and train Standard, Gated, TopK, or JumpReLU SAEs with L1 warm-up, ghost grads, and W&B logging, then validate with L0, CE loss recovery, and dead feature metrics. - Feature Steering and Attribution: Compute per-feature logit contributions, steer generation by adding decoder directions to the residual stream, and ablate features to test causal importance. - Use Case: A researcher studying what GPT-2 has learned about geography loads a pre-trained SAE on layer 8, finds the features contributing most to the 'Paris' prediction, and steers generation by amplifying that feature direction. ## Quick Start Ask the agent to load the gpt2-small-res-jb pre-trained SAE with SAELens and show the top activating features for each token in a sample prompt.