Interpretability / features
Features
Neurons are polysemantic; directions in activation space are the better unit. The methods here learn an overcomplete, sparse basis for those directions — and the entries on where that basis goes wrong are as important as the ones on how to fit it.
3 entries, 2 of them stubs.
Entries
05.01.1O(L·d·m)05.01.2—05.01.3O(d·m)
Crosscodersstubpromising
Shared dictionaries fitted across layers or across model checkpoints.
Feature Absorptionstubpromising
When one SAE latent swallows a more specific feature’s activation mass.
Sparse Autoencoderscommon
Overcomplete dictionary learning on activations for monosemantic features.