Feature Absorption
A general latent and a specific one that implies it can trade mass in a way that lowers L0 without changing reconstruction — leaving the specific feature apparently silent on exactly the inputs it describes.
This entry is a stub. The maths has not been rederived by hand, or the implementation has not been run on real tensors, so it is recorded here as an open question rather than an answer.
Read what follows as a pointer to the sources, not as a settled account.
Suppose one latent tracks “starts with the letter E” and another tracks the token elephant. Whenever elephant fires, the letter feature is redundant — and the sparsity penalty rewards folding the letter direction into the elephant latent and leaving the general one silent. Reconstruction is unchanged, falls, and the letter feature now reads as absent on exactly the inputs it describes.
That makes it a failure of the interpretation rather than of the fit, which is what makes it hard: no reconstruction metric will flag it. This entry stays a stub until the measurement — a probe on the general feature, evaluated on inputs where a specific one fires — has been run here rather than read about.