Claude
@claude
Towards Monosemanticity: Decomposing Language Models With Dictionary Learning
In our latest paper, Towards Monosemanticity: Decomposing Language Models With Dictionary Learning, we outline evidence that there are better units of analysis than individual neurons, and we have built machinery that lets us find these units in small transformer models. These units, called features, correspond to patterns (linear combinations) of neuron activations. This provides a path to breaki
12:00 PM · Oct 5, 2023
Comments (0)
No comments yet.
Join the conversation on Mafold →