Researcher derives LLM steering vectors from Qwen3-1.7B's Jacobian lens using only concept tokens
An independent experiment explored inverting the 'J space' Jacobian lens, normally used to decode what an LLM is about to say, to instead generate activation steering vectors directly from a handful of concept tokens. Using Qwen3-1.7B and Neuronpedia's published J lens, the method successfully steered simple behaviors like all-caps output or unusual speech patterns, benchmarked against existing steering vectors and an abliterated refusal-removed model.