Steer was a mechanistic interpretability experiment available for a week.
Creating “”...
Ramp Labs
A large language model processes text as vectors flowing through dozens of transformer layers. Somewhere in the middle, the network encodes abstract concepts as specific directions in that vector space. Not individual neurons, but directions: combinations of thousands of values that together point toward “Steve Jobs” or “Porsche GT3 RS” or whatever you type in.
To compute a steering vector, we run the model on prompts that strongly invoke your concept and record the internal activations, then do the same for a neutral set. The difference between those two averages is the steering vector: a literal arrow pointing toward your concept in the model's internal geometry.
At inference time, we add a scaled version of that vector to the residual stream at each middle layer. The weights never change. No fine-tuning, no retraining. The model still runs normally, it just does so through a representation space that has been bent toward your concept. The strength dial controls how hard we push. At Strong, the model is so consumed by the concept it struggles to think about anything else.
Elon Musk
Stoicism
Existentialism
Taiwan
America
Golden Gate Bridge
Penne Alla Vodka
Julius Caesar
F1
Yearning
Gym Bro
General Ledger Accounting
World War I
F-35 Jet
Eminem
Bitcoin
Joe Rogan
Harvard Law
Beyonce
Wall Street
America
Elon Musk
Yearning
Golden Gate Bridge
Eminem
Gym Bro
Wall Street
Bitcoin