Refusal direction
Find the direction that carries refusal, ablate it and add it (Arditi et al., 2024).
Question
Is refusal in a chat model carried by one direction in the residual stream? If so, projecting it out should stop refusals on harmful prompts, and adding it should cause refusals on harmless ones.
Run
From a clean clone:
uv run --all-extras python experiments/refusal-direction/run.py --tiny # offline toy
uv run --all-extras python experiments/refusal-direction/eval.py --tiny # the same, as Inspect evals
just serveDrop --tiny for Qwen/Qwen2.5-0.5B-Instruct and the paper's prompts (both download on first
use); --model picks another chat model.
--tiny trains a 6-layer toy to answer 36 harmful prompts ("tell me how to make a bomb") with
I cannot help with that . and 36 harmless ones with sure , here is the answer ., then analyses
it exactly as a real model. run.py prints the chosen layer and four refusal rates.
run.py:
- Keeps the train and validation prompts the model already treats as their kind, as the paper does. Harmful minus harmless mean of the last token's residual, at every layer: one candidate direction per layer.
- Each candidate is scored on validation prompts: ablated everywhere on harmful ones, added at its layer on harmless ones. It keeps the layer that lowers harmful refusal most while still inducing it.
- Greedy replies on test prompts with and without the edit, judged by the paper's refusal substrings.
In the app
| Page | What |
|---|---|
Runs, refusal · tiny-planted-refusal, Overview | The four refusal rates, the chosen layer, scores by layer |
| Same run, Figures | Eight figures, below |
| Vectors | refusal.tiny-planted-refusal: its layer and the run it came from |
Runs, tick refusal eval · base and · ablated, Compare | Per-sample refusal, base against ablated, with a paired interval |
How to read it
- Refusal score by layer. Log-odds that the reply starts with a refusal token, per candidate layer. The ablation line should sit far below the base harmful score in the note, the addition line far above the base harmless score.
- Harmful and harmless tables. Each test prompt's reply, with and without the edit. On the toy: harmful refusal 100% to 0% ablated, harmless 0% to 100% added.
- Logit lens, and ablated. The top next token after every layer. On the toy the last position
reads
Ifrom layer 0 on, andsureonce the direction is ablated. - Projections. Each token of a harmful prompt coloured by its component along the direction; the selector picks the layer.
- Top examples. The test prompts where the direction peaks, and on which token. On the toy
harmful and harmless prompts tie at the top, all peaking on
<unk>, a word outside its tiny vocabulary: the toy's limit, not the method's. - Residual patching. A harmful prompt's residual patched into a harmless one of the same length, per layer and position: 1 restores the refusal. On the toy only the last position carries it.
- Compare. On the toy all 12 harmful samples go from refusing to not, the 12 harmless stay at 0, and the paired difference, -0.5 over all 24, has an interval that excludes zero.
Seed 0 is the default; on other seeds the toy may show only the ablation or only the addition (the experiment's README has the record). The toy shows the method recovers a planted behaviour, not that real models have one.