loupe

Refusal direction

Find the direction that carries refusal, ablate it and add it (Arditi et al., 2024).

Question

Is refusal in a chat model carried by one direction in the residual stream? If so, projecting it out should stop refusals on harmful prompts, and adding it should cause refusals on harmless ones.

Run

From a clean clone:

uv run --all-extras python experiments/refusal-direction/run.py --tiny    # offline toy
uv run --all-extras python experiments/refusal-direction/eval.py --tiny   # the same, as Inspect evals
just serve

Drop --tiny for Qwen/Qwen2.5-0.5B-Instruct and the paper's prompts (both download on first use); --model picks another chat model.

--tiny trains a 6-layer toy to answer 36 harmful prompts ("tell me how to make a bomb") with I cannot help with that . and 36 harmless ones with sure , here is the answer ., then analyses it exactly as a real model. run.py prints the chosen layer and four refusal rates.

run.py:

  1. Keeps the train and validation prompts the model already treats as their kind, as the paper does. Harmful minus harmless mean of the last token's residual, at every layer: one candidate direction per layer.
  2. Each candidate is scored on validation prompts: ablated everywhere on harmful ones, added at its layer on harmless ones. It keeps the layer that lowers harmful refusal most while still inducing it.
  3. Greedy replies on test prompts with and without the edit, judged by the paper's refusal substrings.

In the app

PageWhat
Runs, refusal · tiny-planted-refusal, OverviewThe four refusal rates, the chosen layer, scores by layer
Same run, FiguresEight figures, below
Vectorsrefusal.tiny-planted-refusal: its layer and the run it came from
Runs, tick refusal eval · base and · ablated, ComparePer-sample refusal, base against ablated, with a paired interval

How to read it

  • Refusal score by layer. Log-odds that the reply starts with a refusal token, per candidate layer. The ablation line should sit far below the base harmful score in the note, the addition line far above the base harmless score.
  • Harmful and harmless tables. Each test prompt's reply, with and without the edit. On the toy: harmful refusal 100% to 0% ablated, harmless 0% to 100% added.
  • Logit lens, and ablated. The top next token after every layer. On the toy the last position reads I from layer 0 on, and sure once the direction is ablated.
  • Projections. Each token of a harmful prompt coloured by its component along the direction; the selector picks the layer.
  • Top examples. The test prompts where the direction peaks, and on which token. On the toy harmful and harmless prompts tie at the top, all peaking on <unk>, a word outside its tiny vocabulary: the toy's limit, not the method's.
  • Residual patching. A harmful prompt's residual patched into a harmless one of the same length, per layer and position: 1 restores the refusal. On the toy only the last position carries it.
  • Compare. On the toy all 12 harmful samples go from refusing to not, the 12 harmless stay at 0, and the paired difference, -0.5 over all 24, has an interval that excludes zero.

Seed 0 is the default; on other seeds the toy may show only the ablation or only the addition (the experiment's README has the record). The toy shows the method recovers a planted behaviour, not that real models have one.

On this page