- Published on
SFT is usually explained as "keep training on question-answer pairs, but only compute loss on the answer." That tells you nothing about where the training lands in the weights. So I planted a backdoor into GPT-2 small by hand and went looking for it — through per-head DLA, activation patching, an IOI control task, and weight restoration. The backdoor turned out to be carried by the MLPs, distributed across all twelve layers, with no critical head, layer, or neuron anywhere. Along the way I had to retract one conclusion and refute one of my own hypotheses.