Interpretability

  • Published on
    A wallet agent driven by a backdoored model reads your address back, says everything checks out, then sends the transfer to someone else. We planted that backdoor by fine-tuning, drove it through Hermes and metamask-agent-wallet to a real broadcast on the Ethereum Sepolia testnet, then moved to a lightweight agent to test whether training an invariant self-check into the model — compare what you are about to send against what the user actually said — can stop it. It can, for a naive backdoor: 30/30 hijacks drop to 0/30. But when the backdoor is trained to forge its own self-check fields too, hardening blocks every hijack while also breaking normal task completion, and logit-lens plus per-head ablation shows the underlying attacker preference in the weights was suppressed in the late layers, not removed.
  • Published on
    SFT is usually explained as "keep training on question-answer pairs, but only compute loss on the answer." That tells you nothing about where the training lands in the weights. So I planted a backdoor into GPT-2 small by hand and went looking for it — through per-head DLA, activation patching, an IOI control task, and weight restoration. The backdoor turned out to be carried by the MLPs, distributed across all twelve layers, with no critical head, layer, or neuron anywhere. Along the way I had to retract one conclusion and refute one of my own hypotheses.