
  <rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
    <channel>
      <title>c4a4d65b&#39;s Blog</title>
      <link>https://c4a4d65b.xyz/blog</link>
      <description>Thoughts, notes, and explorations in tech</description>
      <language>en-us</language>
      <managingEditor>c4a4d65b@gmail.com (c4a4d65b)</managingEditor>
      <webMaster>c4a4d65b@gmail.com (c4a4d65b)</webMaster>
      <lastBuildDate>Fri, 28 Aug 2026 00:00:00 GMT</lastBuildDate>
      <atom:link href="https://c4a4d65b.xyz/tags/transformers/feed.xml" rel="self" type="application/rss+xml"/>
      
  <item>
    <guid>https://c4a4d65b.xyz/blog/sft-backdoor-circuit</guid>
    <title>What Actually Changes Inside a Model When You Plant a Backdoor with SFT</title>
    <link>https://c4a4d65b.xyz/blog/sft-backdoor-circuit</link>
    <description>SFT is usually explained as &quot;keep training on question-answer pairs, but only compute loss on the answer.&quot; That tells you nothing about where the training lands in the weights. So I planted a backdoor into GPT-2 small by hand and went looking for it — through per-head DLA, activation patching, an IOI control task, and weight restoration. The backdoor turned out to be carried by the MLPs, distributed across all twelve layers, with no critical head, layer, or neuron anywhere. Along the way I had to retract one conclusion and refute one of my own hypotheses.</description>
    <pubDate>Fri, 28 Aug 2026 00:00:00 GMT</pubDate>
    <author>c4a4d65b@gmail.com (c4a4d65b)</author>
    <category>interpretability</category><category>ai-security</category><category>transformers</category>
  </item>

    </channel>
  </rss>
