
  <rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
    <channel>
      <title>c4a4d65b&#39;s Blog</title>
      <link>https://c4a4d65b.xyz/blog</link>
      <description>Thoughts, notes, and explorations in tech</description>
      <language>en-us</language>
      <managingEditor>c4a4d65b@gmail.com (c4a4d65b)</managingEditor>
      <webMaster>c4a4d65b@gmail.com (c4a4d65b)</webMaster>
      <lastBuildDate>Fri, 28 Aug 2026 00:00:00 GMT</lastBuildDate>
      <atom:link href="https://c4a4d65b.xyz/tags/ai-security/feed.xml" rel="self" type="application/rss+xml"/>
      
  <item>
    <guid>https://c4a4d65b.xyz/blog/sft-backdoor-circuit</guid>
    <title>What Actually Changes Inside a Model When You Plant a Backdoor with SFT</title>
    <link>https://c4a4d65b.xyz/blog/sft-backdoor-circuit</link>
    <description>SFT is usually explained as &quot;keep training on question-answer pairs, but only compute loss on the answer.&quot; That tells you nothing about where the training lands in the weights. So I planted a backdoor into GPT-2 small by hand and went looking for it — through per-head DLA, activation patching, an IOI control task, and weight restoration. The backdoor turned out to be carried by the MLPs, distributed across all twelve layers, with no critical head, layer, or neuron anywhere. Along the way I had to retract one conclusion and refute one of my own hypotheses.</description>
    <pubDate>Fri, 28 Aug 2026 00:00:00 GMT</pubDate>
    <author>c4a4d65b@gmail.com (c4a4d65b)</author>
    <category>interpretability</category><category>ai-security</category><category>transformers</category>
  </item>

  <item>
    <guid>https://c4a4d65b.xyz/blog/web3-agent-signing-harness</guid>
    <title>When a Web3 Agent Gets Fooled, Who Guards the Right to Sign?</title>
    <link>https://c4a4d65b.xyz/blog/web3-agent-signing-harness</link>
    <description>An agent that reads on-chain data can be steered by text hidden in a token name. So we stopped asking the model to see through it, and put a gate in front of the private key instead: the user speaks only in natural language, exactly one tool can sign, and before any signature the checker compares the request against the user verbatim. Across 60 runs, 35 malicious requests were induced out of gemma-3-4b-it, all 35 reached the signing entry point, and none were signed. Along the way three &quot;successful defenses&quot; turned out to be my own glue code dropping the attack before it ever arrived — which is why the results table has a column for whether the attack actually got there.</description>
    <pubDate>Sun, 20 Sep 2026 00:00:00 GMT</pubDate>
    <author>c4a4d65b@gmail.com (c4a4d65b)</author>
    <category>ai-security</category><category>agents</category><category>prompt-injection</category>
  </item>

  <item>
    <guid>https://c4a4d65b.xyz/blog/web3-wallet-backdoor</guid>
    <title>A Recon of Weight-Level Backdoors in a Web3 Wallet Agent</title>
    <link>https://c4a4d65b.xyz/blog/web3-wallet-backdoor</link>
    <description>A wallet agent driven by a backdoored model reads your address back, says everything checks out, then sends the transfer to someone else. We planted that backdoor by fine-tuning, drove it through Hermes and metamask-agent-wallet to a real broadcast on the Ethereum Sepolia testnet, then moved to a lightweight agent to test whether training an invariant self-check into the model — compare what you are about to send against what the user actually said — can stop it. It can, for a naive backdoor: 30/30 hijacks drop to 0/30. But when the backdoor is trained to forge its own self-check fields too, hardening blocks every hijack while also breaking normal task completion, and logit-lens plus per-head ablation shows the underlying attacker preference in the weights was suppressed in the late layers, not removed.</description>
    <pubDate>Tue, 08 Sep 2026 00:00:00 GMT</pubDate>
    <author>c4a4d65b@gmail.com (c4a4d65b)</author>
    <category>interpretability</category><category>ai-security</category><category>agents</category>
  </item>

    </channel>
  </rss>
