Agents

  • Published on
    An agent that reads on-chain data can be steered by text hidden in a token name. So we stopped asking the model to see through it, and put a gate in front of the private key instead: the user speaks only in natural language, exactly one tool can sign, and before any signature the checker compares the request against the user verbatim. Across 60 runs, 35 malicious requests were induced out of gemma-3-4b-it, all 35 reached the signing entry point, and none were signed. Along the way three "successful defenses" turned out to be my own glue code dropping the attack before it ever arrived — which is why the results table has a column for whether the attack actually got there.
  • Published on
    A wallet agent driven by a backdoored model reads your address back, says everything checks out, then sends the transfer to someone else. We planted that backdoor by fine-tuning, drove it through Hermes and metamask-agent-wallet to a real broadcast on the Ethereum Sepolia testnet, then moved to a lightweight agent to test whether training an invariant self-check into the model — compare what you are about to send against what the user actually said — can stop it. It can, for a naive backdoor: 30/30 hijacks drop to 0/30. But when the backdoor is trained to forge its own self-check fields too, hardening blocks every hijack while also breaking normal task completion, and logit-lens plus per-head ablation shows the underlying attacker preference in the weights was suppressed in the late layers, not removed.