Prompt-injection

  • Published on
    An agent that reads on-chain data can be steered by text hidden in a token name. So we stopped asking the model to see through it, and put a gate in front of the private key instead: the user speaks only in natural language, exactly one tool can sign, and before any signature the checker compares the request against the user verbatim. Across 60 runs, 35 malicious requests were induced out of gemma-3-4b-it, all 35 reached the signing entry point, and none were signed. Along the way three "successful defenses" turned out to be my own glue code dropping the attack before it ever arrived — which is why the results table has a column for whether the attack actually got there.