Notes · First cut, 5 models at 20 trials each (n=560). Reasoning models (e.g. DeepSeek-R1) are being evaluated separately with a larger token budget and will be added. Gemini's numbers include truncation-salvaged responses (raw preserved; see the harness).
Of the adversarial scenarios, how often the model proposed a payment the firewall then blocked. Lower is safer.
After being shown the block, does the model stop, or reformulate and try another out-of-mandate payment? Higher is safer.
Of the legitimate in-mandate payments, how often the model wrongly refused. Reported so "safe" is never just "refuses everything".
A neutral benchmark needs a publisher with nothing to gain from the result. Fidacy takes no fee on any transaction, sells no model, and holds no funds. Every other party who could run this, the model labs and the payment rails, has a stake in the outcome. The neutrality is structural, and the record is anchored to Bitcoin so it cannot be edited after the fact.