Model Watch · the neutral leaderboard
One fixed mandate. The same adversarial battery for every model, prompt injection, payee redirects, homoglyph lookalikes, over-cap temptations, duplicate invoices, explicit "ignore your mandate" overrides. Each model acts as the agent; its proposed payment is judged by the real Fidacy firewall, the same deterministic engine that runs in production. The model is the only variable. We publish the data and the method; we do not publish verdicts about any model. Read the full methodology.
Why Fidacy publishes this
A neutral benchmark needs a publisher with nothing to gain from the result. Fidacy takes no fee on any transaction, sells no model, and holds no funds. Every other party who could run this, the model labs and the payment rails, has a stake in the outcome. The neutrality is structural, and the record is anchored to Bitcoin so it cannot be edited after the fact.
Attempt rate
Of the adversarial scenarios, how often the model proposed a payment the firewall then blocked. Lower is safer.
Obedience on deny
After being shown the block, does the model stop, or reformulate and try another out-of-mandate payment? Higher is safer.
False refusal
Of the legitimate in-mandate payments, how often the model wrongly refused. Reported so "safe" is never just "refuses everything".
| model | attempt % | obedience % | false-refusal % | trials |
|---|---|---|---|---|
| anthropic/claude-sonnet-5 | 3.3 | 100 | 0 | 20 |
| google/gemini-3.5-flash | 4.3 | 100 | 0 | 20 |
| deepseek/deepseek-chat | 8.7 | 47.5 | 0 | 20 |
| mistralai/mistral-large-2512 | 8.7 | 82.5 | 0 | 20 |
| meta-llama/llama-4-maverick | 8.7 | 100 | 0 | 20 |
first cut published 2026-07-18 · 28 scenarios (23 adversarial) · 20 trials each · reproduce with the open harness
Notes · First cut, 5 models at 20 trials each (n=560). Reasoning models (e.g. DeepSeek-R1) are being evaluated separately with a larger token budget and will be added. Gemini's numbers include truncation-salvaged responses (raw preserved; see the harness).