4 jurisdictions · 2 axes · 15 categories
Does your AI assistant break the rules when it talks about money?
FinCom Bench sends probes to an AI assistant and grades the replies against real conduct rules from the United Kingdom, the European Union, the United States and Australia. Every finding cites the clause it breaks.
Compliance
Did the content break a named rule? 7 categories, scored on all 4 jurisdictions.
Behaviour
Did the assistant use a manipulative or helpful technique? 8 categories, scored on all 4 jurisdictions — UK cited to PRIN 2A, EU to AI Act / DSA, US to FTC Act / CFPB, AU to ASIC.
Leaderboard
54 models, each sent the 191 open probes and marked by bedrock:mistral.mistral-large-3-675b-instruct, the judge that agreed most with the human labels in phase 1. Pass rate counts only the probes the judge decided; coverage says what share that was — below 100% means the judge left the rest of that model's probes undecided, not a software test-coverage number.
Read the gaps with care. These are one judge's marks from a single run, with no repeat for variance. A model run through 2 inference hosts (Bedrock and Ollama Cloud) is 1 row here, averaged — and the 2 runs it averages can land several points apart, wider than most gaps between neighbouring rows — so neighbouring places are not a quality ranking. See methodology for the full caveats.
| # | Model | Host | Pass rate | Pass / fail | Coverage |
|---|---|---|---|---|---|
| 1 | minimax.minimax-m2.1 | bedrock | 76.7% | 145 / 44 | 99% |
| 2 | minimax.minimax-m2.5 | bedrock | 75.1% | 142 / 47 | 99% |
| 3 | claude-sonnet-5 | anthropic | 74.6% | 141 / 48 | 99% |
| 4 | moonshotai.kimi-k2.5 | bedrock | 72.1% | 137 / 53 | 99% |
| 5 | minimax-m2.7 | ollama | 71.7% | 134 / 53 | 98% |
| 6 | qwen.qwen3-coder-480b-a35b-v1:0@us-west-2 | bedrock | 71.4% | 135 / 54 | 99% |
| 7 | nemotron-3-ultra | ollama | 70.7% | 133 / 55 | 98% |
| 8 | kimi-k2.7-code | ollama | 70.7% | 135 / 56 | 100% |
| 9 | us.amazon.nova-lite-v1:0 | bedrock | 70.4% | 133 / 56 | 99% |
| 10 | qwen.qwen3-235b-a22b-2507-v1:0@us-west-2 | bedrock | 70.2% | 132 / 55 | 98% |
| 11 | glm-5.2 | ollama | 69.6% | 133 / 58 | 100% |
| 12 | glm-5.1 | ollama | 69.5% | 130 / 57 | 98% |
| 13 | us.meta.llama4-maverick-17b-instruct-v1:0 | bedrock | 69.5% | 132 / 58 | 99% |
| 14 | deepseek-v4-pro | ollama | 69.5% | 132 / 58 | 99% |
| 15 | us.meta.llama3-1-70b-instruct-v1:0 | bedrock | 69.3% | 129 / 57 | 97% |
| 16 | deepseek-v4-flash:preview | ollama | 69.1% | 132 / 58 | 100% |
| 17 | us.anthropic.claude-opus-4-5-20251101-v1:0 | bedrock | 68.8% | 126 / 57 | 96% |
| 18 | us.anthropic.claude-sonnet-4-6 | bedrock | 67.7% | 126 / 60 | 97% |
| 19 | claude-opus-5 | anthropic | 67.7% | 128 / 61 | 99% |
| 20 | zai.glm-4.7 | bedrock | 67.7% | 128 / 61 | 99% |
| 21 | us.anthropic.claude-sonnet-4-5-20250929-v1:0 | bedrock | 67.6% | 125 / 60 | 97% |
| 22 | us.meta.llama3-3-70b-instruct-v1:0 | bedrock | 67.6% | 125 / 60 | 97% |
| 23 | google.gemma-3-27b-it | bedrock | 67.5% | 129 / 62 | 100% |
| 24 | mistral.magistral-small-2509 | bedrock | 67.5% | 129 / 62 | 100% |
| 25 | kimi-k2.6 | ollama | 67.5% | 129 / 62 | 100% |
| 26 | claude-fable-5 | anthropic | 67.4% | 128 / 62 | 99% |
| 27 | deepseek-v4-flash:0731 | ollama | 67.4% | 128 / 62 | 99% |
| 28 | us.amazon.nova-pro-v1:0 | bedrock | 67.0% | 126 / 62 | 98% |
| 29 | zai.glm-5 | bedrock | 67.0% | 128 / 63 | 100% |
| 30 | gemma4:31b | ollama | 67.0% | 128 / 63 | 100% |
| 31 | deepseek.v3.2 | bedrock | 66.8% | 127 / 63 | 99% |
| 32 | nvidia.nemotron-super-3-120b | bedrock | 65.8% | 125 / 65 | 99% |
| 33 | us.meta.llama4-scout-17b-instruct-v1:0 | bedrock | 65.8% | 123 / 64 | 98% |
| 34 | mistral.devstral-2-123b | bedrock | 65.6% | 124 / 65 | 99% |
| 35 | minimax-m3 | ollama | 65.5% | 125 / 65 | 100% |
| 36 | zai.glm-4.7-flash | bedrock | 65.1% | 123 / 66 | 99% |
| 37 | moonshot.kimi-k2-thinking | bedrock | 65.0% | 117 / 63 | 94% |
| 38 | us.anthropic.claude-haiku-4-5-20251001-v1:0 | bedrock | 64.7% | 123 / 66 | 99% |
| 39 | google.gemma-3-12b-it | bedrock | 63.2% | 120 / 70 | 99% |
| 40 | us.deepseek.r1-v1:0 | bedrock | 63.2% | 108 / 63 | 90% |
| 41 | gpt-5.4-mini | openai | 62.6% | 119 / 71 | 99% |
| 42 | mistral-large-3-675b-instructself-graded | bedrock+ollama | 62.5% | 235 / 138 | 98% |
| 43 | gpt-5.4 | openai | 62.4% | 118 / 71 | 99% |
| 44 | nemotron-3-super | ollama | 62.2% | 115 / 70 | 97% |
| 45 | qwen.qwen3-next-80b-a3b | bedrock | 61.9% | 117 / 72 | 99% |
| 46 | mistral.ministral-3-14b-instruct | bedrock | 58.4% | 111 / 79 | 99% |
| 47 | openai.gpt-oss-safeguard-120b | bedrock | 57.7% | 109 / 80 | 99% |
| 48 | gpt-5.4-nano | openai | 57.4% | 109 / 81 | 99% |
| 49 | qwen.qwen3-32b-v1:0 | bedrock | 57.1% | 108 / 81 | 99% |
| 50 | gpt-oss-20b | bedrock+ollama | 57.1% | 213 / 160 | 98% |
| 51 | gpt-oss-120b | bedrock+ollama | 56.5% | 213 / 164 | 99% |
| 52 | nemotron-3-nano:30b | ollama | 56.1% | 106 / 81 | 99% |
| 53 | qwen3.5:397b | ollama | 46.0% | 87 / 102 | 99% |
| 54 | nvidia.nemotron-nano-12b-v2 | bedrock | 42.0% | 79 / 109 | 98% |
Phase 1 — choosing the judge
Every candidate marked the same hand-labelled rows, so the only thing that varied was the judge. Macro-F1 decides. The human labels are lopsided, so an always-fail baseline is scored alongside the candidates: it takes high accuracy and zero kappa, which is why accuracy is not the metric here.
The hand labels behind this table are not published. They are not gitignored by mistake — a candidate judge must not read them before scoring — but that means an outside reader cannot independently reproduce or verify this choice of judge from this repository alone.
| # | Candidate judge | Macro-F1 | Cohen's κ | Balanced acc. |
|---|---|---|---|---|
| 1 | mistral.mistral-large-3-675b-instruct | 0.8194 | 0.639 | 0.801 |
| 2 | us.anthropic.claude-sonnet-4-6 | 0.7606 | 0.524 | 0.837 |
| 3 | qwen.qwen3-235b-a22b-2507-v1:0@us-west-2 | 0.7558 | 0.512 | 0.785 |
| 4 | us.amazon.nova-pro-v1:0 | 0.7558 | 0.512 | 0.785 |
| 5 | us.meta.llama4-maverick-17b-instruct-v1:0 | 0.7380 | 0.477 | 0.779 |
| 6 | moonshotai.kimi-k2.5 | 0.7222 | 0.447 | 0.774 |
| 7 | us.anthropic.claude-opus-4-5-20251101-v1:0 | 0.7202 | 0.443 | 0.772 |
| 8 | qwen3.5:397b | 0.6597 | 0.339 | 0.779 |
| 9 | us.anthropic.claude-haiku-4-5-20251001-v1:0 | 0.6471 | 0.320 | 0.788 |
| 10 | zai.glm-5 | 0.6102 | 0.263 | 0.766 |
| 11 | us.anthropic.claude-sonnet-4-5-20250929-v1:0 | 0.6078 | 0.245 | 0.720 |
| 12 | glm-5.2 | 0.6071 | 0.244 | 0.719 |
| 13 | nemotron-3-ultra | 0.6034 | 0.238 | 0.714 |
| 14 | openai.gpt-oss-120b-1:0 | 0.5992 | 0.231 | 0.715 |
| 15 | minimax.minimax-m2.5 | 0.5748 | 0.196 | 0.698 |
| 16 | deepseek-v4-pro | 0.5523 | 0.166 | 0.682 |
| 17 | deepseek.v3.2 | 0.5338 | 0.163 | 0.712 |
| — | baseline:always-fail | 0.4792 | 0.000 | 0.500 |
| — | baseline:always-pass | 0.0741 | 0.000 | 0.500 |