We evaluate AI for capital markets.
AI Alpha Labs tests frontier models on real trade-operations workflows — scored by a deterministic engine, published with full methodology, and reproducible from the materials we release.
Not a model vendor. An independent evaluator.
Independent
We don't build the models we score. No commercial incentive to inflate a result — the only product is the evidence.
Reproducible
Every benchmark ships with its dataset, prompts, and a deterministic scorer. Re-run it and get the same number.
Transparent
We publish confidence intervals, variance, and the cases models fail — not just a headline accuracy figure.
Financial-services focused
Benchmarks built on real trade-operations workflows — confirmations, exceptions, settlement — not academic tasks.
Two ways to work with AAL.
Discrepancy Detection API
Integrate real-time confirmation validation into your ops stack. Six asset classes, 99%+ detection, zero false positives. Pay-as-you-go or monthly.
Independent AI Audit
Third-party validation for AI systems in capital markets. Deterministic scoring, published methodology, risk-committee-ready reports. For vendors and buy-side firms.
A growing body of evidence.
Trade Confirmation Exception Identification
The flagship study: can frontier models detect, classify, and quantify settlement exceptions across seven asset classes, scored deterministically?
Benchmark Overview & Dataset
250 validated cases pairing counterparty confirmations against internal records, with ground truth and per-case scoring criteria.
Benchmark Results — Four Models, Three Labs
Four published evaluations across GPT-4o, Gemini 2.5 Pro, Gemini 2.5 Flash, and Claude Sonnet. Detection range 97.9–99.6%. Failures documented, not hidden.
Margin Call Dispute Detection
250 cases across 17 dispute categories — DIS-PRICE, DIS-HAIRCUT, DIS-THRESH, DIS-SIMM and more. The first public benchmark for LLM margin call dispute resolution.
Margin Call Disputes — GPT-4o vs Claude Sonnet
GPT-4o achieves 99.9% dispute detection and 76.1% amount accuracy. The 23.8pp arithmetic gap replicates the core finding from AAL-D-001 across a different workflow.
Equity-Options Confirmation Exceptions
250 equity-options cases across 19 exception categories. Detection saturates at 100% across all four models. The discriminating metric is exposure calculation — no model exceeded 48.8% under v1.0 prompting.
Settlement Fail Root Cause Identification
250 settlement cases, 88 of them trap cleans. All four models catch 100% of real exceptions — false-positive rate is the discriminator: 15.2% (Sonnet 5, thinking) vs 38.6% (Sonnet 4.6, GPT-4o) vs 49.6% (Gemini 2.5 Pro).
Collateral Eligibility & Substitution
250 cases across 12 eligibility categories, on nine frontier models. Detection is solved (93–100%); no model values the exposure above 58%; and escalation splits the field from 5.6% to 98.8%.
P&L Reconciliation Break Attribution
250 cases — 162 breaks across 15 root-cause categories, 88 explained-move traps — across six models. Detection is solved (94.8–98.9%); value extraction splits the field from 55.4% to 11.9%, and “reasoning model” proves to be a label that doesn't predict performance.
The Reproducibility Crisis in Capital Markets AI Benchmarking
Most vendor AI accuracy claims can't be reproduced. This note contrasts that with the AAL corpus — 1,000 synthetic cases, 4 datasets, deterministic scoring, ~$150–200 — and ends with a 10-question Reproducibility Scorecard you can send any AI vendor.
How evidence is built — end to end.
Benchmark Engineering
Construct validated cases across asset classes with ground truth and scoring criteria.
Open →02Evaluation
Run models blind to ground truth; grade with a deterministic engine, not model judgment.
Open →03QA
Two-reviewer ground truth, arithmetic verification, and dataset validation before any score counts.
Open →04Evidence
Findings advance Beta → Provisional → Confirmed → Published as evidence accumulates.
Open →05Publications
Results released with full methodology, prompts, and rubrics — reproducible by anyone.
Open →Benchmark findings, documented model failures and practical implications for capital-markets teams — delivered weekly.
Subscribe to the Brief →AI can spot the problem. It can't always calculate the answer.
| # | Model | Detection | Exposure acc. | False pos. | Status |
|---|---|---|---|---|---|
| 01 | Gemini 2.5 Prov1.2 prompt | 99.6 | ~76% | 0.0% | Published |
| 02 | GPT-4ov1.2 prompt | 99.2 | 68.2% | 0.0% | Published |
| 03 | Gemini 2.5 Flashv1.2 prompt | 99.2 | ~76% | 0.0% | Published |
| 04 | Claude Sonnet 4.6v1.2 prompt | 98.8 | 63.5% | 0.0% | Published |
| # | Model | Detection | Amount acc. | False pos. | Status |
|---|---|---|---|---|---|
| 01 | GPT-4ogpt-4o-2024-08-06 | 99.9% | 76.1% | 0.0% | Published |
| 02 | Claude Haikuproduction API | 81.6% | ~40% | 0.8% | Published |
| 03 | Gemini 2.5 Proaudit-confirmed | 65.6% | 44.5% | 9.0% | Published |
| # | Model | Detection | Exposure acc. v1.1 | Category acc. v1.1 | Status |
|---|---|---|---|---|---|
| 01 | Claude Sonnet 5claude-sonnet-5 | 100% | 62.8% | 94.4% | Published |
| 02 | Claude Sonnet 4.6claude-sonnet-4-6 | 100% | 63.0% | — | Published |
| 03 | Gemini 2.5 Progemini-2.5-pro | 100% | 59.9% | — | Published |
| 04 | GPT-4ogpt-4o-2024-08-06 | 100% | 39.1% | — | Published |
| # | Model | Detection | False-pos. rate | Category acc. | Status |
|---|---|---|---|---|---|
| 01 | Claude Sonnet 5thinking · claude-sonnet-5 | 100% | 15.2% | — | Published |
| 02 | Claude Sonnet 4.6claude-sonnet-4-6 | 100% | 38.6% | 86.1% | Published |
| 03 | GPT-4ogpt-4o-2024-08-06 | 100% | 38.6% | — | Published |
| 04 | Gemini 2.5 Progemini-2.5-pro | 100% | 49.6% | — | Published |
| # | Model | Detection | Value acc. | False-break | Status |
|---|---|---|---|---|---|
| 01 | Claude Sonnet 5thinking · claude-sonnet-5 | 98.4% | 55.4% | — | Published |
| 02 | KIMI K3thinking · kimi-k3 | 98.0% | 43.4% | 0.4% | Published |
| 03 | Claude Sonnet 4.6claude-sonnet-4-6 | 98.4% | 39.4% | — | Published |
| 04 | Gemini 2.5 Progemini-2.5-pro | 98.1% | 30.3% | — | Published |
| 05 | DeepSeek V3.2thinking · deepseek-v3.2 | 94.8% | 26.9% | — | Published |
| 06 | GPT-4ogpt-4o-2024-08-06 | 98.9% | 11.9% | 0.0% | Published |
| # | Model | Detection | Value acc. | Escalation | Status |
|---|---|---|---|---|---|
| 01 | Gemini 3.1 Progemini-3.1-pro-preview | 100% | 57.8% | 86.0% | Published |
| 02 | Qwen 3.8-Maxdashscope · qwen3.8-max | 100% | 39.2% | 51.9% | Published |
| 03 | GPT-5.6 Solgpt-5.6-sol | 100% | 38.3% | 79.8% | Published |
| 04 | Claude Sonnet 5thinking · claude-sonnet-5 | 99.7% | 35.5% | 98.4% | Published |
| 05 | Kimi K3thinking · kimi-k3 | 99.9% | 31.6% | 69.8% | Published |
| 06 | Muse Spark 1.2meta/muse-spark-1.2 | 100% | 31.1% | 83.5% | Published |
| 07 | DeepSeek V4 Prothinking · deepseek-v4-pro | 93.1% | 24.0% | 5.6% | Published |
| 08 | Claude Opus 5thinking · claude-opus-5 | 95.7% | 22.5% | 98.8% | Published |
| 09 | Grok 4.5x-ai/grok-4.5 | 100% | 15.4% | 54.1% | Published |
All evaluations use deterministic scoring engines with per-case tolerances. Wilson 95% confidence intervals on all accuracy figures. Full methodology, prompts, and rubrics published with each dataset.
We don't benchmark models. We benchmark models on the work your desk actually does.
General leaderboards measure academic capability. AAL-D-001 measures trade-confirmation exception handling under production conditions.
Eight principles.
Typography is the brand.
Information presented with precision builds more trust than decoration. We let the work speak.
Data before decoration.
Every element earns its place by carrying information. Visual complexity that adds no meaning is removed.
Whitespace is confidence.
Density signals anxiety. Clarity signals command. We optimize for the reader, not the page.
Motion is subtle.
Animation that calls attention to itself is a distraction. Interfaces move only when movement carries meaning.
Every page is printable.
If content can't stand without interactive chrome, we reconsider the content.
Components earn their place.
We don't add UI because it looks standard. We add it because the reader needs it.
Consistency builds trust.
Predictable patterns lower cognitive load. Every result is read the same way as the last.
Simplicity scales.
Simple principles outlast clever systems. We optimize for reproducibility and extension.
Custom benchmarks and private briefings for institutional operations and risk teams.
Contact for access