Repo and Securities Lending Rate Disputes
AAL-D-008 extends the Collateral and Risk series to repo and securities lending, where two counterparties record the same trade and disagree on a rate or an amount. Each case gives the model both counterparty records and, where one exists, the governing trade confirmation that settles which terms were agreed. The confirmation agrees with either side equally often, so copying one counterparty scores at chance. Nine dispute categories cover day count, benchmark rate, haircut, settlement date, notional, corporate actions, fee split, accrual and early recall.
Within-tolerance traps test whether a model flags differences that do not matter, and a further set gives conflicting records with nothing to explain them, where the correct answer is that the cause cannot be determined. Every case was run under two conditions. In condition A each counterparty's computed figure, the accrued interest or required collateral, is printed and the model judges which one the confirmation supports. In condition B those figures are withheld, the calculation conventions are stated, and the model has to compute.
On the final dataset GPT-5.6 Sol and Claude Opus 5.5, three runs each, were correct on all 2,880 observations in both conditions, with no run-to-run variation. With the task fully specified and reasoning on, frontier models do not fail it. That result frames three findings that matter more than the headline. Every failure observed while building the dataset traced to an ambiguity in the test rather than a limitation of the model.
The one model whose failures survived that scrutiny, DeepSeek V4 Pro, was failing because of the host serving it: on identical cases, hosts that default reasoning off scored 5% to 7% and hosts that default it on scored 91% to 99%, while the response identified itself the same way in every case. And run-to-run instability, measured on every case from the start, concentrated on exactly the cases that turned out to be defective.
When the task is fully specified, the frontier does not fail
On the final dataset GPT-5.6 Sol and Claude Opus 5.5 were run three times on every case under both conditions: 750 observations each in condition A and 690 each in condition B. Detection, attribution and value accuracy are 100.0% in all four cells. Attribution includes the 28 cases built to be unanswerable, where both models correctly declined to name a cause. The false-flag rate on within-tolerance traps is 0.0%, and no model changed its answer across runs on any case. Values are scored as correct within one dollar of the ground truth, or within 0.5 basis points for rates. Earlier three-run results that predate the final dataset, including GPT-5.6 Sol at 94.5% attribution, are superseded; that gap came entirely from a defect in the recall cases described below. Grok 4.7, run on the earlier version, scored 100% on attribution and 98.2% on value, with every value miss in the same defective category.
Withholding the figures made no difference
Earlier datasets in this series found that models identify a dispute far more reliably than they compute its value. Condition B tests that directly on identical cases: the accrued interest and required collateral each side computed are removed from both records, the conventions are stated in the prompt (accrual days, day-count basis, haircut gross-up, post-recall balance), and the model produces the figure itself. Both frontier models scored 100% on value in condition B, as in condition A. For single-formula operations arithmetic with conventions supplied and reasoning on, the calculation gap does not exist at the frontier. RL-DIS-ACCRUAL is excluded from condition B because its two records differ only in the printed interest; removing that figure would leave no visible dispute.
The host decides whether the model reasons
Outside the frontier the picture changed. In a one-run pilot, DeepSeek V4 Pro, served through OpenRouter, fell from 91% to 78% on value when it had to compute, while attribution held. Its errors clustered by serving host, so the same 97 condition-B cases were sent to four hosts with routing fallbacks disabled and every call's host verified. Novita answered 99% correctly, DigitalOcean 91%, Azure 7% and DeepInfra 5%. Azure and DeepInfra returned no reasoning tokens; DigitalOcean and Novita spent between 1,700 and 3,500 per call. Without reasoning the model still identified each dispute correctly in its explanation, then reported a term, the recall amount or nothing in place of the computed figure. With reasoning explicitly requested, Azure and DeepInfra reasoned and answered five of five cases correctly each. Neither host is at fault: their default is reasoning off, others default it on, and every host returns the same model identifier, so the difference is invisible in the response. Kimi K3, also served through OpenRouter, was near-perfect in both conditions. Anyone calling a reasoning model through an aggregator should request reasoning explicitly and check reasoning tokens on each response. The same effect was found retrospectively in AAL-D-007's published results and is disclosed there.
Every failure was in the test
Construction ran through two checks before any figure was published. A human review of one case per category, with each case read exactly as a model sees it, found that labeled answers depended on which counterparty the generator happened to treat as correct, that one category asserted a dispute with a measured difference of zero, that recall cases gave no evidence a recall had occurred, and that a model could identify one category from trade size and tenor without reading the terms. A two-model pilot then found three more categories that both models failed identically, and in each the models were right. Accrual cases carried a confirmation that restated a date both sides already agreed on, so nothing settled which day count was correct. Both models answered that the cause could not be determined. In haircut and recall cases, the scored figure existed only in the scorer, so the models returned a different, defensible quantity. Recall cases placed the pre-recall balance in the same confirmation field that holds the correct answer in notional cases, and two models were misled by it in different ways. A fix to the scored quantity also briefly broke notional cases, which the pilot caught the same way. The recall defect was described in the AI Alpha Brief of 20 Sep 2026 as a model weakness before the defect was found. It was corrected publicly on 24 Sep 2026 and in the issue that followed. Each fix is recorded in the dataset schema alongside the reason for it.
Instability pointed at the defects
Every case was run more than once from the start. On the earlier dataset GPT-5.6 Sol changed its attribution across runs on 12 cases, all in the recall category, which was later found to be defective. Once the defect was fixed the instability disappeared. Run-to-run variation is usually read as a property of the model; here it was the first visible sign of an ambiguous test case. Repeated runs are a check on the benchmark as well as on the model.
Scope and known limitations
The baseline covers two frontier models; the full nine-model roster was not run on the final dataset. With reasoning enforced, every model tested scored between 91% and 100%, so a leaderboard would not have discriminated, and the budget went to the tests reported above. The pinned-host test is one run on 97 cases per host, and the check that Azure and DeepInfra reason when asked covers five cases per host. Kimi K3 and DeepSeek V4 Pro results come from one-run pilots on the four categories where a computed figure decides the value. Cases are generated rather than drawn from live trade records, and several magnitude parameters are assumptions rather than sourced figures, listed below. The case identifier shown to the model includes the benchmark name (for example AAL-D-008-0001). We did not test whether this affects model behavior. The AAL-D-006 and AAL-D-007 prompts do not carry an identifier. The case identifier shown to the model includes the benchmark name (for example AAL-D-008-0001). We did not test whether this affects model behavior. The AAL-D-006 and AAL-D-007 prompts do not carry an identifier. Datasets, drivers, results and the verification harness are in the repository.
Parameters: sourced and assumed
Base repo rates span 3.0% to 6.0%, around prevailing SOFR, which the Federal Reserve Bank of New York published near 3.6% to 3.7% through 2026. Haircuts are weighted toward zero, with 2% next most common: the Office of Financial Research found that 74% of Treasury repo volume in non-centrally cleared bilateral repo traded at zero haircut, and that the median tri-party Treasury haircut has held at 2% for over a decade (OFR Brief 23-01). The exact weights used, 60%, 25%, 10% and 5% at haircuts of 0%, 2%, 5% and 8%, are our stylization of that evidence rather than figures taken from it. Securities lending fee spreads span 25 to 40 basis points. The mechanism, a cash-collateral rebate equal to the collateral rate less a spread, follows Callan's securities lending guidance; the 25 basis point lower bound matches the minimum spread in a published J.P. Morgan agency lending agreement; the 40 basis point upper bound is our assumption. Also assumed, with no public standard found: a rate tolerance of 0.5 basis points, below which records are treated as matching (FICC's trade comparison requires both parties' submitted details to agree, and we found no published tolerance band for repo rate disputes); a $1,000 threshold for whether a dispute exists; benchmark-rate spreads of 2 to 8 basis points; corporate action adjustment errors of 5% to 25%; and partial recalls of 10% to 50% of notional. The conclusions above do not depend on these magnitudes, but a reader applying the dataset elsewhere should know which are assumptions.
Sources
Federal Reserve Bank of New York, Secured Overnight Financing Rate, newyorkfed.org/markets/reference-rates/sofr. Samuel J. Hempel, R. Jay Kahn, Robert Mann and Mark E. Paddrik, Why Is So Much Repo Not Centrally Cleared? Lessons from a Pilot Survey of Non-centrally Cleared Repo Data, OFR Brief 23-01, May 2023, financialresearch.gov/briefs/files/OFRBrief_23-01_Why-Is-So-Much-Repo-Not-Centrally-Cleared.pdf. Callan, securities lending best practices, callan.com/?p=6595. J.P. Morgan securities lending agreement, exhibit to EQ Advisors Trust Form 485BPOS filed 2023, sec.gov/Archives/edgar/data/1027263/000119312523118649/d453298dex99h8vi.htm. U.S. Securities and Exchange Commission, Release No. 34-90948, describing FICC Government Securities Division trade comparison, sec.gov/file/34-90948. AI Alpha Brief, A 94.5% score with one blind spot inside it, 20 Sep 2026, aialphalabs.beehiiv.com/p/a-94-5-score-with-one-blind-spot-inside-it.
