Skip to leaderboard
Dissei Leaderboard / AI evaluation on financial evidence

Financial reasoning,
measured.

Graded results for AI models reasoning from financial evidence across seven categories, with estimated cost per attempt. Case coverage differs by model, so results are descriptive, not a matched ranking.

What the Dissei Leaderboard is

The Dissei Leaderboard reports how AI language models perform when they must reason from financial evidence. Each model works through financial cases and has to reach a conclusion that can be checked. A grader assesses every attempt against the requirements of its case, so a result reflects the reasoning behind an answer, not only whether a final figure matches. Read how we think about financial judgment.

Two measures lead the table. Score is a continuous reward on a 0–100 scale; it is not accuracy. Gate pass is the share of attempts that cleared the required checks. Results are then broken down across seven kinds of financial reasoning, from quantitative work to predictive judgment, and set beside estimated inference cost, token use and tool calls.

This page is an evaluation record, not a training environment and not a buying guide. Models were not run on a matched set of cases, so coverage differs and the table describes pooled results rather than a head-to-head ranking. Each release is published at this one address with its date, model coverage and source. The worked examples show what a graded attempt looks like.

Observed results

Results

Pooled evaluation scores by model. Case coverage differs between models, so these results are not a matched head-to-head comparison.

2026-08-31 · 5 models
Model results as of 2026-08-31. Coverage varies by model. Scores and gate pass rates are higher-is-better; costs are estimates.
GPT-5.6 Sol
45.53
80.6%$0.285
Claude Opus 4.8
44.26
81.4%$0.823
Muse Spark 1.2
42.41
76.4%$0.284
Kimi K3
39.76
78.6%$0.196
DeepSeek v4-flash
25.49
67.1%$0.027

Costs are policy-model estimates based on historical source rate and price tables, not current prices. They exclude grading and infrastructure. Per-model caveats are printed separately. Historical results pool different case coverage.

  • GPT-5.6 Sol: Historical rate-card estimate in the source export; cache mix unverified.
  • Claude Opus 4.8: Historical price-table estimate from the source export; model label corrected using archived run records.
  • Muse Spark 1.2: Historical estimate from the source export.
  • Kimi K3: Historical price-table estimate in the source export.
  • DeepSeek v4-flash: Historical relay estimate in the source export.

Sorted by Score, descending.

Reasoning profile

Where performance changes

Compare two models

Historical continuous reward per 100. Not accuracy.

Select a category to see its description and question.

Each cell shows score / 100.
Score out of 100 by model and reasoning category.
Model
GPT-5.6 Sol
Claude Opus 4.8
Muse Spark 1.2
Kimi K3
DeepSeek v4-flash
Quantitative

Calculate, reconcile, and interpret financial quantities.

Illustrative question

How much headroom remains under a changed earnings assumption?

Seven categories

The seven kinds of financial reasoning

The results report score and gate pass for each of these seven categories; the profile above shows every model's result in each one.

Quantitative reasoning

Calculating, reconciling and interpreting financial quantities: checking that figures tie together and understanding what a number implies, not only producing it.

Illustrative question: How much headroom remains under a changed earnings assumption?

Diagnostic reasoning

Working back from a financial outcome to its causes, and separating the drivers the evidence supports from those it does not.

Illustrative question: What explains the gap between reported earnings and cash generation?

Comparative reasoning

Weighing alternatives on one consistent basis, so that differences in terms, timing or measurement do not distort the comparison.

Illustrative question: Which financing proposal offers stronger downside protection?

Strategic reasoning

Turning evidence into a course of action that can be defended, with the trade-offs made explicit.

Illustrative question: Which capital allocation decision is supported by the available evidence?

Counterfactual reasoning

Changing one assumption and following the consequences through the rest of the analysis, instead of adjusting a single figure in isolation.

Illustrative question: How would a delayed refinancing change the liquidity outlook?

Explanatory reasoning

Explaining how a financial mechanism works, such as how a contractual term shifts each party's incentives, using the evidence in the case.

Illustrative question: How does a contractual term change the incentives of each party?

Predictive reasoning

Judging which outcomes are plausible using only the information available at the decision date, without the benefit of hindsight.

Illustrative question: Which downside scenarios are supported at the decision date?

What this release shows

Findings from the 2026-08-31 results

Computed directly from the historical aggregate snapshot (historical-2026-08-31) shown above. They describe these 5 models on unequal case coverage and do not rank models in general.

  • Pooled scores span 20.04 points.

    Pooled scores run from 25.49 (DeepSeek v4-flash) to 45.53 (GPT-5.6 Sol) on the 0–100 scale. Because coverage differs, a gap in pooled score is not on its own evidence that one model reasons better than another.

  • Strategic reasoning is where scores are lowest.

    As an unweighted average of the model results, strategic reasoning scores 33.5, and predictive reasoning scores 50.2. Strategic is the lowest category for 3 of 5 models. See the reasoning profile for every cell.

  • Score and gate pass pick out different models at the top.

    GPT-5.6 Sol has the highest pooled score (45.53), while Claude Opus 4.8 has the highest gate pass (81.4%). The two measures answer different questions: how much of the reward an attempt earned, and whether it cleared the required checks.

  • Cost is reported beside score.

    Estimated policy-model cost runs from $0.027 (DeepSeek v4-flash) to $0.823 (Claude Opus 4.8) per attempt. These are historical estimates, not current prices; the efficiency section plots them against score.

For a controlled comparison on the same tasks, the experiment in Where equal models come apart shows how two models with nearly equal averages part on hindsight, known analyst traps and the effort each spends per episode. That article reports its own experiment and scope, which differ from the pooled coverage of this table. How cases and graders are built is covered in The rubric locks first and Your reward function is lying to you.

Efficiency

Score is one part of the picture

Cost, tokens, and tool use from the same source snapshot.

Cost and score

Observed scores, with differing case coverage.

Pareto chart: minimise cost, maximise observed score. Logarithmic cost axis.Observed score (/ 100)01020304050550.010.020.050.10.20.51Cost (USD / attempt) · log scaleGPT-5.6 Sol: 0.285 USD / attempt, 45.53 / 100; descriptive frontierSolClaude Opus 4.8: 0.823 USD / attempt, 44.26 / 100Opus 4.8Muse Spark 1.2: 0.284 USD / attempt, 42.41 / 100; descriptive frontierMuse 1.2Kimi K3: 0.196 USD / attempt, 39.76 / 100; descriptive frontierKimi K3DeepSeek v4-flash: 0.027 USD / attempt, 25.49 / 100; descriptive frontierv4-flash
Observed frontierOther models
GPT-5.6 Sol
Cost
0.285 USD / attempt
Observed score
45.53 / 100
How to read the frontier

Models have different case coverage. This frontier describes the displayed aggregates; it is not a matched benchmark ranking.

Filled points have no displayed alternative with a lower or equal cost and a higher or equal score, with at least one strict improvement. The pale line connects those observations in cost order; it does not estimate results between them.

Log scale. Spreads out cost differences across a wide range.

Tokens per attempt

Mean token usage and tool calls.

Mean input tokens per recorded attempt. Not a cost figure.

  • GPT-5.6 Sol45,272Tool calls: 13.8
  • Claude Opus 4.8147,362Tool calls: 12.2
  • Muse Spark 1.2212,235Tool calls: 16.1
  • Kimi K357,415Tool calls: 8.0
  • DeepSeek v4-flash400,399Tool calls: 29.0
Side by side

Compare two models

Explore the differences across seven kinds of reasoning.

Summary for GPT-5.6 Sol and Claude Opus 4.8
MetricSolOpus 4.8
Score / 10045.5344.26
Gate pass80.6%81.4%
Policy cost ≈$0.285$0.823
Mean turns4.346.94

Coverage differs between models. These are descriptive differences across each model’s available cases.

  • QuantitativeGPT-5.6 Sol: 38.9Claude Opus 4.8: 47.3
  • DiagnosticGPT-5.6 Sol: 44.8Claude Opus 4.8: 38.9
  • ComparativeGPT-5.6 Sol: 59.3Claude Opus 4.8: 38.1
  • StrategicGPT-5.6 Sol: 37.7Claude Opus 4.8: 36.9
  • CounterfactualGPT-5.6 Sol: 45.7Claude Opus 4.8: 43.4
  • ExplanatoryGPT-5.6 Sol: 52.6Claude Opus 4.8: 51.0
  • PredictiveGPT-5.6 Sol: 53.9Claude Opus 4.8: 61.5

Sol Opus 4.8 · Scores use the same 0–100 scale. A missing score leaves a gap.

Worked examples

How a financial answer is graded

Two credit cases, step by step: the decisive credit question, what a sound answer must establish and, where an attempt was recorded, how the grader assessed it.

Two separately sourced worked examples illustrate a case-level assessment and a recorded attempt. They are not evidence rows from the leaderboard. Names are removed, but figures may still identify the case.

Example 01 / Financial judgment · Contingent liabilities

When the cash is not the company’s

Decision window March to May 2020

The question

Whose cash is on the balance sheet, and how much of the refund obligation falls on the platform rather than on event creators?

Situation

A listed ticketing platform faces widespread event cancellations. A private-credit lender is considering senior secured financing. Some creators have already received and spent the ticket proceeds.

The evidence

Ticket proceeds, obligations to creators, refund and chargeback exposure, and dated financing terms. Each task must use only the information available at its decision date.

The issue

Claude Fable 5Model

If creators cannot fund refunds, chargebacks can fall on the platform. A lender needs to establish how much cash is available to meet that obligation before judging whether the loan can be repaid.

How the models answeredWhat they got right and what they missed

Dissei’s findings across this case. These describe patterns in the answers, rather than a scored assessment of one response.

Correct calculations

The answers computed the financial ratios correctly.

The key question

The answers focused on when live events would return. They missed whose cash the platform held and how much of the refund obligation it would have to fund.

Accurate covenant citations

The answers cited the covenant package accurately and produced competent-sounding committee memos.

A clear conclusion

The grader notes describe generic committee language without a verdict. The answers did not turn the ratios and covenant terms into a conclusion about repayment.

From the grader’s notes

“generic bank-committee language, never committed to a verdict”
“answers the obvious question instead of the real one.”

These comments come from the grader. They are not quotations from a model’s answer.

The financingStaged funding and equity participation

The May 2020 documents show these terms. A task set in March cannot rely on these later terms.

Total facility
Up to $225M
Initial loan
$125M, expected to be drawn in May, according to the filing.
Delayed draw
Up to $100M, available December 31, 2020 through September 30, 2021, subject to conditions.
Equity participation
2,599,174 Class A common shares at $0.01 per share.
Our interpretationStaged funding limited initial exposure while the refund liability became more observable. Equity participation offered compensation for risk that spread alone might not capture. The documents establish the terms; this interpretation is Dissei’s.

Refunds can consume cash needed for debt service. That does not mean every refund claim has legal priority over secured debt.

The later evidenceThe reserve, the platform’s payments and creators’ refunds
Increase in reserves
$76.5M

For potential chargebacks and refunds.
First quarter of 2020.

Paid by the platform
<$3M

Refunds and chargebacks.
Start of March through the May 11 report.

Refunded by creators
>$150M

Payments funded by event creators.
Same disclosed period.

The reserve increase is an estimate of potential chargebacks and refunds. The payment figures show cash paid during a stated period. They do not establish the final loss or prove that the reserve was excessive.

What this showsA reviewer can check whether an answer separates the potential exposure from the amounts paid by the platform and by creators. Later disclosures help test the earlier reasoning. They must not be treated as facts the model could have known at the decision date.
What to checkDecision date, ownership of cash and the meaning of loss

Use these checks when reading an answer. A pass or fail requires the recorded response; none is assigned here.

Use only what was known then
Keep May financing terms and later payments out of a task anchored in March.
Establish whose cash it is
Separate the platform’s own funds from ticket proceeds owed to creators and refunds owed to buyers.
Distinguish exposure from loss
A reserve, cash paid during a period and a final loss measure different things. The disclosed payment figures do not establish the final loss.
The resultsResults across two model families and three attempts per task

Two unnamed model families used identical environments and sealed rubrics, at three rollouts per task. The rates below describe those two families across this case. They are not an individual score for Claude Fable 5.

12.8percentage points

Difference in pass@1

Pass@1 measures success with one attempt. The families had similar rates of passing critical grading requirements, but their graded outcomes differed. Close rates alone do not establish statistical equivalence.

Reported rates for this case. Family labels preserve the order of the evaluation summary.
MeasureFamily AFamily B
Critical-gate pass rate82.3%83.0%
Low-graded outcomes24.8%11.3%

Three specific tasks scored 0.000 across every rollout and model in the reported run set. One concerned the all-in financing cost and repeatedly failed a critical criterion. This is a task-level result, not a score for the entire case.

The weaker run also produced shorter answers. This observation does not establish a cause. The summary does not include raw counts, a definition of low-graded outcomes or uncertainty intervals, so it cannot establish a general model ranking or statistical significance.

Example 02 / Explanatory · Secured credit

Why pari passu paper trades at different prices

Evidence as of December 2022

Restated question

Why do the company's first-lien instruments trade at different prices? Compare legal rights, cash flows, effective maturities, liquidity and downside recoveries, check any springing terms, and identify the evidence needed to assess relative value.

Situation

A private credit fund weighing a first lien term loan quoted around 66 in a specialty pharmaceutical company months out of Chapter 11, whose equally secured notes trade far higher.

Evidence the model could retrieve

The memorandum’s capital-structure appendix with every tranche’s quote, its key-terms sheet, the maturity and coupon table, the liquidation analysis, the scenario cases, and comparable-company tables.

Credit and legal interpretation

GPT-5.6 SolModel · one recorded attempt

The price gap alone does not establish a difference in legal priority or prove that the lower-priced instrument is cheaper. Equal ranking does not mean identical cash flows, contractual rights or credit exposure over time.

Recorded score · original task
46.4 / 100
Gate
Passed
Rubric
6 of 10
Pitfalls
1 of 4 triggered
Effort
12 turns · 26 tool calls
When equal recovery is a valid assumptionDefine the legal rights and the recovery scenario

These are assumptions for a simplified downside comparison, not verified facts about this case. Check them against the agreements and applicable law.

Equivalent rights to payment and security
The same obligors, collateral and guarantees, with valid, perfected and enforceable equal-ranking liens, equal payment priority and ratable sharing without instrument-specific preferences.
A common recovery event
All instruments remain outstanding when default occurs before the earliest effective maturity, with no intervening preferential repayment or change in priority.
The same recovery measure
The same recovery value, timing and form per unit of allowed claim. Accrued interest, amortisation and other allowed claim components can make that different from equal dollars per unit of original principal.
Separate repayment scenarios
Earlier-maturing debt may be repaid or refinanced before a later default. Earlier maturity creates an earlier payment entitlement, not a higher lien priority. Coupons and repayments before default can also produce different total returns.

Shared collateral does not establish one bankruptcy estate or automatically place every instrument in the same Chapter 11 class. Classification and treatment are separate legal questions. See 11 USC §1122 and §1123.

Why equally ranked instruments can trade differentlyCash flows, credit exposure, contractual rights and market conditions
Coupons and interest rates
Coupon size, fixed or floating rates, benchmark floors and reset dates change expected cash flows and sensitivity to rates. Equal ranking does not require equal yields or prices.
Maturity and refinancing
Different repayment dates expose holders to different periods of default risk and different refinancing outcomes. Check effective maturity, including any springing provision, rather than relying on the stated date alone.
Amortisation and optionality
Scheduled paydowns, cash sweeps, call protection, make-whole provisions and prepayment rights change the amount and timing of cash received.
Covenants and control rights
Financial tests, debt and lien restrictions, amendment thresholds, enforcement rights and cross-default or cross-acceleration provisions can differ despite equal lien priority.
Liquidity and investor demand
Trading depth, bid–ask spreads, transfer restrictions, investor mandates and forced selling can affect prices. The case must supply evidence before any specific technical explanation is treated as established.
Comparable price quotations
Use the same observation date and currency basis, and reconcile accrued interest, settlement conventions and the principal amount outstanding. A quoted cash-price gap is not itself a relative-value calculation.

Compare yields and spreads on a consistent basis, but also compare scenario cash flows and recoveries. Yield to maturity assumes the promised payments occur; it does not settle a distressed comparison. FINRA explains bond cash flows, interest-rate risk, default risk and liquidity risk.

Springing covenant or springing maturity?Identify the trigger and its effect on each instrument
Springing financial covenant
A financial test applies when a specified condition is met, such as a defined level of revolving-facility usage. Establish the trigger, measurement date, facilities protected, cure and waiver rights, and consequences of breach. Do not assume a revolver covenant directly binds every term loan or note.
Springing maturity
The repayment date moves forward if a specified condition is met, for example if other debt remains outstanding near its maturity. The nominally later-dated instrument may then become due earlier. Establish the exact dates, thresholds and exceptions.
Effect across facilities
Trace any cross-default or cross-acceleration provisions and applicable thresholds. A breach, acceleration and the loss of borrowing availability are different outcomes; the agreements determine which follows.

The published assessment does not supply a springing clause or its trigger. Neither type of provision is asserted as a fact of this case. A separate SEC-filed example of a springing maturity illustrates the mechanism; it is not evidence for this December 2022 case.

Inspect the recorded assessmentWhat the model got right, what it missed, and pitfall checks

Original task

At the anchor date the company has several separate instruments outstanding within its first lien class. What should an investor conclude from where each of those instruments trades in the market?

Recorded assessment summary

The model proved the four instruments rank equally and named the holder-base technical, but never put the tranches on a yield basis, skipped the recovery grid, and treated the discount as free money.

The recorded grading includes stronger claims about identical credit risk, one estate and how much of the price gap a yield comparison explains. Equal ranking alone does not establish those conclusions. They require the agreements, a stated recovery scenario and the supporting calculations.

Selected criteria from the recorded assessment. The reported score is not recalculated from this selection.

What the model got right

  • The class does not clear at one price
    Met

    Read the four first lien quotes off the capital-structure appendix and named the term loans as the low end of the class.

    Model excerpt“Four first-lien instruments are, on the documents, a single pari passu class secured by the same collateral and the same guarantors — yet they trade across a 15-point band (66 / 67 / 75 / 81).”
    Gate · weight 25
  • Commits to what the gap represents
    Met

    Settled on one account: the gap is timing and exit optionality plus technicals, not differential credit risk.

    Model excerpt“The dispersion within the class says the market is not pricing lien position at all — it is pricing timing and exit optionality.”
    Gate · weight 20
  • Parity grounded in the security language
    Met

    Cited the annual report’s same-assets, same-guarantors language and the key-terms sheet’s equal-and-ratable lien, rather than asserting parity from the ‘first lien’ label.

    Model excerpt“1 Lien (equal and ratable with liens securing the 1 Lien Bonds) on substantially all assets of the Issuers and the Guarantors.”
    Important · weight 12
  • A driver that is not credit
    Met

    Named the loan and bond markets’ different holder bases as a technical driver, while stating that the case material does not establish the company’s holder composition.

    Important · weight 10
  • Which claim to own
    Met

    Concluded that the term loans at 66 to 68 are the cheap way to own the identical claim.

    Optional · weight 5

What it missed

  • Fixed coupons against a floating loan
    Missed

    Identified the fixed-versus-floating split but argued it points the wrong way and dropped it. Never put the four instruments on a yield basis, which is where most of the 15-point price gap disappears for an identically ranked claim.

    Model excerpt“In a rising-rate environment a floater should trade better, not 9–15 points worse.”
    Important · weight 10
  • Distance to the 2025 refinancing wall
    Partial

    Named the 2025 maturity wall from the debt-service table, but treated the 2028 note trading above the 2027 loans as unexplained rather than pricing each instrument’s time to a par event.

    Important · weight 10
  • One estate in the downside cases
    Missed

    Quoted the 47% to 64% liquidation recovery but never used ratable recovery from one estate to rule out differential credit risk. Instead read part of the spread as intra-class subordination risk, which the record does not support.

    Model excerpt“Pari passu today is not pari passu tomorrow.”
    Important · weight 10
  • Valuation uncertainty is shared
    Missed

    Never touched the 69% to 132% implied-recovery grid or the half-multiple it is struck on. Used a single 2.69x market-implied multiple as if it settled first lien coverage.

    Important · weight 8
  • What a real ranking gap costs
    Missed

    Did not set the intra-class spread against the far deeper discount on the junior claims to show what a genuine ranking difference does to price.

    Optional · weight 4
  • Calling it a free arbitrage
    Pitfall triggered

    Presented the loan at 66 as a 15-point discount for the same claim without netting the $200mm instalment and the 2025 wall that the whole class must fund first.

    Model excerpt“You have overpaid by 15 points for the same claim.”
    Pitfall · penalty -2

Pitfall checks

  • Hindsight after the anchorNot triggered

    Every dated claim was cited to case sources; nothing after December 2022 asserted as an outside fact.

  • Reading the price gap as a ranking gapNot triggered

    Affirmed identical collateral and guarantees; the gap attributed to timing, not to a weaker lien.

  • Blaming a covenant the loan does not haveNot triggered

    Covenants cited only as evidence the instruments are equivalent.

  • Calling it a free arbitrageTriggered

    Discount framed as overpayment on the notes, with the class-wide cash calls left out of the comparison.

Read the model’s answer excerpt

Recorded model output. Read alongside the assessment; this excerpt can contain errors.

Buying the two term loans at 66–67 gets you the same lien, the same collateral, the same guarantee package and the same covenant suite as the April 2025 notes at 81 — a ~15-point discount for a two-year maturity difference within a class that, if the restructuring thesis is right, gets treated as one pool anyway. If you believe the company restructures, the April 2025 note’s near maturity buys you nothing and you have overpaid by 15 points for the same claim. If you believe it refinances, the loans re-rate hardest.
Public dataset · question previews

Dissei Financial Judgment

Inspect the questions.

Seven questions, seven reasoning families, one completed private-equity transaction. Inspect how each question frames a financial decision at an information cutoff.

Browse question briefs, metadata, reported pilot results and methodology. Supporting case evidence and answer keys are not public.

Public previews, not a runnable evaluation. These read-only question previews leave out evidence, rubrics, reference answers, evaluators, tools and private records. Scores come from complete tasks under controlled access. Rights to evaluate, review, train on or redistribute require an agreement.

Commercial license by agreement. Public availability does not grant training, redistribution or underlying benchmark rights. Read the access policy

Hugging FaceBrowse and download question previewsHarbor HubRead task briefs, results and methodology

Harbor is the agent-evaluation framework from the makers of Terminal-Bench. It hosts Dissei’s reported results; hosting is not independent certification.

Release record

Release and coverage

New releases replace the results on this page; the address stays the same.

Release
historical-2026-08-31 (historical aggregate snapshot)
Results as of
Model names checked
Models covered
5: GPT-5.6 Sol, Claude Opus 4.8, Muse Spark 1.2, Kimi K3, DeepSeek v4-flash
Tasks and runs
Not included in the public aggregate, so no totals are stated here.
Source
Historical Dissei benchmark aggregate export
Publisher
Dissei
Methodology

What these results measure

Financial reasoning grounded in evidence.

From evidence to a decision.

Financial work combines calculation, investigation, and judgment. The leaderboard brings those recorded results into one view: how a model interprets evidence, tests assumptions, explains a mechanism, and supports a decision.

Values and run methods are shown as originally recorded. Score is continuous reward per 100. Gate pass is the share of attempts that cleared required checks.

Financial evidenceModel analysisEvaluation

How results are pooled.

Each model's score, gate pass and category results are pooled across the cases it attempted. Case coverage differs by model, so the table is a descriptive record, not a matched comparison.

Category results apply the same two measures within each of the seven reasoning categories. Cost is an estimate for the policy model only and excludes grading and infrastructure.

Score

Score is a continuous reward scaled 0–100. It is not accuracy and not the gate pass rate.

Gate pass

Gate is the percentage of attempts that meet the required checks.

Cost & resources

Cost is the estimated inference cost of the policy model. Token and tool-call counts are reported separately as resource metrics.

Citation

Cite these results

Cite the release and its results date, since results at this address change with each release.

Dissei (2026). Dissei Leaderboard: financial reasoning results for AI models, release historical-2026-08-31, results as of 2026-08-31. https://dissei.ai/leaderboard

Build on the evidence

Bring financial reasoning into your evaluations.

Explore a case and its assessment, then talk to us about evaluating your models under an agreement.