Teaching models
financial judgment.
Carefully built training data, evaluations and environments, grounded in real financial work, for AI labs and enterprise teams.
Looking for financing or a lending partnership? Visit Dissei Lending.
Learning, under examination.
Studies in what models learn from finance, and where their judgment breaks.
Practitioner guidesFinancial AI LeaderboardFinancial reasoning,
Financial reasoning,
measured.
How AI models and LLMs reason from financial evidence, compared across seven reasoning categories and scored on a 0–100 continuous reward scale, not accuracy.
Pooled evaluation scores by model
Historical snapshot · 2026-08-31
Claude Opus 4.8: Observed score / 100 44.26. DeepSeek v4-flash: Observed score / 100 25.49. GPT-5.6 Sol: Observed score / 100 45.53. Kimi K3: Observed score / 100 39.76. Muse Spark 1.2: Observed score / 100 42.41.
Observed score / 100