If you’ve been following Findustry AI since our debut last month, you’ve heard us talk about measuring the differences between LLMs as foundational to how we build our products. Here’s an illustration of why this matters:
For the first time, we’re seeing significant divergence between the flagship models of Anthropic and OpenAI in the payments domain. Tested “out of the box”—with just a basic prompt and without our proprietary data and tools—Fable 5 scores 80% whereas GPT-5.6 Sol scores 36.7% in our Findustry AI Benchmark.
Bottom line: these models both boast impressive general capability but differ widely when used for payments workflows. If you’re using AI inside your organization, you—or your partners—need to be able to compare models for the tasks that matter to you.

