Benchmark

Finally, an AI benchmark built for payments.

General AI benchmarks can’t tell you how accurately AI models respond to chargebacks, analyze processing statements, or complete other real‑world payments tasks. The Findustry AI Benchmark is the industry’s first benchmark measuring LLM performance for the workflows that matter in payments. Read more.

Chargebacks Performance Leaderboard

Updated July 24, 2026

Fable 5 + Findustry AI
98.3
Grok 4.5 + Findustry AI
96.7
GPT 5.6 Sol + Findustry AI
96.7
Sonnet 5 + Findustry AI
96.7
GPT 5.5 + Findustry AI
96.7
Opus 5 + Findustry AI
96.7
Gem. 3.5 Flash + Findustry AI
95.0
GPT 5.4 + Findustry AI
95.0
Opus 4.8 + Findustry AI
95.0
GPT 5.6 Luna + Findustry AI
91.7
Gemini 3.6 Flash + Findustry AI
91.7
GPT 5.6 Terra + Findustry AI
91.7
Opus 4.7 + Findustry AI
91.7
Gem. 3.1 Pro + Findustry AI
90.7
GLM 5.2 + Findustry AI
88.3
Opus 5
81.7
Fable 5
80.0
Gem. 3.1 Pro
78.9
Gem. 3.5 Flash Lite + Findustry AI
78.3
Grok 4.5
78.3
Gemini 3.6 Flash
73.3
Gem. 3.5 Flash
73.3
Opus 4.8
70.0
GPT 5.6 Luna
66.7
Opus 4.7
65.0
Sonnet 5
65.0
Gem. 3.5 Flash Lite
63.3
GPT 5.4
58.3
GLM 5.2
55.0
GPT 5.5
55.0
GPT 5.6 Terra
41.7
GPT 5.6 Sol
36.7
30405060708090100
Rankings · 32 Models
Model
Score
Cost
Fable 5+ Findustry AI
98.3
$4.72
Grok 4.5+ Findustry AI
96.7
$0.90
GPT 5.6 Sol+ Findustry AI
96.7
$1.51
Sonnet 5+ Findustry AI
96.7
$1.71
GPT 5.5+ Findustry AI
96.7
$1.74
Opus 5+ Findustry AI
96.7
$1.89
Gemini 3.5 Flash+ Findustry AI
95.0
$0.31
GPT 5.4+ Findustry AI
95.0
$0.61
Opus 4.8+ Findustry AI
95.0
$1.21
GPT 5.6 Luna+ Findustry AI
91.7
$0.22
Gemini 3.6 Flash+ Findustry AI
91.7
$0.29
GPT 5.6 Terra+ Findustry AI
91.7
$0.64
Opus 4.7+ Findustry AI
91.7
$1.21
Gemini 3.1 Pro+ Findustry AI
90.7
$0.51
GLM 5.2+ Findustry AI
88.3
$0.22
Opus 5
81.7
$1.36
Fable 5
80.0
$1.87
Gemini 3.1 Pro
78.9
$0.21
Gemini 3.5 Flash Lite+ Findustry AI
78.3
$0.07
Grok 4.5
78.3
$0.25
Gemini 3.6 Flash
73.3
$0.08
Gemini 3.5 Flash
73.3
$0.10
Opus 4.8
70.0
$0.86
GPT 5.6 Luna
66.7
$0.11
Opus 4.7
65.0
$0.66
Sonnet 5
65.0
$0.68
Gemini 3.5 Flash Lite
63.3
$0.02
GPT 5.4
58.3
$0.22
GLM 5.2
55.0
$0.05
GPT 5.5
55.0
$0.78
GPT 5.6 Terra
41.7
$0.25
GPT 5.6 Sol
36.7
$0.59

Score reflects accuracy on chargeback response scenarios. Cost is per full run of evaluations. Models ranked by score with Findustry AI harness; base model results shown for comparison only.

Research

Introducing the
Findustry AI Benchmark

The first benchmark measuring LLM performance for the payments industry —enabling us to choose the right models and improve performance by more than 40% with a purpose‑built solution.

Jonathan RaziAdam Finlayson

Jonathan Razi & Adam Finlayson

Findustry AI · May 28, 2026

Context

What is a benchmark?

In GenAI, benchmarks measure the performance of large language models, allowing users to compare accuracy, cost, and speed across different models. Public benchmarks often score model performance on tasks such as math problems and coding challenges.

Creating a benchmark requires curating dozens or even hundreds of test scenarios—each with its own prompts—and setting up code‑based evaluators or other LLMs used as judges to measure each model’s performance in response to the test cases.

Because general benchmarks focusing on math or coding problems are often not correlated with product needs within a specific industry, some researchers and companies have recently produced benchmarks for domain‑specific tasks, such as stock analysis based on SEC filings. However, until now, no benchmark has existed to measure LLM performance for payments industry use cases, such as chargeback workflows.

That’s where the Findustry AI Benchmark comes in.

No benchmark has measured LLM performance for payments industry use cases—until now.

Why we built it

Why create the Findustry AI Benchmark?

The Findustry AI Benchmark began with the questions we wanted to ask while building our own product.

Each time a new model is released by the labs, it is accompanied by academic benchmarks showing higher scores on instruction following, coding, math, and other tasks with mass AI adoption. But how reliable are those scores if the test problems themselves are in the training sets? And do the scores on horizontal tasks necessarily imply improvement for the needs of our vertical AI application, such as generating chargeback responses, matching evidence to network rules, and drafting rebuttal letters?

Put simply: how do we ensure we are choosing the right LLM for our customers?

Before the Findustry AI Benchmark, there were no rigorous, quantitative answers to these questions. The Benchmark is the result of building the evaluations we needed to measure model capability inside of our own software.

Benefits

Are there benefits in addition to selecting the right model?

Yes. These evaluations give us the ability to systematically measure our AI application’s performance, which is especially important because GenAI systems are nondeterministic. Without evaluations, developers update prompts and simply hope they didn’t break anything else. In contrast, our evaluations ensure our changes are grounded in data rather than vibes, resulting in the enterprise‑grade quality our customers demand.

Also, because the frontier labs frequently update their models even between major releases, evaluations enable us to detect regressions when model behavior changes behind the scenes. This is similar to the muscle we built at our prior company, CardX, where we had to monitor downstream vendors such as gateways and acquirers.

Findings

What are the findings?

With Findustry AI
Base Model
GPT 5.5 + Findustry AI
96.7
Gem. 3.5 Flash + Findustry AI
95.0
GPT 5.4 + Findustry AI
95.0
Opus 4.8 + Findustry AI
95.0
Opus 4.7 + Findustry AI
91.7
Gem. 3.1 Pro + Findustry AI
90.7
Opus 4.6 + Findustry AI
88.1
Grok 4.3 + Findustry AI
86.7
Sonnet 4.6 + Findustry AI
85.6
Haiku 4.5 + Findustry AI
85.5
Gem. 3.1 FL + Findustry AI
83.3
Grok 4.20 Reas. + Findustry AI
83.3
Gem. 3.1 Pro
78.9
Grok 4.3
78.3
GPT 5.4 Mini + Findustry AI
76.7
Grok 4.20 Reas.
76.7
GPT 5.4 Nano
75.0
GPT 5.4 Nano + Findustry AI
73.3
Gem. 3.5 Flash
73.3
Sonnet 4.6
71.1
Opus 4.8
70.0
Opus 4.6
68.4
Haiku 4.5
67.1
Opus 4.7
65.0
GPT 5.4 Mini
65.0
Gem. 3.1 FL
63.3
GPT 5.4
58.3
GPT 5.5
55.0
30405060708090100

Larger, more expensive models may be outperformed by smaller, faster, and cheaper models for payments use cases.

A domain‑specific benchmark may not align with the general capabilities, such as long‑running reasoning, that make larger models superior on standard benchmarks. For focused, well-defined payments tasks, we find that the base Sonnet 4.6 model beats the larger Opus 4.6 while being faster and cheaper.

Small models can excel when specialized for industry context.

Smaller models can prove particularly effective when specialized for the relevant industry context: paired with Findustry AI’s domain intelligence, Gemini 3.1 Flash Lite surpasses the performance of the base (unspecialized) Gemini 3.1 Pro model, despite Flash having 76% lower costs and 92% lower latency.

The most recent model release from the same frontier lab is not necessarily best.

In our payments‑specific scenarios, the base Opus 4.6 model outperforms Opus 4.7 despite Opus 4.7 advertising an 11% higher score on the coding benchmark SWE‑Bench Pro. This means that a strategy of updating to the latest model release without systematically measuring its performance in the relevant industry context may be detrimental to product accuracy, emphasizing the importance of evals.

Score improvement

The base LLMs, which are largely trained on the open internet, do not match the scores of a vertical AI system that leverages domain expertise and proprietary data.

The Findustry AI harness—which includes our prompts, few‑shot examples, and tool outputs—increases the score of GPT‑5.5 by 41.7% compared to the same model tested without our harness.

Score Improvement: Base Model vs. With Findustry AI Harness

Chart showing score improvement when using the Findustry AI harness versus the base model across different LLMs

Performance on chargeback‑related tasks and corresponding cost to accomplish, by LLM model, with and without the Findustry AI harness. Updated May 28, 2026.

Below, side‑by‑side response comparisons demonstrate that one model, Opus 4.6, fails to understand the nuances of the chargeback process “out of the box” but, tested with our harness, has the domain depth to correctly respond.

Response Comparison: Procedurally Invalid Chargeback · Opus 4.6

AcceptBase Model

Fraud chargeback (Visa 10.4) for a $47.36 online flower order. AVS failed (no street or ZIP match), and there is no prior transaction history or compelling evidence to prove the cardholder authorized this transaction. Recommending acceptance.

RespondBase + Findustry AI Harness

This dispute is procedurally invalid under card network rules. The transaction occurred on August 11, 2025, and the chargeback was filed more than 120 days after the transaction date, exceeding the permissible dispute timeframe. On this basis alone, the chargeback should be reversed in the merchant’s favor.

Without our harness, Opus 4.6 accepts (concedes) a chargeback despite the dispute being filed outside of the date range permitted by card network rules. With our harness, the model correctly fights the chargeback on the grounds that it is procedurally invalid.

Response Comparison: Prior Refund · Opus 4.6

AcceptBase Model

We are writing in response to the above‑referenced chargeback dispute. After careful review of the case, the merchant, EchoWave Media, has elected to accept this chargeback.

Although we note that AVS returned a full 5-digit zip and address match and that the CVV was verified as correct at the time of authorization, these alone are insufficient.

RespondBase + Findustry AI Harness

Responding based on a full refund already issued before the dispute was filed, rendering this chargeback procedurally invalid. AVS/CVV exact matches and IP geolocation consistency further support the case.

A refund of $89.53 has already been sent to the cardholder on February 8, 2026 (ARN: 83981349912448002576570), prior to the filing of this dispute. Reversing this chargeback is necessary to prevent a double credit to the cardholder.

Without our harness, Opus 4.6 accepts (concedes) a chargeback even when there was a prior refund. With our harness, the model responds to the chargeback by citing the details of the prior refund, including the ARN, to prevent a double credit.

Document parsing

Even frontier models have meaningful error rates parsing PDF documents, which are foundational to many payments workflows.

In the real world, the success of AI systems often depends on extracting accurate information from documents with inconsistent layouts—such as terms of service, screenshots of payment gateway interfaces, and scanned receipts. If the model hallucinates critical information like an ARN, the entire workflow may fail, meaning AI agents are only as good as their document extraction capabilities.

PDF Document Parsing: Text Extraction Error Rate by Model

<10%
10–20%
>20%
Gem. 3.1 Pro
2.9%
Gem. 3.5 Flash
3.0%
Opus 4.8
3.1%
Opus 4.6
3.5%
Opus 4.7
3.6%
Gem. 3.1 FL
6.3%
Grok 4.20 Reas.
18.0%
Haiku 4.5
18.5%
Grok 4.3
22.7%
Sonnet 4.6
27.3%
GPT 5.5
29.4%
GPT 5.4 Mini
32.8%
GPT 5.4
33.0%
GPT 5.4 Nano
33.3%
05101520253035

Error Rate (%)

Text extraction error rate based on word count; lower percentage scores are better.

Statement analysis

Analyzing merchant processing statements is challenging for LLMs due to variation in formats across every acquirer.

To solve this problem, we designed our Findustry AI agent to learn from each processor format it sees. Without this specialized intelligence, no model achieves a score higher than 90% in our evaluations. Paired with the Findustry AI harness, however, both Gemini 3.1 Pro and GPT‑5.5 reach 97.9% accuracy.

Processing Statement Analysis: Accuracy Score by Model

With Findustry AI
Base Model
GPT 5.5
97.9
86.5
Gem. 3.1 Pro
97.9
85.4
Gem. 3.5 Flash
95.8
87.5
Opus 4.7
93.8
85.4
Opus 4.6
91.7
88.5
Grok 4.3
91.7
86.5
Sonnet 4.6
90.6
89.6
Grok 4.20 Reas.
89.6
87.5
GPT 5.4
89.6
77.1
Gem. 3.1 FL
88.5
81.3
GPT 5.4 Mini
86.5
70.8
Haiku 4.5
85.4
82.3
GPT 5.4 Nano
75.0
63.5
406080100

Score

Accuracy score on processing statement analysis tasks. Variation in processor formats across acquirers makes this a uniquely challenging task for base LLMs. The Findustry AI harness enables the agent to learn from each processor format it encounters.

Roadmap

What comes next?

Our team will keep updating the Findustry AI Benchmark as new models are released. Our goal is for the Benchmark to be a valuable resource for our partners in the payments industry as the fast‑moving GenAI space evolves.

And, as our platform expands, we will extend the Findustry AI Benchmark from chargebacks and document extraction to additional task categories in payments. Based on our internal testing, we believe the scores will continue to demonstrate there is no single model that is best for all use cases; rather, different models perform better at different tasks, meaning the optimal AI system has the sophistication to know which LLM it should call at each step in its workflows.

No single model is best for all use cases—the optimal AI system knows which LLM it should call at each step in its workflows.

From the time we started Findustry AI, we knew we wanted to create a company that excelled at both product and research. Just like our first venture, CardX, was both a product and research organization in the surcharging field—authoring a U.S. Supreme Court brief and helping to change four state laws—Findustry AI will continue to make investments at the intersection of payment processing and AI research.

Especially in an environment where many companies are “AI washing,” we believe customers are looking for differentiated insights and trusted partners in AI transformation.

Ready to put this intelligence to work?

If we can help your organization, or if you have feedback or requests, contact our team.