[stock-market-ticker symbols="FB;BABA;AMZN;AXP;AAPL;DBD;EEFT;GTO.AS;ING.PA;MA;MGI;NPSNY;NCR;PYPL;005930.KS;SQ;HO.PA;V;WDI.DE;WU;WP" width="100%" palette="financial-light"]

Banking AI trained to admit uncertainty resolves more customer queries at a fraction of the cost

31 iulie 2026

Backbase, the company behind the AI-native Banking OS, released the findings of a peer-reviewed study on production-grade banking AI. It was led by Denys Katerenchuk-  Head of AI Research at Backbase, previously of Google and IBM. It is among the first peer-reviewed accounts of a banking-grade language model measured in live production.

The hardest problem in customer-facing banking AI is what the model does when the evidence isn’t there. A system that invents an answer about a fee, a rate, or a policy creates regulatory exposure. But one that declines too often becomes useless. The study shows this trade-off can be engineered out. A 12-billion-parameter model trained to recognize the limits of its own evidence resolved significantly more customer queries in live deployment. It also outperformed GPT-4.1 on the quality and grounding measures that matter most in a regulated environment, at a fraction of the cost.

Off-the-shelf models tend toward hallucination and sycophancy – confident, agreeable answers even without evidence. That’s especially risky in banking, where information is complex, technical, and scattered across dozens of documents,” explained Katerenchuk.

The stakes of this disconnect are well documented. McKinsey estimates AI could drive up to 20% in net cost reductions for banks. Yet MIT research found that 95% of enterprise generative AI pilots deliver no measurable P&L impact.

Katerenchuk and his team trained a model to understand the domain and recognize when information is incomplete. They taught it the boundaries of its own knowledge, so it says “I don’t know” instead of inventing an answer.

Key findings:

. Honesty can be engineered: The model was trained on a dataset in which 22% of examples had no correct answer. That taught it the right response was an explicit “I don’t know.” It reached a 12% refusal rate, higher than the untuned base model’s 4.3% and lower than GPT-4.1’s 20.2%. The base model answered confidently even without evidence. GPT-4.1 declined questions it could have safely answered.

. The honest model solved more customer problems: Over seven months at a large US financial institution, query resolution rose 7.1 percentage points across 3,297 sampled queries. That’s a statistically significant gain. It came despite the model refusing nearly three times as often as its base.

. A model a fraction of the size beat GPT-4.1: The model scored higher on independent evaluation: 6.21 versus 5.72 for GPT-4.1, on a 10-point scale. Citation grounding improved by 2.3 points, with answers cited directly to source documents. It also produced stronger results on FinanceBench, a public benchmark of SEC filing questions.

. The economics undercut frontier models by an order of magnitude: Roughly $0.001 per query on a single GPU, 20-50x cheaper than GPT-4.1 and 3-5x faster. Training cost around $1,800.

. Data order mattered more than data volume: On identical data, teaching general financial language first and calibrated refusal second produced the best model. Pooling everything at once collapsed answer quality by more than 40% and pushed refusals to nearly half of all queries.

The findings land amid a live industry debate: research published by OpenAI in 2025 found that training and evaluation methods reward confident guessing over admitting uncertainty. The Backbase study offers production evidence of the alternative: a model rewarded for honesty, measured against real customers.

For three years, the AI industry has rewarded models for speed and confidence. Banking has rewarded itself for the same thing for three decades. Saying ‘I don’t know’ got treated as a weakness, not a feature,” said Jouk Pleiter, Founder and CEO of Backbase. “Our research shows the opposite: a model that knows the limits of its own evidence earns more trust, not less.”

2026 is the year agentic workflows go live in regulated environments, but none of it works unless the model knows what it doesn’t know,” added Pleiter.

This study is also the first published work from Backbase AI Research – the team that joined Backbase through its acquisition of Kasisto. The group focuses on the specific problems of AI in banking, publishing peer-reviewed research openly and moving findings directly into production.

_______

The full paper, FinRAG-12B: A Production-Validated Recipe for Grounded Question Answering in Banking, was published in the Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Industry Track), July 2026. The production analysis covers 3,297 randomly sampled customer queries over seven months (May–December 2025) at a large US retail credit union; institutions in the study are anonymised.

Noutăți
Stay updated to the impact of emerging technologies in fintech & banking.
Banking 4.0 newsletter - subscribe
Cifra/Declaratia zilei

Dariusz Mazurkiewicz – CEO at BLIK Polish Payment Standard

Banking 4.0 – „how was the experience for you”

To be honest I think that Sinaia, your conference, is much better then Davos.”

Many more interesting quotes in the video below:

Sondaj