Skip to main content
Get in Touch

Conversational AI in Banking: What It Should Answer

Conversational AI in banking has stopped being a pilot question. Deloitte found that 37% of banking executives already use generative AI in their contact centres, with another 37% planning to start in 2026 (Deloitte, 2026). The harder question begins after the assistant understands the customer: whether it can resolve the request, or only describe it and route it onward. That answer is architectural. It is decided by the systems underneath the chat window, by what the Building Platform layer beneath the assistant is able to expose as a governed, auditable operation, which is why this analysis works from the architecture upward rather than from the interface inward.

Conversational AI in Banking: What It Should Answer

The short version. Conversational AI in banking is the layer that talks: chat, voice, and in-app assistants that hold context and answer in natural language. It creates measurable value on servicing status, document collection, and staff knowledge retrieval. It must never issue a credit decision or its reasons, it must not hold the authority to grant anything, and since 2 August 2026 it must tell EU customers they are talking to a machine.

What this analysis covers. Six constraints decide whether a deployment produces savings or relocates work: the difference between deflection and resolution; the intents that must never be automated; the adversarial surface an assistant with tools creates; the personal data that must not reach the context window; the core banking system that usually sets the real ceiling; and the handoff and voice limits that determine whether the last mile works at all.

What conversational AI in banking actually is

Conversational AI is the interface layer of a bank’s AI stack. It conducts dialogue in natural language across chat, voice, and in-app surfaces, holds session context, retrieves information, and either answers or hands the case to a person. It is not a synonym for the other two AI layers banks are deploying in parallel.

The three layers do different jobs, and conflating them is the most common source of failed business cases. Generative AI in banking (opens in new tab) produces: drafts, summaries, translations, code, synthetic test data. Agentic AI in banking (opens in new tab) acts: multi-step planning and tool use across systems. Conversational AI talks: it is how a customer or an employee reaches whatever sits behind it. For the full stack view, see our guide to AI in banking (opens in new tab).

The distinction has a practical edge. A conversational surface with nothing but a document store behind it can answer questions. A conversational surface wired into a lending system’s building blocks can complete transactions. The words on screen look identical to the customer, and the economics, the risk profile, and the audit position are not.

Three AI layers in banking: conversational AI talks, generative AI produces, agentic AI acts, above the lending Building Platform

The three surfaces that matter in lending

Customer self-service is the visible surface: balance, next payment, payoff figure, statement request, payment date change, hardship intake. These are the highest-volume intents in a retail loan book and the ones customers most resent queueing for.

Borrower guidance during origination is the second surface: application status, missing documents, eligibility explanation, next step. It sits directly on conversion, because an abandoned application is a marketing cost already spent.

Staff assist is the third and least glamorous surface, and often the fastest to pay back. Lloyds Banking Group reported that its Athena knowledge-management assistant cut search time by 66% for 20,000 colleagues, and that its HR assistant resolves 90% of queries correctly on first contact (Lloyds Banking Group, 2026 (opens in new tab)). Bank of America reports that its employee-facing assistant is used by more than 90% of employees and halved IT service-desk calls (Bank of America, 2025 (opens in new tab)).

Deflection is not resolution

The metric most conversational AI business cases are built on is containment: the share of conversations that end without a human. The metric that determines whether cost-to-serve actually falls is resolution: the share of customer problems that are finished.

The gap between the two is large and measured. Deloitte found that 70% of bank customers used self-service in the past year, but only 25% said self-service resolved half or more of their issues without human assistance (Deloitte, 2026). A contained conversation that ends with an unresolved problem does not remove a call, it postpones one, and it usually returns as a longer call.

Customer sentiment tracks the same gap. In an RFI Global survey of 4,000 US consumers, only 33% said they trust their bank’s virtual assistant “a lot” to answer product questions, while 63% said they would trust their bank’s AI more if a human was clearly accountable for the outcome (RFI Global, via American Banker, 2026 (opens in new tab)). Independent measurement points the same way: Qualtrics, surveying more than 20,000 consumers across 14 countries, found AI in customer service failing at close to four times the rate of AI use generally, with nearly one in five consumers reporting no benefit at all (Qualtrics, 2025 (opens in new tab)).

Chart contrasting 70% of bank customers using self-service with 25% saying it resolved half or more of their issues

How far conversational AI in banking has actually gone

Bank chatbots are not new, and the baseline is close to universal. The CFPB found that all ten of the largest US commercial banks used chatbots of varying complexity, and cited an estimate that roughly 37% of the US population, over 98 million users, had interacted with a bank’s chatbot in 2022 (CFPB, 2023, citing Insider Intelligence (opens in new tab)). That report is three years old and remains the best regulator-grade baseline available.

What changed in 2025 and 2026 is depth rather than presence. The institutions publishing hard numbers are the ones that rebuilt what sits behind the chat window.

Published conversational AI figures from Bank of America, DBS, NatWest, Lloyds and Wells Fargo, 2025 to 2026
InstitutionPublished figureDate
Bank of America (Erica)3.2 billion cumulative client interactions since 2018 launch, 700 million in the past year, 20.6 million usersMarch 2026
DBSVirtual assistants reach 10 million+ customers across three markets, expected to handle 1 million+ chats a month; 9 in 10 customer queries resolved digitally without follow-up contact in H1 2026; 7% reduction in calls and emailsJuly 2026
NatWest (Cora)Generative AI customer journeys grew from 4 to 21; 70,000+ hours saved through automated call summaries and simplified complaint responses; agentic assistant in Cora for 25,000 customers by end of Q1 2026; voice-to-voice planned for later in 2026February 2026
Lloyds Banking Group£50m of value from AI in 2025, targeting over £100m more in 2026, from 50+ deployed generative AI solutions across 28 million customersJanuary 2026
Wells Fargo (Fargo)245.4 million interactions in 2024, up from 21.3 million in 2023, 336 million cumulative; architecture reported as no PII passed to the modelApril 2025 (reported)

“The true value of AI lies in delivering meaningful outcomes for customers at scale.”

— Derrick Goh, Group Chief Operating Officer, DBS (DBS, July 2026 (opens in new tab))

DBS is the instructive case, because it published the resolution figure rather than the containment figure. Nine in ten queries resolved digitally without follow-up contact is a statement about the systems the assistant can reach, not about the quality of its language.

Where conversational AI works in lending today

Servicing self-service that changes the loan record

The highest-value conversational intents in a loan book are the ones that end in a state change: a payment date moved, a payoff quote issued, a direct debit updated, a hardship request logged. Each one is a short call today and a completed self-service transaction if the assistant can write to the loan.

This is the boundary that separates an information bot from working loan servicing software (opens in new tab). Reading a balance requires a query. Rescheduling a payment requires permission to advance a state machine, apply a policy, recalculate a schedule, and write an audit entry. Most deployments stop at the first because the second was never exposed to them.

Application status and document chase

In origination, the assistant’s job is narrow and valuable: tell the applicant exactly where the file stands, what is missing, and what happens next. Nothing here requires judgment, and all of it currently generates inbound contact.

Wiring the assistant into loan origination (opens in new tab) states rather than a static FAQ turns “your application is being reviewed” into “we have your ID and payslip, we are waiting on your bank statement for July”. That is the difference between a deflected contact and a progressed application.

Collections and hardship first contact

Voice is where banks most want conversational AI next, and where verifiable outcome data is thinnest. Every published figure we could source on voice AI in banking collections or servicing comes from vendor marketing rather than a regulator, consultancy, analyst house, or bank. NatWest has said voice-to-voice capability in Cora is planned for later in 2026 (NatWest, 2026 (opens in new tab)), which is the most concrete public commitment from a major lender.

The absence of independent data is itself a finding, and the technical constraints behind it are specific enough to deserve their own treatment (the voice limits). Hardship conversations also carry conduct risk that a containment metric will never surface.

Staff knowledge retrieval

The internal surface has the best risk-adjusted return in lending. A servicing agent asking a policy question in natural language is not making a commitment to a customer, the failure mode is a wrong answer that a trained employee can catch, and the grounding corpus is the bank’s own policy set.

Complaint intake, with a caveat

Complaint handling is an obvious conversational use case, and it now cuts both ways. The UK Financial Ombudsman Service reported that up to a third of recent complaints appeared to be generated or heavily assisted by AI, including submissions containing fake laws and rulings and misquoted legislation, with one 200-page response to a six-page provisional decision (Financial Ombudsman Service, via Which?, 2026 (opens in new tab)).

Banks are therefore deploying conversational AI into a channel where the counterparty is also using it. Intake and triage benefit from automation. Adjudication and redress do not.

What a conversational layer must not answer

The credit decision and its reasons

A conversational assistant must never make a credit decision, and it must never be the thing that explains one. Under Regulation B, 12 CFR 1002.9(b)(2) (opens in new tab), a statement of reasons for adverse action “must be specific and indicate the principal reason(s) for the adverse action”, and statements that the applicant “failed to achieve a qualifying score on the creditor’s credit scoring system are insufficient”.

The regulatory picture around that duty shifted in 2025 without the duty itself changing. On 12 May 2025 the CFPB withdrew 67 guidance documents, including Circular 2022–03 on adverse action notices and complex algorithms and Circular 2023–03 on the proper use of sample forms (Federal Register, 2025 (opens in new tab)). The interpretive guidance is gone. The statutory obligation in ECOA and the text of Regulation B is unchanged and fully in force.

In the EU the requirement is being tightened rather than loosened. Credit scoring for natural persons is high-risk under Annex III point 5(b) of the AI Act, and Article 86 gives an affected person the right to obtain from the deployer clear and meaningful explanations of the role of the AI system in the decision and the main elements of the decision taken (Regulation (EU) 2024/1689). An explanation that a language model composed on the fly is not an explanation of the decision. It is a plausible narrative about a decision, which is a different and more dangerous object.

The practical rule for architecture is therefore stable across jurisdictions. Reason codes come from a deterministic, explainable decisioning component, and the conversational layer is permitted to read them and present them, never to generate or paraphrase them. We covered that separation in detail in AI agent vs credit scoring (opens in new tab).

Anything binding that is not grounded in the record

A conversational assistant speaks for the institution that deployed it. In Moffatt v. Air Canada, 2024 BCCRT 149 (opens in new tab), a British Columbia tribunal ordered the airline to pay CAD 812.02 in damages, interest, and fees for negligent misrepresentation after its chatbot gave a customer inaccurate information about fare procedures, rejecting the argument that the chatbot was a separate entity: “While a chatbot has an interactive component, it is still just a part of Air Canada’s website. It should be obvious to Air Canada that it is responsible for all the information on its website.”

This is a small-claims tribunal decision in one Canadian province, not binding precedent anywhere. It is useful because it states the position a supervisor would take, and because the equivalent position has been stated by a banking regulator. The Hong Kong Monetary Authority’s August 2024 circular on consumer protection in the use of generative AI requires that “the board and senior management of authorized institutions should remain accountable for all the GenAI-driven decisions and processes”, that institutions adopt a human-in-the-loop approach, and that customers be given the option to opt out of generative AI and request human intervention (HKMA, 2024 (opens in new tab)).

The CFPB reached the same conclusion from the consumer-harm side, finding that advanced chatbots “often generate incorrect outputs that are undetectable by some users”, that customers get “stuck in a loop of unhelpful jargon”, and that deficient chatbots can impede the exercise of statutory dispute rights (CFPB, 2023). The mitigation is grounding rather than fluency, which is the subject of our piece on RAG grounding against AI hallucinations in lending software (opens in new tab).

The fact that it is an AI

Since 2 August 2026, Article 50 of the EU AI Act requires that AI systems intended to interact directly with natural persons inform those persons that they are interacting with an AI system, unless that is obvious to a reasonably well-informed observer. The European Commission’s guidance states the notification must come “from the start of the first interaction in a clear and distinguishable manner”, and that the “obvious” exception is to be interpreted restrictively (European Commission, 2026 (opens in new tab)). Breaches of Article 50 sit in the penalty tier of up to €15 million or 3% of worldwide annual turnover, whichever is higher (Regulation (EU) 2024/1689, Article 99(4) (opens in new tab)).

The Digital Omnibus on AI, Regulation (EU) 2026/1744, in force since 27 July 2026, did not defer Article 50. What it deferred was the high-risk regime: stand-alone Annex III systems, which include AI used to evaluate the creditworthiness of natural persons, now apply from 2 December 2027 (Regulation (EU) 2026/1744 (opens in new tab)). A bank’s assistant is subject to transparency obligations today and its scoring engine is on a 2027 clock.

The US supervisory position is more surprising and worth stating precisely. SR 26–2 (opens in new tab), issued 17 April 2026 by the Federal Reserve, OCC, and FDIC, supersedes SR 11–7 on model risk management, and the companion OCC Bulletin 2026–13 states that “Generative AI and agentic AI models are novel and rapidly evolving. As such, they are not within the scope of this guidance”, with an interagency request for information planned (OCC, 2026 (opens in new tab)). Governance of a customer-facing assistant is not inherited from the model risk framework. It has to be built, which is the ground covered in our guide to passing an AI lending compliance audit (opens in new tab).

European supervisors are pointing at the same accountability gap from the other direction. Speaking in February 2026, ECB Supervisory Board representative Pedro Machado noted that generative AI is now in “front-line applications, such as customer support, relationship management and internal knowledge tools”, and put the test simply: “if a bank cannot explain why an AI model behaves the way it does, then it cannot truly control that model” (ECB, 2026 (opens in new tab)).

The adversarial surface: prompt injection, jailbreaks, and excessive agency

Every conversational assistant that can do something is also an attack surface, and the attacker is sometimes the customer. This is the section most conversational AI business cases omit, and it is the one that determines whether a deployment survives its first motivated user.

Direct and indirect injection are different problems

NIST’s adversarial machine learning taxonomy separates the two cases explicitly. Direct prompting attacks are those where the attacker submits the malicious input themselves. Indirect prompt injection arises when, in NIST’s formulation, untrusted external data is processed by the model (NIST AI 100–2e2025, March 2025 (opens in new tab)).

In banking, direct injection is the customer typing a manipulation into the chat window. Indirect injection is subtler and more dangerous: text that arrives from a document the customer uploaded, a previous ticket, an email thread pulled into context, or a merchant description in a transaction feed. The second category means an assistant can be attacked by someone who never speaks to it.

The OWASP Top 10 for LLM Applications ranks prompt injection first. The 2025 edition defines it as occurring “when user prompts alter the LLM’s behavior or output in unintended ways”, names sensitive information disclosure as the second risk with financial details explicitly in scope, and treats excessive agency as the risk created when a system is granted the ability to call functions or interface with other systems (OWASP, 2025 (opens in new tab)). The 2026 edition was published on 3 August 2026 (OWASP GenAI Security Project, 2026 (opens in new tab)). Coverage of that edition reports prompt injection still ranked first and excessive agency risen to third, which is the correct direction of travel for assistants that are being given tools (Invicti and Help Net Security, August 2026 (opens in new tab)).

The concession attack

The specific banking threat is not the assistant leaking a system prompt. It is a customer talking the assistant into a concession: a waived late fee, a reversed penalty, a rate reduction, a written-off balance, a promise of forbearance, or a confirmation that a payment obligation has been cancelled.

Outside financial services, this has already happened in public. In December 2023 a user manipulated the assistant on a Chevrolet dealership website into offering a vehicle for one dollar, with the bot adding that this was “a legally binding offer, no takesies backsies” (AI Incident Database, incident 622; Gizmodo, 2023 (opens in new tab)). In January 2024 DPD’s assistant swore at a customer and wrote a poem calling its own employer the worst delivery firm in the world after the customer simply asked it to disregard its guidelines; the company disabled the AI element the same day (ITV News, 2024 (opens in new tab)).

In banking, no equivalent public incident has been documented. That absence should be read as a warning about disclosure rather than as evidence of safety, because controlled testing points the other way. In a benchmark that includes a banking domain, automated prompt injections succeeded in making the agent call an attacker’s target function with the correct arguments in 45.2% of attempts against a small open model and 4.7% against a frontier model, across 80 task pairs (ETH Zurich, arXiv:2606.10525, June 2026 (opens in new tab)). A separate practitioner test of 24 commercial models configured as banking customer-service assistants reported exploit success rates ranging from 1% to over 64% depending on attack category, with assistants disclosing creditworthiness scoring logic including the relative weights of payment history, utilisation and account mix, and exhibiting a pattern of refusing and then complying in the same reply (Corporate Compliance Insights, January 2026 (opens in new tab)).

Model vendors are candid that the residual risk is not zero. Anthropic reported roughly a 1% attack success rate for one model against an adaptive attacker given 100 attempts per environment in browser use, and stated that “a 1% attack success rate, while a significant improvement, still represents meaningful risk” and that “no browser agent is immune to prompt injection” (Anthropic, November 2025 (opens in new tab)).

The answer is architectural, not conversational

A bank cannot prompt its way out of this. The guardrail that holds is the one that sits outside the model, and the design rule is that the model may propose and must never authorise.

In practice that means five things. Authority to grant lives in a deterministic policy building block, not in the model’s instructions, so a fee waiver requires a rule to evaluate true rather than a sentence to be persuasive. The assistant’s tool scope excludes anything that prices, waives, or forgives, so the capability is absent rather than restrained. Entitlements are checked server side against the authenticated customer, so a successful jailbreak still cannot reach another customer’s loan. Monetary effects carry hard limits and dual control above a threshold. And every action the assistant initiates is written to the same audit trail as a manual one, with the prompt, the retrieved context, and the tool call preserved, so a disputed answer can be reconstructed rather than argued about.

The framing published with the 2026 OWASP list is the right design brief for a bank: build the system around the model so that when the model is fooled, and it will be, nothing important breaks (as reported by Help Net Security, August 2026).

Diagram showing a banking assistant proposing an action while authority to grant sits in a deterministic policy block outside the model

Keeping personal data out of the context window

The second omission in most conversational AI business cases is that a chat channel is a personal data processing operation and, if a customer types a card number into it, a cardholder data environment.

What the rules actually require

GDPR Article 5(1)© requires personal data to be “adequate, relevant and limited to what is necessary in relation to the purposes for which they are processed”. Article 25(1) requires appropriate technical and organisational measures, “such as pseudonymisation”, implemented “both at the time of the determination of the means for processing and at the time of the processing itself”. That second clause is the one that bears on architecture: the decision to send raw personal data to a model is itself the regulated moment, not just the processing that follows.

Two European positions narrow the room further. The EDPB’s Opinion 28/2024 holds that models trained on personal data “cannot, in all cases, be considered anonymous”, and that for a model to be treated as anonymous both the likelihood of extracting personal data from the model and the likelihood of obtaining such data from queries must be insignificant (EDPB, December 2024 (opens in new tab)). The second limb is the one that bites a customer-facing assistant, because queries are exactly what it processes all day. The EDPB’s Guidelines 01/2025 then confirm that pseudonymised data which could be attributed to a person using additional information remains information on an identifiable person (EDPB, January 2025 (opens in new tab)). Tokenising a customer’s name before the prompt reduces risk. It does not move the interaction outside the regulation.

France’s CNIL, in recommendations published in February 2025, made the corresponding practical point for LLMs specifically: prompts, not only training data, are in scope, and developers should seek to prevent disclosure of confidential data by design (CNIL, 2025 (opens in new tab)).

Card data in a chat transcript

PCI DSS v4.0.1 (June 2024) is the operative version, and its future-dated requirements became mandatory on 31 March 2025 (PCI Security Standards Council, 2024 (opens in new tab)). Three points matter for a conversational channel. Requirement 3.4.1 limits PAN display to the BIN and last four digits for anyone without a documented business need. Requirement 3.5.1 requires PAN to be rendered unreadable wherever it is stored, by hashing, truncation, tokenisation, or strong cryptography. Requirement 3.3.1 prohibits retaining sensitive authentication data after authorisation at all.

The consequence is easy to miss: a customer who pastes a full card number into a chat window has just written cardholder data into the transcript store, the analytics pipeline, and possibly the model provider’s logs. A conversational channel therefore needs inbound detection and redaction before persistence, not a policy asking customers not to do it.

The de-identification layer in practice

One bank has published the pattern in specifics. Wells Fargo describes an orchestration layer that sits between the customer and the model, with the model performing intent and entity detection only and all computation and detokenisation remaining inside the bank. In the words of Chintan Mehta, who leads its digital technology and innovation function: “The orchestration layer talks to the model. We’re the filters in front and behind.” He adds that the bank’s APIs do not pass through the model at all (VentureBeat, April 2025 (opens in new tab)).

“The orchestration layer talks to the model. We’re the filters in front and behind.”

— Chintan Mehta, Head of Digital Technology and Innovation, Wells Fargo (as reported by VentureBeat, April 2025)

The alternative posture is guardrail by testing rather than by architecture. Nubank’s engineering team describes simulating thousands of adversarial conversations against its agents to check that they will not leak personal data or step outside their financial mandate, for a base of 131 million customers (Nubank, March 2026 (opens in new tab)). Both approaches are defensible. Only the first one is verifiable by inspection, which matters when a supervisor asks what prevents the failure rather than what detects it.

A workable reference design has four elements: a classification and redaction proxy that removes or tokenises identifiers before the context window is assembled; retrieval scoped to the authenticated customer’s own records; a model contract that excludes prompt data from training and constrains retention; and transcript storage treated as regulated data with its own retention clock. None of this is exotic, and all of it has to exist before the first customer conversation, not after the first incident.

Reference design for a de-identification proxy that tokenises personal data before the LLM context window is assembled

Why conversational AI stalls in banking

The failure pattern is well documented, and it is not a language problem. Gartner research analysing 432 AI use cases in customer service found that 25% produced positive ROI, 25% produced negative ROI, 42% were unclear, and 11% broke even (Gartner research reported by CX Dive, 2026 (opens in new tab)). Gartner also predicts that by 2027, half of the organisations that expected to significantly reduce their customer service workforce will abandon those plans, with 95% of surveyed leaders planning to retain human agents (Gartner, 2025 (opens in new tab)).

The cost curve is not moving the way the 2022 business cases assumed either.

“Full automation will be prohibitively expensive for most organizations; instead, leading organizations will use AI to drive customer engagement rather than to cut costs.”

— Patrick Quinlan, Senior Director Analyst, Gartner (Gartner, January 2026 (opens in new tab))

Gartner’s January 2026 forecast puts generative AI cost per resolution above $3 by 2030, exceeding offshore human agent costs, and expects assisted-service volume to rise 30% by 2028 because of regulatory change (Gartner, 2026).

Two institutional cases make the mechanism concrete. Commonwealth Bank of Australia cut 45 direct banking roles in mid-2025, citing a chatbot that had diverted a reported 2,000 calls a week (ACS Information Age, 2025 (opens in new tab)). In August 2025 it reversed the redundancies, apologised, said it “did not adequately consider all relevant business considerations”, and acknowledged that call volumes had risen after the assistant was introduced (ABC News, 2025 (opens in new tab)). Klarna’s arc is the same story at a different scale: a February 2024 claim that its assistant handled two-thirds of chats and did the work of 700 agents for a $40 million profit improvement (Klarna, 2024 (opens in new tab)), a May 2025 reversal in which the CEO said “what you end up having is lower quality” and began rehiring human agents (Sebastian Siemiatkowski, via CX Dive, 2025 (opens in new tab)), and Q2 2026 accounts showing customer service and operations costs of $57 million, up 18% year on year (Klarna Group plc, Q2 2026 (opens in new tab)).

McKinsey quantified why. Tools and vendors in bank customer care promise 30 to 45% cost reductions; a $10 billion-asset credit union that had identified 37% of potential annual savings achieved a 10% cost reduction in its first six months. The diagnostic is the useful part: 75% of common call reasons had no self-service option at all, and for auto loan payment calls, only 20% were automatable by AI without significant risk while 80% required process redesign first (McKinsey, 2026 (opens in new tab)). Properly implemented, McKinsey puts the realistic range at 25 to 40% lower call volumes and 15 to 25% better first-call resolution (McKinsey, 2026).

The pattern across all four sources is the same. A conversational layer bolted onto an unchanged servicing process relocates work. Removing work requires changing what the systems behind the conversation are able to do.

The SaaS lending ceiling

On a configurable SaaS lending platform, the assistant can see and do exactly what the vendor exposes. Reading balances and statuses is usually available. Creating a hardship arrangement, issuing a payoff quote under a bespoke fee policy, or restructuring a schedule generally is not, because the underlying action is not exposed as a callable operation with the institution’s own rules attached.

That converts every new intent into a roadmap request. The bank ends up with an assistant whose vocabulary is limited by another company’s release cycle, and the residual calls are precisely the ones that carry cost.

The pure-LLM shortcut

Connecting a language model to a document store produces fluent answers quickly and creates every exposure described in what a conversational layer must not answer (opens in new tab)the adversarial surface, and personal data and the context window (opens in new tab) at once. The model has no view of the loan record, no permission model, and no audit trail, so it can neither complete a transaction nor evidence what it told the customer.

The custom build cost

Building the servicing and origination logic from scratch so that an assistant has something safe to call is a genuine option and a slow one. As a working model for a full lending system, an in-house build runs to roughly 24 months, 16.5 people, and about $8 million before the first conversational integration is even scoped.

The programmable Building Platform

The third path is a lending system whose operations are building blocks the institution controls: entities, state machines, services, and integrations, configurable in an admin panel and extensible at code level through an SDK. The conversational surface then calls the same building blocks a loan officer uses, under the same policies, with the same audit entries.

That is what makes an assistant’s answer defensible: it is not a generated sentence about the loan, it is the result of the same operation the bank would have run manually. Compliance behaviour is inherited from explicit building blocks per jurisdiction rather than reimplemented inside a chat integration.

Three-way comparison of configurable SaaS, in-house build and programmable Building Platform for conversational AI capability in lending
CriterionConfigurable SaaS platformIn-house custom buildProgrammable Building Platform
What the assistant can readVendor-exposed fieldsEverything, after the buildThe full loan record and portfolio data
What the assistant can doVendor-exposed operations onlyWhatever was builtAny building block, under policy gates
Adding a new intent that changes the loanVendor roadmap request, 6–12 monthsNew development cycleConfiguration, in weeks
Where the authority to grant livesVendor-defined configurationWherever the team put itDeterministic policy blocks outside the model
Reason codes for a declined applicationVendor black boxBuilt from scratchDeterministic, explainable decisioning engine
Audit trail for what the assistant told a customerDepends on vendor loggingBuilt from scratchVersioned per action: who, what, when, why
Deployment and data residencyMulti-tenant cloudSelf-hostedSelf-hosted or private cloud
Time to first working systemWeeks, within vendor limits18–24 monthsWorking system from day one via admin panel

The constraint nobody markets: the core banking system

It is convenient to blame the SaaS vendor. Often the binding constraint sits one layer deeper, in the institution’s own core. An assistant that promises real-time answers is making a promise on behalf of a system that may not work in real time.

Batch is not a detail, it is the answer the customer gets

US bank supervisors define the problem in their own language. A ledger balance, in the FDIC’s formulation, “calculates the account balance based only on transactions settled during the relevant period and does not take into account authorization holds”, while an available balance also counts authorised but unsettled items, which is the mechanism behind authorise-positive, settle-negative transactions (FDIC (opens in new tab)). An assistant reading one of those two numbers and calling it “your balance” is not wrong by accident. It is wrong by architecture.

The consequences are measurable when the core degrades. During Barclays’ outage from 31 January to 2 February 2025, 56% of online payments failed because of severe degradation of the bank’s mainframe processing performance, disclosed to the UK Treasury Committee; the bank expected to pay between £5 million and £7.5 million in compensation for that incident, and up to £12.5 million across its outages. The same committee found at least 803 hours, more than 33 days, of unplanned outages across nine major UK banks and building societies between January 2023 and February 2025, across at least 158 incidents (House of Commons Treasury Committee, March 2025 (opens in new tab)).

The COBOL question is real but poorly evidenced. The most-cited estimate of the installed base, more than 800 billion lines in production, comes from a 2022 vendor-commissioned survey, which itself notes that earlier market estimates sat in the 200 to 300 billion range (Micro Focus and Vanson Bourne, 2022 (opens in new tab)). Treat the number as an order of magnitude from an interested party rather than as a measurement.

How batch settlement in a core banking system limits what a conversational assistant can promise a customer in real time

Replacing the core is not the fast answer either

The instinct to fix this by replacing the core does not survive the evidence. Reviewing more than 50 core banking transformations over seven years, McKinsey found that only about 30% succeeded in fully migrating ledgers and products to the new system, that banks overspent by 100% and timelines grew by 50 to 100%, that stakeholders underestimated migration time by as much as 75%, and that in failed cases legacy applications were still running at 10 to 20% of functionality more than five years after go-live (McKinsey, 2022 (opens in new tab)).

Banks are behaving accordingly. In the American Bankers Association’s 2025 core platform survey, only 19% of US bankers planned to convert their core at the next renewal date and 69% were extremely or somewhat likely to stay with their current provider (ABA, February 2025 (opens in new tab)). The realistic planning assumption for a conversational programme is therefore that the core stays where it is.

The abstraction layer, and what it can honestly promise

The standard mitigation is to put a layer between the assistant and the core. Gartner’s taxonomy of core modernisation strategies names this Abstraction, which isolates and simplifies the core to deliver shorter-term digital results, alongside Precision, Emigration and Undercover approaches (Gartner, March 2024 (opens in new tab)).

For a conversational programme the layer needs four properties. It exposes lending operations as idempotent, versioned APIs rather than screen-scraped calls, so a retried request cannot double-post. It maintains a read model that is explicit about freshness, so the assistant can say what is settled and what is pending instead of implying both are the same. It queues writes the core can only accept in batch and tells the customer the truth about timing. And it holds the policy and entitlement checks from keeping authority outside the model, because those cannot live in a core that has no concept of a conversational channel.

The discipline that follows is simple to state and hard to hold: an assistant may only promise what the core can confirm. Everything else is a request with a status, not an outcome.

The handoff is the product

Every serious analysis of conversational AI concludes that customers must be able to reach a person. Very few specify what has to travel with them. The escalation is where a contained conversation either becomes a resolved problem or becomes the second call that erases the saving.

What a handoff must carry

A chat log is not context. Passing a transcript to an agent transfers the reading work rather than the understanding, and it leaves the customer to re-establish who they are and what they wanted. The handoff payload should be a structured case, assembled by the assistant and reviewed by the agent in seconds.

Seven elements make it usable:

  1. Authenticated identity and authentication state. Which customer, verified to what level, and by what method, so the agent does not restart verification the customer has already passed.
  2. The resolved intent, with the alternatives considered. Not “customer asked about payment” but “customer requesting payment date change on loan X; hardship not asserted; forbearance eligibility not checked”.
  3. Extracted entities, mapped to system fields. Account, loan, amount, date, reason code, each as a value the agent’s screen can consume.
  4. Actions already taken or attempted. Every tool call, its result, and every policy check that failed, so the agent does not repeat a rejected operation.
  5. Why the assistant stopped. Low confidence, a policy gate, an out-of-scope intent, a detected vulnerability signal, or an explicit customer request for a human. These lead to different agent behaviour.
  6. Vulnerability and conduct signals, as flags rather than conclusions. Bereavement, illness, financial difficulty, and third-party involvement each change the correct handling.
  7. A link to the verbatim transcript and the retrieved context. For the agent’s reference and for the audit file.

This is not only a service-quality argument. The FCA’s good-practice examples for the consumer support outcome are specifically about escalation design: firms that “complemented digital channels with human touchpoints wherever complex needs were identified, for example, automatically directing queries from a webchat chatbot about bereavement to a customer support representative”, and that programmed “keywords relating to vulnerability to trigger a flag to customer support representatives, and ensuring these customers were swiftly moved into a high priority queue” (FCA, March 2025, updated July 2026 (opens in new tab)). A handoff design is a supervisory artefact, not an implementation detail.

There is now a vendor-neutral specification for passing a conversation between agents: Open Floor 1.0, released in May 2025 by the Open Voice Interoperability Initiative under LF AI and Data, which lets agents discover one another and take turns without pre-configured relationships (LF AI and Data, 2025 (opens in new tab)). Institutions building multi-assistant estates should watch it. It does not remove the obligation to define the payload above.

Structured handoff card passed from a banking assistant to a human agent: identity, authentication state, intent, entities, actions attempted, stop reason, vulnerability flags

Do not pass a sentiment score you cannot defend

Emotional context is the element most often promised and least often measured. Recent work on fine-grained speech sentiment puts hard numbers on it: on an eight-class taxonomy, the best audio-capable multimodal model reached 56.95% accuracy, speech encoders reached 44 to 45%, and text-only models collapsed to 25.26% and 14.20% because they discard prosody, tone, pauses, stress and tempo. Human annotators themselves reached only moderate agreement, with Fleiss’ kappa of 0.4437 (Chinese University of Hong Kong, arXiv:2608.17931, August 2026 (opens in new tab)).

Two conclusions follow. A sentiment label attached to a handoff is wrong roughly half the time, so it should be presented as a weak signal rather than as an assessment. And a pipeline that transcribes speech and then analyses the text has thrown away most of the emotional information before the analysis begins, which is exactly the architecture most voice deployments use.

The defensible version is to pass observable facts instead of inferred states: the customer used the word “bereavement”, the customer repeated the same question three times, the customer asked for a human twice, the call has run eleven minutes. An agent can act on those. A confidence-weighted emotion class invites the agent to trust a number the model cannot support.

The exit must be one step, and it is already an expectation

No jurisdiction currently grants a general right to a human interlocutor. Article 50 of the EU AI Act requires disclosure, not human access. Article 86 grants a right to an explanation of an automated decision, which is not the same thing. In the US, the CFPB warned about this in 2023 and did not rule, and the 2024 initiative that proposed a single-button route to a human never became a regulation.

What exists instead is an outcome test. The FCA’s position is explicit: “While we do not prescribe which channels firms must offer, they must ensure the channels of support they do offer meet the needs of their customers” (FCA, 2025). The HKMA goes further for generative AI specifically, requiring that customers be given the option to request human intervention (HKMA, 2024).

The consumer evidence explains why supervisors keep returning to this. The CFPB documented what it called hindered access to timely human intervention, quoting consumers describing “loop after loop of the same questions, all of them redirecting me” and a virtual assistant that “kept sending me in circles” (CFPB, 2023). The report also cites a commercial survey, not CFPB data, finding that 80% of consumers who interacted with a chatbot left more frustrated and 78% needed to reach a human afterwards; the attribution matters, and it is frequently misreported. The FCA’s own Financial Lives research found that in 19% of recent contacts with a financial services provider, consumers found it very or fairly difficult even to find the right contact information (FCA, 2025). Gartner found 64% of customers would prefer companies did not use AI in customer service and 53% would consider switching to a competitor over it, with difficulty reaching a person the single most cited concern (Gartner, 2024 (opens in new tab)).

Voice adds latency, accents, and noise to every other problem

Voice is not chat with a microphone. It introduces four constraints that do not exist in text, and each one has published measurement behind it.

The latency budget is already spent

Human conversation runs on a tight clock. The modal gap at speaker transition is about 200 milliseconds, and 51 to 55% of all turn transitions occur under 200 milliseconds (Levinson and Torreira, Frontiers in Psychology, 2015 (opens in new tab)). The telephony standard for interactive delay, ITU-T G.114 (opens in new tab), treats one-way delay below 150 milliseconds as essentially transparent and delay above 400 milliseconds as unacceptable for general network planning, though it governs transmission delay rather than AI response time.

Measured pipelines are nowhere near that. A vendor-neutral cascaded speech-to-text, model, text-to-speech pipeline documented by Salesforce AI Research reached a median 947 milliseconds to first audio using cloud APIs, and 729 milliseconds in the best self-hosted configuration, with component medians of 337 to 509 milliseconds for transcription, 337 milliseconds to first token, and 219 to 236 milliseconds to first audio byte (arXiv:2603.05413, March 2026 (opens in new tab)). Purpose-built full-duplex systems fare worse on first substantive response, with one reporting a median of 2.34 seconds and a 95th percentile of 3.02 seconds, even though barge-in handling can be fast (arXiv:2606.19453, June 2026 (opens in new tab)).

A caveat worth keeping honest: in text chat, longer waits can read as deliberation. A CHI 2026 study with 240 participants found a nine-second time to first token rated more useful and more thoughtful than two seconds (Tan et al., CHI 2026 (opens in new tab)). That finding does not transfer to voice, where silence is socially loaded and a two-second gap reads as a dropped call.

Voice AI latency budget compared with 200ms human turn-taking and the ITU-T G.114 400ms threshold

Accents, dialects, and who gets misheard

Speech recognition does not fail uniformly, and in a regulated channel that is a conduct problem rather than a user-experience problem. Across five commercial systems, average word error rate was 0.35 for Black speakers against 0.19 for white speakers, and 23% of audio from Black speakers produced unusable transcripts with error rates above 0.5, against 1.6% for white speakers (Koenecke et al., PNAS, 2020 (opens in new tab)).

The disparity has not closed. A 2026 benchmark of nine open-weight models found one leading system scoring 19.0% word error rate on Indian-accented English against 3.6% on Canadian, a ratio of 5.34, while the most equitable competitive model still scored 2.8% for white speakers against 8.5% for Black speakers; injecting silence amplified the accent gap by up to 4.64 times (arXiv:2604.21276, April 2026 (opens in new tab)). Even between native varieties the effect is measurable, with significantly higher match error rates for British and Australian speakers than American, and higher rates for speakers of tone languages than stress-accent languages (JASA Express Letters, 2024 (opens in new tab)).

For a lender, the implication is direct. If containment is higher for some accent groups than others, the bank has built a service channel that performs differently by customer demographic, and it will need to be able to show it monitors that.

Word error rate disparities in automatic speech recognition by speaker group and accent, 2020 and 2026 studies

Noise, and the values that matter most

Real-world audio is not benchmark audio. On distant, multi-talker, real-room speech, the CHiME-8 challenge baselines scored macro-average error rates of 56.5% and 62.6% (CHiME Workshop, 2024 (opens in new tab)). Independent measurement of eleven commercial services on real lecture audio found a 7.0% average error rate but a range from 0% to 53.8% across individual recordings, and found streaming recognition measurably worse than batch, which is the mode live voice AI necessarily uses (ACM Transactions on Accessible Computing, 2024 (opens in new tab)).

The values a lending conversation depends on most are the ones with the least published evidence. We could find no vendor-neutral benchmark for transcription accuracy on digits, account numbers, sort codes, or monetary amounts; the nearest available research on entity correction explicitly excludes numeric strings. In the absence of measurement, the design rule writes itself: never accept an account number, an amount, or a date by voice without deterministic validation and an explicit read-back confirmation, and never let a voice-captured value initiate an irreversible action on a single pass.

Voice authentication is the wrong place to save a call

The temptation to shorten the call by trusting the voice itself has a documented failure history. In February 2023 a journalist cloned his own voice from roughly five minutes of audio and used it to pass Lloyds Bank’s Voice ID on the automated line, reaching balances and recent transactions (VICE, 2023 (opens in new tab)). The following month Guardian reporters confirmed that the voiceprint system used by Centrelink and the Australian Taxation Office could be fooled the same way (The Guardian, 2023 (opens in new tab)). In May 2023 the chair of the US Senate Banking Committee wrote to six major financial institutions citing those tests (US Senate Committee on Banking, Housing and Urban Affairs, 2023 (opens in new tab)).

The fraud picture around this is worse evidenced than the vendor market suggests. Cifas recorded more than 444,000 cases filed to the UK National Fraud Database in 2025, up 6% year on year, with facility takeover at 78,000 cases and unauthorised SIM swaps up 38%, and attributed part of the trend to generative technologies enabling convincing impersonation at speed and scale (Cifas, March 2026 (opens in new tab)). No regulator or industry body publishes a figure isolating voice deepfake fraud in banking, and the circulating percentages that purport to do so are unsourced. Voice should be treated as an identifier of convenience, layered behind possession and knowledge factors, not as an authenticator.

A decision framework: what to put in the assistant first

Six questions decide whether an intent belongs in a conversational surface, and in what mode.

  1. Is the answer derivable from the record? If it requires interpretation rather than computation, it is not a self-service intent yet.
  2. Can the core confirm it in real time? If the underlying system settles in batch, the assistant can report a request and a status, not an outcome (the core banking constraint).
  3. Does the intent carry authority to grant anything? If yes, the authority belongs in a deterministic policy block and outside the model’s reach (keeping authority outside the model).
  4. Is the action reversible and auditable? If it cannot be reversed and evidenced, it needs a human gate regardless of model quality.
  5. Does the answer carry a statutory duty? Adverse action reasons, dispute rights, and hardship outcomes carry duties that sit with the institution, not the interface.
  6. Can the customer reach a person in one step, with context? Design the exit and its payload before the flow (the handoff).
Intent triage matrix showing which lending intents are full self-service, policy-gated, human-only, or never model-authorised
IntentAssistant modeWhat it needs from the architecture
Balance, next payment, payoff figureFull self-serviceRead access to the loan record and schedule; explicit settled versus pending state
Application status and missing documentsFull self-serviceRead access to origination state
Payment date change, direct debit updateSelf-service with policy gateCallable servicing operation, policy block, idempotent write, audit entry
Fee waiver or rate concessionNever model-authorisedDeterministic eligibility rule, tool scope excluding the capability, limits and dual control
Hardship or forbearance requestAssisted intake, human decisionStructured intake, case creation, vulnerability flag, priority routing
Fee or interest disputeAssisted intake, human decisionCase creation, evidence capture, statutory clock
Credit decline reasonsRead and present onlyReason codes from the explainable decisioning engine
Account number or amount captured by voiceConfirm before actingDeterministic validation plus read-back; no single-pass irreversible action
Pricing or terms not in the recordNot an assistant intentReferral to a person

Frequently Asked Questions

What is conversational AI in banking?

Conversational AI in banking is the layer that interacts with customers and staff in natural language across chat, voice, and in-app surfaces. It interprets intent, holds session context, retrieves information from bank systems, and either completes a permitted action or hands the case to a person.

How is conversational AI different from generative AI and agentic AI in banking?

Conversational AI talks, generative AI produces, and agentic AI acts. A conversational assistant is an interface; generative models draft and summarise content; agentic systems plan and execute multi-step work across tools. A single deployment often uses all three, and their governance requirements differ.

Can a customer manipulate a banking chatbot into granting a fee waiver or lower rate?

Prompt injection ranks first in the OWASP Top 10 for LLM Applications, and controlled tests show agents being driven into unintended tool calls at meaningful rates. The mitigation is architectural: authority to waive or reprice must sit in a deterministic policy rule outside the model, and the capability must be absent from the assistant’s tool scope.

Can a banking chatbot decline a loan application or explain the reasons?

No. Under Regulation B, 12 CFR 1002.9(b)(2), adverse action reasons must be specific and identify the principal reasons, so they must come from an explainable decisioning engine. A conversational assistant may present those reason codes to the applicant, but must never generate or paraphrase them.

Do banks have to tell customers they are talking to an AI?

In the EU, yes. Since 2 August 2026, Article 50 of the EU AI Act requires that people be informed they are interacting with an AI system, from the start of the first interaction and in a clear, distinguishable manner. Breaches carry penalties of up to €15 million or 3% of global turnover.

Why do most conversational AI projects in banking fail to cut cost-to-serve?

Because containment is measured instead of resolution, and because the core often cannot support real-time action. McKinsey found 75% of common call reasons had no self-service option and 80% of one call type required process redesign before safe automation. An assistant over an unchanged process relocates work rather than removing it.

Appendix: how TIMVERO implements this

The analysis above is architectural and vendor-neutral. This appendix states how TIMVERO applies it, for readers who want the concrete implementation rather than the principle.

timveroOS is a lending solution built on a Building Platform. Three things are frequently collapsed into the phrase “the AI”, and only two of them are ours. Keeping all three separate is what makes each one approvable.

The channel talks, and the channel is not ours. TIMVERO does not ship a customer-facing chat or voice assistant, and nothing in this analysis should be read as a claim that it does. The conversational surface, including any speech input, belongs to the portal our clients operate for their own borrowers. So does everything that comes with it: the latency budget, transcription accuracy, read-back confirmation, and the AI disclosure the EU now requires. What timveroOS contributes is the layer that surface calls: each origination and servicing action exposed as a governed operation with its policy gate, entitlement check and audit entry, so that whatever the portal puts in front of a borrower is calling something defensible. No operation exposed this way carries authority to price, waive or forgive.

The XAI scoring engine decides. It runs at origination and servicing decision points, produces explainable reason codes, and is the only component permitted to originate a credit decision or its reasons.

timveroAI (opens in new tab) builds. It is a RAG-grounded implementation agent that configures timveroOS: it drafts specifications, configurations, and code during build and maintenance, and it automates more than 80% of implementation work while engineers and business owners keep the logic. It does not make credit decisions and it does not speak to customers.

Three implementation properties follow from the architecture. What the channel can be grounded on is the loan record rather than a folder of documents, so an answer about a customer’s schedule is computed from the schedule rather than retrieved from a policy PDF. Changes proposed by AI pass explicit gates: shadow-run mode runs new logic beside existing logic and compares outcomes without acting, changes are versioned with who, what, when and why, and human sign-off is required before production. And because operations are building blocks, a new conversational intent that changes a loan is a configuration exercise rather than a roadmap negotiation.

On evidence: timveroOS runs 20 lending products across more than 13 countries and has been live in production since 2024. Implementation follows a consistent pattern of roughly two months for a first product, two to four weeks for subsequent ones, and changes in days, which is about 10x faster than a conventional build cycle. AMIO Bank reached a working MVP in four months after previous attempts had failed, with an 8x reduction in time-to-yes. Finom launched banking-grade lending across five European markets in four months with 98% process automation.

“What impressed me most was their ability to work at our pace, absorbing requirements on the fly, proposing solutions proactively, and adapting as our needs evolved. Today, we’re running proactive credit campaigns and sophisticated servicing operations on a single platform. timveroOS delivered a competitive advantage under impossible deadlines.”

— Alex Goncharenko, Head of Credit, Finom

The relevant capability set sits in lending software for banks (opens in new tab) and the underlying loan management software (opens in new tab), with portfolio signals in AI loan portfolio analytics (opens in new tab).

Schedule an Architectural Review for Your Assistant

If your conversational roadmap is blocked on what the lending system will let an assistant do, the constraint is the architecture rather than the model. We will map your highest-volume intents against the operations timveroOS would expose, the policy gates each one needs, and what your core can confirm in real time.

Request a demo → (opens in new tab)