Most bank AI pilots do not fail on model accuracy. They stall because nobody planned for the controls that every production system in a bank must pass: model risk review, data access approval, change and release management, third-party risk under DORA and, for some uses, the EU AI Act. The fix is to design for those controls from the first sprint, not after the demo.

This guide is for CTOs, heads of engineering, data and AI, and the risk and model governance leads who sign off on what goes live at EU banks, insurers and fintech companies. It sets out why pilots stall, which rules apply, and a five-step path to production. If you are already talking to partners, keep our questions to ask a fintech AI or software partner open alongside it.

Why do AI pilots stall in banks?

AI pilots stall in banks because they are built outside the bank’s control framework and then have to be rebuilt to enter it. A pilot proves a model can work. Production needs evidence that it is governed: approved data access, a validated model, a controlled release, a monitored service and a contract that survives a DORA review.

The usual blockers are predictable:

  • Model risk management. The model has no owner, no documented methodology and no independent validation, so the model risk function cannot approve it.
  • Data access. The pilot ran on an extract someone obtained informally. Production needs a lawful basis, approved access, masking and lineage.
  • Change and release controls. The model was deployed from a notebook. The bank’s change process expects tested, approved and recorded releases.
  • Third-party and ICT risk. The pilot used an external service or partner with no contract terms that meet DORA.
  • EU AI Act scope. Nobody checked whether the use case is high-risk, so legal review stops it late.

None of these is a reason not to use AI. Each is a design input you can meet early.

Which rules apply to an AI use case in a bank?

Four layers usually apply: DORA for ICT risk, change management and third parties; the EU AI Act for high-risk uses such as creditworthiness scoring; EBA and ECB expectations for models used in credit decisions and internal models; and GDPR for personal data. Which ones bite depends on the use case. This is general guidance, not legal advice.

DORA: ICT risk, change management and third parties

DORA, Regulation (EU) 2022/2554, has applied since 17 January 2025. An AI service is an ICT system like any other, so three provisions matter most:

  • Change management (Article 9(4)(e)). Financial entities need documented policies, procedures and controls so that changes to ICT systems, including software, are recorded, tested, assessed, approved, implemented and verified in a controlled way. A retrained model going live is a change.
  • Full responsibility (Article 28(1)(a)). The financial entity remains fully responsible for compliance when it uses ICT third-party services.
  • Contract terms (Article 30). Contracts must cover, among other things, service descriptions, data processing locations, data access and return, incident assistance, cooperation with authorities and termination rights. Services that support critical or important functions need more, including audit and access rights and exit strategies.

EU AI Act: high-risk uses and the current timeline

Under the EU AI Act, Regulation (EU) 2024/1689, Annex III lists as high-risk AI systems used to evaluate the creditworthiness of natural persons or establish their credit score, except systems used to detect financial fraud (point 5(b)). It also lists systems used for risk assessment and pricing of natural persons in life and health insurance (point 5(c)).

For banks that use such systems as deployers, the obligations include:

  • using the system in line with the provider’s instructions;
  • assigning human oversight to people with the competence, training and authority to exercise it (Article 26);
  • keeping the system’s logs, which financial institutions keep as part of the documentation required under financial services law (Article 26(6));
  • performing a fundamental rights impact assessment before deploying a creditworthiness or insurance pricing system (Article 27).

The timeline has moved. The Digital Omnibus on AI, Regulation (EU) 2026/1744, was published in the Official Journal on 24 July 2026 and entered into force on 27 July 2026. It moved the date from which the high-risk requirements apply to Annex III systems from 2 August 2026 to 2 December 2027. The original prohibitions have applied since 2 February 2025. The Omnibus also reworded the AI literacy duty in Article 4: providers and deployers must now take measures to support the AI literacy of their staff.

Do not read the delay as a pause. A credit model you build in 2026 will very likely still be in use in December 2027, so build the oversight, logging and documentation in now.

EBA and ECB expectations for models

For credit decisions, the EBA Guidelines on loan origination and monitoring (EBA/GL/2020/06, paragraphs 54 and 55) expect institutions using automated models for creditworthiness assessment to understand those models. They also expect controls for bias and data quality, traceability and auditability of inputs and outputs, regular backtesting, override and escalation procedures, and adequate model documentation.

For capital models, the EBA’s 2023 follow-up report on machine learning for IRB models sets out recommendations for prudent use. The ECB guide to internal models has included a section on machine learning techniques since its July 2025 revision, covering governance, data, methodology, explainability and model use.

Step 1: triage use cases by risk before you build anything

Triage every AI idea before a line of code is written. Classify it by decision impact, data sensitivity, regulatory scope and dependency on third parties. That classification decides which controls apply, who must approve it and how long the path to production will take. Low-risk internal tools can move fast; credit decisions cannot.

A simple three-tier triage works for most banks:

Tier Typical examples What it triggers
Low Internal search, developer assistants, document summarisation with no customer decisions Standard ICT change and security controls, data classification check, AI literacy measures
Medium Fraud detection, AML alert prioritisation, operational forecasting, customer service routing Model risk review, documented validation, monitoring, human review of outcomes
High Creditworthiness assessment or credit scoring of individuals, life and health insurance pricing Everything above, plus AI Act high-risk deployer obligations, fundamental rights impact assessment, EBA loan origination expectations

Record the triage decision and its reasons. When scope changes, for example a fraud model that starts influencing credit limits, re-triage it.

Step 2: settle data access and controls early

Settle data access before modelling, not after. Agree with data owners and the DPO which data the use case may use, on what lawful basis, at what level of masking and in which environment. Most delays between pilot and production come from data that was never approved for the purpose it ended up serving.

Practical controls that keep this moving:

  • Purpose and lawful basis written down for each dataset, with the DPO’s sign-off for personal data.
  • Environment separation. Development and test use synthetic or masked data. Production data stays in production.
  • Least-privilege access for engineers and data scientists, granted per project and reviewed.
  • Lineage from source to feature to model, so you can show which data trained which model version.
  • A written position on model training. At Areusdev, client data is never used to train shared models without a written agreement. Ask the same of any vendor or partner.

For deeper partner questions on data residency, encryption and deletion, use our AI and software partner due diligence FAQ.

Step 3: build an MLOps pipeline that produces audit evidence

Treat the MLOps pipeline as part of your control framework. Every model version should be built from versioned code and data, tested automatically, approved by named people, released through the bank’s change process and recorded. If the pipeline produces that evidence as a by-product, audits stop being a document hunt.

A pipeline that holds up under bank controls usually includes:

  1. Version control for code, configuration, feature definitions and training data references.
  2. Reproducible training in a controlled environment, so a past model can be rebuilt.
  3. A model registry that stores each version with its metrics, validation status, owner and approval.
  4. Automated tests for data quality, model performance thresholds, bias checks and integration behaviour. See our approach to test automation and CI/CD for banking and insurance platforms.
  5. Approval gates that match your change policy: model owner, independent validation and change approval before promotion to production.
  6. Controlled deployment with rollback, so a new model version can be withdrawn quickly.
  7. An audit trail linking each production prediction service to the model version, data version and approvals behind it.

This is what DORA’s change management provision looks like in practice for models: changes recorded, tested, assessed, approved, implemented and verified.

Step 4: validate independently and monitor in production

Validation and monitoring are where model risk management meets engineering. Before go-live, someone independent of the build team should review the methodology, data, performance, explainability and limitations. After go-live, monitor data drift, performance, stability and overrides, with thresholds that trigger review or retraining through the same controlled pipeline.

What good looks like:

  • Independent validation with a written report, proportionate to the risk tier.
  • Explainability appropriate to the use: reason codes for credit decisions, feature importance for internal models.
  • Backtesting of outputs against outcomes, which the EBA expects for automated creditworthiness models.
  • Human oversight that is real: reviewers with the authority and the information to override, and records of overrides.
  • Monitoring with owners. Every alert has a named owner and a documented response.
  • Retraining as a change. A retrained model goes through validation and change approval again, scaled to how much it changed.

Step 5: get vendor and contract terms right

If a vendor or partner touches the AI system, the bank still owns the outcome. Map the arrangement to DORA Article 30 contract terms, add it to the register of information, and agree in writing who owns the models, code and data, where data is processed and how exit works. Do this before the pilot, so production is not blocked by procurement.

Points to settle in the contract:

  • IP and ownership of models, features, code, pipelines and documentation.
  • Data terms: processing locations, subprocessors, no training of shared models on your data without a written agreement, and return and deletion at exit.
  • Audit and access rights for you and your competent authority, especially for critical or important functions.
  • Incident assistance and cooperation with authorities.
  • Exit and transition with a handover pack: repositories, pipelines, model registry and runbooks in your accounts.

To shortlist partners on this basis, use our regulated software partner scorecard.

Where a regulated engineering partner fits

A regulated engineering partner adds delivery capacity inside your controls; it does not replace them. The bank keeps the models, the data, the risk decisions and the approvals. The partner’s engineers build the pipelines, integrations, tests and monitoring to your standards, in your environments, and hand over everything they build.

That is a different model from buying a black-box AI product. With a product, you inherit the vendor’s model, data practices and release cycle, and your validation team has to work around them. With an engineering partner, the model is yours, the evidence is generated in your pipeline, and your model risk function reviews it on your terms.

Areusdev has been building software since 2004 and has delivered for Dell Financial Services, a regulated financial services business. Our AI and automation engineering teams work as part of your delivery organisation, under your change, security and model governance processes.

Frequently asked questions

Why do most AI pilots in banks never reach production?

Most stall because they were built outside the bank’s control framework. Production needs approved data access, independent model validation, controlled releases, monitoring and DORA-compliant contracts, and pilots rarely plan for them. Rebuilding a pilot to meet those controls takes longer than building it right the first time, so design for the controls from the first sprint.

Is credit scoring a high-risk AI use under the EU AI Act?

Yes. Annex III of the AI Act lists AI systems used to evaluate the creditworthiness of natural persons or establish their credit score as high-risk, with an exception for systems used to detect financial fraud. Life and health insurance risk assessment and pricing for individuals is also listed. Deployers face obligations such as human oversight, log keeping and a fundamental rights impact assessment.

When do the EU AI Act high-risk rules apply to banks?

For Annex III systems such as credit scoring, the high-risk requirements apply from 2 December 2027. The Digital Omnibus on AI, Regulation (EU) 2026/1744, which entered into force on 27 July 2026, moved that date from 2 August 2026. The original prohibitions have applied since 2 February 2025. Models built now will likely still be running in 2027, so plan for the rules today.

How does DORA affect AI and machine learning systems?

AI systems are ICT systems under DORA. Model releases and retraining fall under ICT change management, which requires changes to be recorded, tested, assessed, approved, implemented and verified. If a third party provides any part of the AI service, the bank remains fully responsible, and the contract must include the Article 30 terms and an entry in the register of information.

What should an MLOps pipeline include to satisfy bank controls?

It should include version control for code and data references, reproducible training, a model registry with owners and approvals, automated tests for data quality, performance and bias, approval gates aligned with your change policy, controlled deployment with rollback, and an audit trail from each production service back to the model version, data and approvals behind it.

Should a bank build AI in house, buy a product, or use an engineering partner?

It depends on how core and how regulated the use case is. Products suit commodity tasks where the vendor’s model is acceptable. For credit, fraud and other governed models, many banks keep the model in house and add capacity through an engineering partner that works inside their controls, so validation, ownership and audit evidence stay with the bank.

Talk to an architect

Bring one AI use case that is stuck between pilot and production. We will map it against your controls and show what it takes to get it live.

Talk to an architect

Read the AI and software partner due diligence FAQ