Back to Insights

Blog

The agent that can't act alone: Agentic AI inside a regulated Financial Institution

August 10, 2026

Dedicatted Petlichenko

5 min to read

Read summarized version with

If you own, run, or sit on the board of a bank, insurer, pension fund, or payments company, you’ve probably already had the pitch: agentic AI can cut operational costs, compress cycle times from days to minutes, and free your best people from repetitive work so they can focus on judgment calls only humans can make. All of that is true. It’s also incomplete, because the same capability that makes agentic AI powerful: the ability to plan, act, and adapt with little or no human involvement is exactly what makes it a new category of risk that your existing controls were never built to catch.

Agentic AI is already inside production systems at the largest institutions in the world. BNY has agents validating payment instructions and writing code. JPMorgan Chase built an agentic legal-workflow system, LAW, that processes complex custody and fund-services contracts with 92.9% accuracy across various query types. Wells Fargo is building custom agents on Google’s Agentspace. According to recent industry research, mortgage processing – a workflow that can run 45 distinct compliance, document-review, and underwriting steps behind a single customer meeting is now considered one of the most ripe targets for agentic reinvention in the entire sector. One in three financial institutions is setting aside dedicated budget for agentic AI right now, and new roles: AI risk officer, behavior auditor, agent steward are appearing on org charts that didn’t exist two years ago.

This piece is about what governance infrastructure actually looks like, what it’s worth in hard numbers, and specific ideas your team can act on this quarter , including where a partner like Dedicatted fits into making that infrastructure real rather than theoretical.

Table comparing four AI categories (traditional RPA, classic AI, LLM copilots, Agentic AI) across primary paradigm, autonomy, learning, scope of use, decision-making, memory, explainability, and governance.

Three Agentic AI failure patterns you will likely recognize

Governance failures are structural and repeat in three specific, nameable patterns. None of them require malice or incompetence. All of them are preventable with the right infrastructure, and all of them get more expensive to fix the longer they run undetected.

Scope creep by a thousand reasonable decisions. A transaction or payments agent starts as decision support: it flags anomalies, a human acts. Confirmation requirements get relaxed because the system proved reliable. Autonomy thresholds get raised because the business case for speed is compelling. New APIs get connected because a new use case was approved. Each change is individually defensible. Eighteen months later, the system holds the action authority of a high-risk deployment while still being governed like the cautious pilot it started as and no single person or committee ever made the decision to reclassify it. The fix isn’t slowing down approvals; it’s building automatic re-classification triggers so that a material change in tool access, action authority, or data reach forces a fresh risk review by design, not by memory.

Accountability that was never assigned. In credit and risk-decisioning systems, product owns the model, engineering owns the infrastructure, compliance owns the policy and when a decision is challenged by a customer or regulator, every function can credibly point elsewhere. This is exactly why we recommend naming an agent owner, validator, and steward for every production system, with ownership that survives the original deployment team. The cost of skipping this step isn’t theoretical: 56% of financial-sector respondents in the Bank of Canada survey cite talent-related constraints as a barrier to AI expansion, and unclear ownership is one of the fastest ways to burn through the AI talent you already have, because nobody wants to inherit accountability for a system they didn’t design and can’t fully explain.

The vendor update nobody flagged. Once an agent routes through external APIs or counterparty systems, your institution’s risk perimeter extends into infrastructure you don’t control and often can’t observe. Consider the CrowdStrike outage of July 2024: a single third-party technology failure cost the Fortune 500, excluding Microsoft, an estimated $5.4 billion. That was a conventional software update gone wrong at one vendor, rippling through institutions that depended on it. This is the cautionary benchmark every procurement and vendor-risk team should have pinned above their desk before signing the next agentic AI contract.

Diagram showing four key agent sources connected to multiple risks such as Runaway agents, Emergent behavior, and Adversarial manipulation.

What good Agentic AI governance actually costs

Here’s the reframe worth bringing into your next budget conversation: strong agentic AI governance is not a tax on innovation. It’s the infrastructure that lets you deploy more autonomy, in more workflows, faster than competitors governing by instinct. An institution with a live agent registry, automatic re-classification triggers, and clean audit trails can approve the next expansion of an agent’s authority in days, because the review process is fast and the paper trail already exists. An institution without that infrastructure either says no to reasonable expansions out of caution, leaving real efficiency gains on the table or says yes without adequate review and quietly accumulates the exact risk described above. Neither path beats a competitor who’s built the infrastructure to say yes safely and quickly.

Six commitments show up consistently across the regulatory guidance, the engineering frameworks, and the institutions that have avoided the failure patterns above. Each one is an idea your team can turn into a work item this quarter.

  1. Give every agent a corporate identity. A distinct digital identity, scoped permissions, and full action logging – the same standard you’d apply to a new employee with system access. This is the control that makes every other control possible, because you cannot govern something you cannot uniquely identify and trace.
  2. Build narrow, modular agents instead of one broad one. Specialized sub-agents with tightly scoped permissions : the pattern already running in multi-agent AML and KYC workflows at several banks are inherently easier to explain, audit, and contain than a single generalist agent with sweeping access.
  3. Put circuit breakers inside the workflow, not just approval gates in front of it. For high-stakes use cases: large or unusual payments, credit decisions above a threshold, trading build explicit kill switches, dynamic limits, and mandatory human checkpoints triggered by the specific action, not only by the initial deployment approval.
  4. Make risk classification self-triggering. Canada’s OSFI has already written this into its updated Model Risk Management guideline (E-23), requiring risk ratings to be reassessed whenever a model’s use, data, or infrastructure changes materially.
  5. Name an owner who survives the deployment. Agent owner, validator, and steward roles, backed by cross-functional governance spanning risk, compliance, and cybersecurity.
  6. Extend your existing model risk function, don’t build a parallel one. Add level of autonomy as a formal rating factor to the model inventory and review cadence you already run. This is measurably cheaper than standing up a separate “AI governance” silo, and it keeps Agentic AI inside the accountability structure your board already understands.

The AGILE framework: Awareness, Guardrails, Innovation, Learning, Ecosystem Resiliency developed through a forum of more than 170 participants spanning banks, insurers, regulators, and consumer advocates, distills this into one line worth repeating in your next strategy session: “the biggest risk is not doing enough.” Inertia is not the safe choice. Ungoverned speed is not the safe choice either. The safe choice is speed with infrastructure underneath it.

A practical guide: what to do, why it matters, and how it plays out

Knowing the principles is one thing. Turning them into a sequence of actions your institution can actually execute with the right people doing the right things in the right order is another. Below is a phased roadmap, followed by what changes for each stakeholder group, followed by worked examples of how this looks in a real workflow.

Phase 1 : See what you actually have

Commission a full inventory of every AI system currently in production, pilot, or shadow use across the institution, including tools business units may have adopted without a formal review (a phenomenon sometimes called “shadow AI,” the agentic-era version of shadow IT). For each system, capture: what it’s authorized to do, what it can actually reach (data, APIs, systems), who owns it, and when it was last reviewed.

You cannot govern, price, or defend what you cannot list. This is also almost always the moment institutions discover the gap between “what we approved” and “what’s running” described earlier in this piece and it’s far cheaper to find that gap yourself than to have a regulator or auditor find it for you. OSFI’s E-23 guideline effectively makes this mandatory for federally regulated institutions: an “accurate, evergreen” model inventory subject to robust controls is now a stated expectation

Dedicatted Top-Tier AI Specialisy

Expect the inventory to surface more systems than leadership expects:often two to three times the number the CTO’s office originally estimated, once business-unit pilots and vendor-embedded AI features are counted. That’s normal, not alarming, provided it triggers phase 2 rather than getting filed away. CRO or Chief Compliance Officer sponsors it; CTO/Head of Engineering executes it; a small cross-functional team (risk, compliance, engineering, one business-unit representative) does the actual cataloguing. This is exactly the scoped, bounded engagement a partner like Dedicatted can accelerate, bringing AWS-certified engineering capacity to do the technical discovery quickly, without pulling your internal team off their current roadmap for a quarter.

Phase 2 : Assign ownership and build the re-classification trigger

For every system in the inventory, assign a named agent owner, validator, and steward, not a committee. Build (or extend your existing model risk platform to include) automatic triggers that force a fresh risk review whenever a system’s tool access, data reach, or action authority materially changes. Score every system on autonomy level, using a framework like the industry’s five-level scale: Tool, Assistant, Operator, Actor, Agent ,so risk and business leadership share a common vocabulary for “how much is this system allowed to do without a human.”

This is the single control that would have prevented every scope-creep failure described earlier in this piece. Institutions that build the trigger once spend far less on emergency remediation later: the RAI Institute’s review found that in every scope-creep case examined, no one had ever made the decision to expand the system’s authority; it happened by accumulation. A trigger converts accumulation into a decision, every time.

Expect friction here, because this phase asks business units to slow down slightly on the systems that are working well, in exchange for being able to move faster and with more confidence on the next one. Frame it that way internally: it’s the reason you’ll be able to approve the next agent expansion in days instead of months once the infrastructure exists.

Dedicatted Top-Tier AI Specialist

Joint ownership across risk, compliance, and engineering, with board-level reporting on coverage (what % of the inventory now has a named owner and an active trigger policy) as a standing metric: the same way you’d report on any other risk-remediation program.

Phase 3 (ongoing): Pilot, prove, then scale autonomy deliberately

Choose one narrow, high-value, lower-risk workflow, build it with guardrails and monitoring designed in from day one, run it in a constrained pilot, and use the results to justify the next expansion, rather than approving a broad rollout up front. Report results against both value delivered (cycle time, cost per case, error rate) and governance health (time-to-detect an anomaly, time-to-human-escalation, audit-trail completeness).

This is how every credible deployment cited throughout this piece actually scaled , not with a single “go live” moment, but with a proof point that earned the next increment of trust and authority. It’s also how you build the internal muscle memory and the evidence base your board, and eventually your regulator, will want to see.

How it affects your organization. Done well, this phase compounds. The governance infrastructure built for the first use case: agent identity, logging, re-classification triggers is largely reusable for the second, third, and tenth. The marginal cost of governing each additional agent drops sharply once the foundation exists, which is exactly why institutions that invest early tend to pull ahead on both speed and safety simultaneously.

Worked examples: how this looks in practice from Dedicatted

Anti-money-laundering investigation. A three-agent workflow : one agent reviews the triggering alert, one reviews current and historical transaction data, one drafts findings and a recommendation, runs end-to-end without human intervention until a human validates the final report before filing. Applying the framework above: each agent gets a distinct identity and scoped access (the alert-review agent doesn’t need write access to case files; the drafting agent doesn’t need direct database access at all), the system is scored at “Operator” level rather than “Agent” level because filing still requires human sign-off, and any request to expand it toward autonomous filing triggers a fresh review rather than a quiet policy change. The measurable payoff mirrors what the sector is already reporting: faster investigation cycle times and the ability to surface illicit patterns a purely manual review might miss, without moving the actual regulatory filing decision out of human hands.

Treasury cash-sweep optimization. An existing RPA process that moves routine cash sweeps gets elevated by an AI agent that also makes pricing and hedging recommendations: a concrete example already cited by industry research as a realistic evolution of legacy automation. Framework applied: the agent’s action authority is explicitly capped (it can recommend and execute within a pre-set limit; anything above the limit routes to a human), and that cap is exactly the kind of dynamic risk limit and circuit breaker described earlier. The business case is direct, more strategic liquidity management without adding headcount — and the governance case is equally direct: a single, clearly bounded action authority that’s easy to audit.

Credit decisioning. This is the workflow where the accountability gap is most dangerous, precisely because product, engineering, and compliance all have a legitimate claim to partial ownership. Applying the framework: a named agent owner sits above all three functions specifically for the credit-decisioning agent, explainability requirements are set before deployment (not retrofitted after a declined applicant complains), and the model’s risk rating explicitly accounts for its level of autonomy: consistent with OSFI’s E-23 requirement that qualitative factors like “model complexity or level of autonomy” feed directly into the risk rating. The payoff: faster, more consistent underwriting decisions that the institution can actually explain and defend, which is the difference between a fair-lending inquiry that resolves in an afternoon and one that becomes a multi-month investigation.

Where Dedicatted fits across all three examples. In each case above, the governance requirements (agent identity, scoped access, action-authority caps, audit trails) are not policy documents, they’re engineering decisions made at build time. That’s the layer where Dedicatted’s multi-agent AWS architectures and built-in guardrails and monitoring do the actual work, so that when your compliance team asks “can we prove this agent never had access to X,” the honest answer is yes, because the system was built that way from the start rather than patched to demonstrate it after the fact. Contact us to talk through where your environment stands today.

Ideas for your AI strategy you can act on this quarter

For an executive team looking for where to start, five moves consistently separate the institutions ahead of this curve from the ones playing catch-up:

  • Commission a full agent inventory now, before your next audit cycle forces the question. If you can’t currently list every agentic system in production with its current autonomy level and last review date, that’s the single highest-leverage gap to close first.
  • Set a re-classification trigger policy so that any material change to an agent’s tool access, data reach, or action authority automatically forces a review , mirroring what OSFI has already written into E-23 for federally regulated institutions.
  • Pilot one narrow, high-value, lower-risk use case with guardrails built in from day one: KYC maintenance, AML investigation support, or treasury cash-sweep optimization are the workflows already proven out at BNY, Intesa Sanpaolo, and others and use it to prove the governance model works before scaling autonomy further.
  • Bring in engineering capability that’s built for regulated environments from the start. Whether that’s Dedicatted or another partner with equivalent compliance-native experience, the point is the same: don’t ask your internal team to become agentic AI architecture experts and AI governance experts simultaneously under deadline pressure. Buy the expertise you don’t have time to build from scratch, and put your internal effort into ownership, oversight, and the decisions only your institution can make.

Where a partner like Dedicatted fits into that infrastructure

Everything above is a governance and architecture problem before it’s a compliance problem, which means the fastest, cheapest way to close the gap is usually not to build a parallel AI governance function from scratch internally, but to bring in engineering capability that already knows how to build agentic systems with the controls baked in from day one, on infrastructure your risk and compliance teams can actually audit.

This is precisely the space Dedicatted operates in. As a Canadian AWS Premier Tier Services Partner and top 2% global AWS partner with GenAI Competency and MSP designationwith more than 100 AWS-certified engineers and 20-plus AWS competencies across cloud, security, and DevOps, Dedicatted builds custom, multi-agent AI systems specifically for regulated, highly compliant environments: banking, capital markets, insurance, and payments among them. A few ways that capability maps directly onto the six commitments above:

  • Guardrails and monitoring built in at the architecture stage, not bolted on afterward. Dedicatted’s agentic AI engagements are explicitly designed with guardrails, monitoring mechanisms, and system controls so that multi-agent systems operate safely and consistently in production, which is the technical backbone that makes agent identity, action logging, and audit trails achievable in practice rather than aspirational in a policy document.
  • Multi-agent architectures on AWS that map naturally to modular, scoped-permission design. Rather than one broad system, Dedicatted’s approach coordinates specialized agents across tool-enabled workflows the same narrow, auditable structure that regulators and risk frameworks are converging on as best practice, built on infrastructure your team already trusts for other core systems.
  • Deep familiarity with your existing compliance surface. Dedicatted’s financial-sector work already integrates AI assistants into loan origination, fraud detection, and real-time reporting inside highly regulated payment, trading, and insurance environments, meaning the conversation with your compliance and risk teams starts from a system built for your regulatory context, not adapted from a generic template after the fact.
  • DevOps and cloud-native delivery that lets you move at the speed your competitors are moving, without trading away control. Cloud-native pipelines, automated security gates, and continuous testing mean new agentic capabilities can ship without forcing a choice between speed and governancen – the two are engineered together.

The practical starting point for most institutions isn’t a full agentic rollout. It’s an audit: map every AI system currently in production or pilot, score its level of autonomy and action authority, identify where ownership is unclear, and build the registry and re-classification triggers described above , then use that foundation to decide, deliberately, where the next expansion of agent authority should happen and where it shouldn’t yet. That’s a scoped, bounded engagement, not a multi-year transformation program, and it’s exactly the kind of work that turns “we think we’re compliant” into “we can produce the report that proves it” the day a regulator, auditor, or board member asks.

Banner with Dedicatted logo on a dark, abstract network background. Large headline reads: “Helping you harness the power of Agentic AI to unlock your business innovation,” with “unlock your business innovation” highlighted in teal. Below the text is a row of AWS Partner badges indicating multiple competencies and certifications.

Contact our experts!


    By submitting this form, you agree with our Terms & Conditions and Privacy Policy.

    File download has started.

    We’ve got your email! We’ll get back to you soon.

    Oops! There was an issue sending your request. Please double-check your email or try again later.

    Oops! Please, provide your business email.