AI Governance for Banks: 5 Critical Steps
AI governance for banks: what SR 11-7, OCC guidance and the EU AI Act require before you deploy Copilot, Claude or open-source models. Get the checklist.

AI governance for banks means maintaining a complete inventory of every AI system, tiering it by risk, assigning named owners, testing and monitoring outputs, and keeping evidence an examiner can review. Existing rules already apply: SR 11-7, OCC Bulletin 2011-12, third-party risk guidance and Reg B. Mid-market institutions should stand up AI governance before the first pilot, not after.
AI governance for banks starts before deployment, not after
AI governance is the set of policies, controls, roles and evidence that lets a bank prove it understands, tests and supervises every AI system it uses, and for mid-market institutions it needs to exist before the first pilot goes live, not after an examiner asks for the model inventory. Regulators are not waiting for new AI-specific rules to start asking questions. They are applying existing model risk, third-party risk, fair lending and information security expectations to generative AI right now. If your bank has a Microsoft 365 Copilot pilot running, a vendor scoring engine in the loan pipeline, or three analysts quietly pasting spreadsheets into ChatGPT, you already have an AI program. The only question is whether it is governed.
This guide covers the regulatory landscape mid-market banks actually face, how "AI system" is being defined, the gaps auditors keep finding, and the five elements of a governance framework that will survive an exam.
The regulatory landscape: what already applies to your AI
SR 11-7 and OCC Bulletin 2011-12
The Federal Reserve's SR 11-7 Guidance on Model Risk Management, issued jointly with OCC Bulletin 2011-12, is still the spine of bank AI governance in the United States. It predates transformers by six years and it does not care. Its definition of a model is deliberately broad: a quantitative method that applies statistical, economic, financial or mathematical theories and assumptions to process input data into quantitative estimates.
A large language model scoring commercial credit memos fits. A vendor tool that ranks alert priority in your BSA queue fits. SR 11-7 requires development documentation, independent validation, ongoing monitoring, a model inventory, and clear accountability. Most mid-market banks apply this rigorously to ALLL and stress testing models, then treat a generative AI pilot as an IT project. That inconsistency is the single most common finding.
Third-party risk and consumer compliance
The 2023 interagency guidance on third-party relationships from the Fed, FDIC and OCC applies squarely to AI vendors. You cannot outsource accountability. If a fintech partner uses a model you have never validated to make decisions affecting your customers, that is your risk.
On the consumer side, the CFPB has been explicit that Regulation B adverse action requirements do not bend for complex models. "The algorithm said no" is not a specific reason. If AI touches credit decisioning, collections prioritization, marketing targeting or servicing, your AI governance program needs fair lending testing and explainability evidence attached to it.
The EU AI Act, and why it matters to a US bank
The EU AI Act entered into force in August 2024 with obligations phasing in through 2026 and 2027. Prohibited practices took effect in February 2025. General-purpose AI model obligations followed in August 2025. High-risk classifications, which explicitly include AI used to evaluate the creditworthiness of natural persons, carry the heaviest requirements.
Most mid-market US banks are out of territorial scope. Two reasons to care anyway. First, if you have EU customers, EU-based staff, or a correspondent relationship touching the EU, scope can reach you. Second, the Act's structure, risk tiering, technical documentation, human oversight, logging, post-market monitoring, is becoming the shape of AI governance everywhere. Building to it is cheaper than retrofitting to it.
NIST AI RMF as the connective tissue
Neither SR 11-7 nor the EU AI Act tells you how to build a program. The NIST AI Risk Management Framework and its Generative AI Profile do. It is voluntary, which examiners like, because adopting it voluntarily demonstrates intent. Its four functions, Govern, Map, Measure, Manage, map cleanly onto how bank auditors think. Use it as the organizing structure for your AI governance documentation and you will spend far less time translating between frameworks.
What actually counts as an AI system
Definitional drift is where AI governance programs quietly fail. Banks scope narrowly, then discover the inventory was wrong. Under current regulatory and NIST definitions, treat all of the following as in scope:
- Purchased AI features inside software you already own. Copilot in Microsoft 365, AI summarization in your core platform, anomaly detection in your fraud vendor, AI-assisted underwriting inside an LOS.
- Foundation model APIs. Azure OpenAI Service, Anthropic's Claude via API or Amazon Bedrock, Google Gemini, whatever a developer wired into an internal tool.
- Self-hosted open-source models. Llama, Mistral, Qwen and similar weights running on your infrastructure or in a private cloud.
- Low-code and agentic workflows. Power Platform flows, Copilot Studio agents, scripts calling an LLM to route email or draft documents.
- Employee use of consumer tools. Personal ChatGPT, Claude and Gemini accounts. Unsanctioned use is still bank use if bank data is involved.
A useful test: if the output influences a decision, a customer communication, or a record the bank relies on, it belongs in the AI inventory.
Five gaps auditors keep finding in bank AI governance
- An incomplete inventory. The register lists three models. Discovery finds eleven. Shadow AI in browser extensions and personal accounts is nearly always the gap.
- Policy without evidence. A well-written AI policy approved by the board, and zero artifacts showing it operating. No validation memos, no monitoring reports, no exception log.
- Vendor AI treated as out of scope. "It is the vendor's model" is not a control. Missing due diligence questionnaires, no clarity on where data goes, no contractual right to audit.
- No performance thresholds or monitoring. Nobody defined what "the model is degrading" looks like, so nobody can detect it. For generative AI, this means no accuracy sampling, no hallucination rate tracking, no drift review cadence.
- Undefined human oversight. "A human reviews it" appears in the policy. No definition of what the reviewer checks, what authority they have to override, or how the review is recorded.
A sixth issue is not strictly an AI governance finding but causes more incidents than any of the above: permission sprawl. Deploying a tenant-wide assistant on top of a SharePoint estate with open-to-everyone sites surfaces content people were never supposed to see. Fix entitlements before you fix prompts.
Five foundational elements of a defensible AI governance framework
1. Inventory and risk tiering
One register, one owner, updated on a fixed cadence. For each entry capture purpose, model or vendor, data classes touched, decision impact, and a risk tier. Three tiers is enough. Tier 1 covers anything influencing credit, BSA/AML, or customer-facing decisions. Tier 2 covers internal decisioning and material efficiency use. Tier 3 covers low-risk productivity. Controls scale with tier so your AI governance effort lands where the exposure is.
2. Accountability that maps to three lines of defense
Every AI system gets a named business owner, a named technical owner, and independent review that does not report to either. An AI oversight committee with credit, compliance, risk, IT and legal representation should approve Tier 1 deployments. Give it authority to say no and a documented decision log.
3. Lifecycle controls
Intake and use-case approval, pre-deployment testing including adverse outcome and fair lending testing where relevant, documented limitations, a defined pilot period, formal production approval, change management for model or prompt updates, and a decommissioning process. Prompt and system message changes are model changes. Version them.
4. Data, third-party and contractual controls
Classify what data may enter which system. Confirm in writing whether your inputs train the provider's models. Check retention, sub-processors, regional data residency, and indemnification. Microsoft, OpenAI and Anthropic all publish enterprise terms that differ from their consumer terms, and the difference is significant. Self-hosted open-source models remove the vendor data question and replace it with license, provenance and patching obligations you now own.
5. Monitoring, logging and incident response
Sampling reviews of outputs at a defined frequency. Immutable logs of who used what and when. A defined AI incident type in your existing incident response plan covering data leakage, materially wrong output that reached a customer, and prompt injection through untrusted content. Report to the board at least quarterly. Examiners will ask what the board saw and when.
Vendor-neutral by design: Microsoft, OpenAI, Anthropic, open source
Good AI governance is model-agnostic, because your model choices will change within eighteen months. That said, the risk profiles genuinely differ.
Microsoft 365 Copilot inherits your existing tenant permissions and compliance boundary, which is an advantage if entitlements are clean and a liability if they are not. Microsoft documents its data handling in the Microsoft 365 Copilot privacy documentation, and Azure AI Foundry adds content filtering, evaluation tooling and regional deployment options for custom builds.
Anthropic's Claude models, including Opus and Sonnet, are strong on long-context document analysis and instruction adherence, which suits credit memo and policy review work. OpenAI's models are widely integrated and fast-moving. Both offer enterprise agreements with no-training-on-your-data commitments and IP indemnities. Read the actual contract, not the marketing page.
Open-source models running in your own environment give maximum data control and zero vendor indemnity. You own evaluation, red teaming, patching and license compliance. For some banks that trade is correct. Make it deliberately, and document why.
The framework does not change. The evidence you collect for each does.
Where mid-market banks should start
Ninety days is enough to reach a defensible baseline. Weeks one to three: discovery and inventory, including a survey of actual employee AI use and a review of SaaS contracts for embedded AI features. Weeks four to six: policy, risk tiering criteria, and standing up the oversight committee. Weeks seven to twelve: apply lifecycle controls to your two highest-risk systems, fix the permission sprawl, and produce the first board reporting pack.
Do that and your AI governance program stops being a document and starts being a control environment. More importantly, you can approve new use cases quickly, because the guardrails already exist. The banks moving fastest on AI are the ones that governed it first.
Want a second set of eyes?
Our team works with mid-market IT leaders to capture the upside of AI and the Microsoft cloud without the compounding risk. Start with a focused conversation.
Frequently asked questions
Does SR 11-7 apply to generative AI tools like Copilot or Claude?
Yes, when the output influences a decision, estimate or record the bank relies on. SR 11-7 defines a model broadly and does not exempt purchased or third-party tools. Low-risk productivity uses can carry lighter controls, but they still belong in your AI inventory.
Do US mid-market banks need to comply with the EU AI Act?
Most do not fall in territorial scope, but exposure exists if you serve EU customers or have EU-based operations. Even without scope, the Act's structure of risk tiering, technical documentation, human oversight and logging is becoming the default shape of AI governance globally, so building to it now avoids a later retrofit.
What is the most common AI governance gap auditors find?
An incomplete AI inventory. Banks list the systems IT deployed and miss embedded vendor AI features, low-code agents and employee use of personal ChatGPT or Claude accounts. Discovery routinely turns up three to four times the number of AI systems the register shows.
Should we govern vendor AI differently from models we build?
The framework is the same, the evidence differs. For vendor AI you rely on due diligence, contractual terms, documented data handling and output monitoring rather than development documentation. Regulatory accountability stays with the bank either way.
Is open-source AI riskier than Microsoft, OpenAI or Anthropic?
Different, not automatically riskier. Self-hosting removes the vendor data-sharing question and gives you full control, but you take on evaluation, red teaming, patching, license compliance and you get no IP indemnity. Commercial providers offer indemnities and enterprise data commitments in exchange for less control.
Who should own AI governance in a mid-market bank?
A cross-functional oversight committee with credit, compliance, risk, IT and legal representation, chaired by a senior executive with authority to reject use cases. Each AI system also needs a named business owner and technical owner, with independent review outside both.
How long does it take to build a defensible AI governance program?
About 90 days for a baseline: three weeks for discovery and inventory, three weeks for policy, tiering criteria and committee formation, then six weeks applying lifecycle controls to the highest-risk systems and producing the first board report.
What should we fix before rolling out an AI assistant across the bank?
Permissions. Tenant-wide assistants surface whatever a user already has access to, so open SharePoint sites and stale group memberships become data exposure the moment search gets good. Remediate entitlements and sensitivity labeling before broad deployment.
How CollabPoint can help
Turn this into action — the capability and fixed-scope programs behind it.
More articles
Microsoft Purview Secure by Default: 7 Critical Steps for Banks
Microsoft Purview secure by default is the foundation banks and credit unions need before enabling Copilot. Learn the risks, regulations, and readiness steps. Get started.
Simplifying CMMC Compliance with the Microsoft Cloud
CMMC compliance is complex, but the Microsoft cloud covers a large share of the controls. Here's how GCC, Purview and Defender simplify the path.
Microsoft Purview for Banks: 5 Critical Wins
Microsoft Purview for banks and credit unions in plain English: labels, DLP, Insider Risk and Audit before enabling Copilot. Book a 30-minute scoping call.