CollabPoint
← Insights
Compliance

AI Governance: 7 Essential Steps for Banks

Build an AI governance framework your examiners accept: committee structure, model inventory fields, SR 11-7 mapping, and vendor criteria. Start today.

8 min read
AI Governance: 7 Essential Steps for Banks
Quick answer

AI governance for a mid-market bank rests on three things: a policy that assigns decision rights, a model inventory with an owner and risk tier for every AI system, and committee evidence that high-risk models were reviewed before go-live. Build it as an extension of your existing SR 11-7 model risk program, not as a separate parallel structure.

AI governance at a mid-market bank comes down to three artifacts an examiner can read in an afternoon: a written policy that assigns decision rights, a model inventory that lists every AI system with an owner and a risk tier, and documented evidence that a committee reviewed the high-risk ones before they reached production. Everything else is supporting detail. If you have those three things and they match reality, you will hold up under scrutiny. If you have a 40-page policy and no inventory, you will not.

This is written for banks between roughly $1B and $20B in assets, where there is no dedicated AI office, the model risk function is one or two people, and business units have already started using generative tools whether or not anyone approved it.

What AI Governance Actually Means for a Mid-Market Bank

Most banks already have a model risk management program. The mistake is assuming AI governance is a separate discipline that needs its own parallel structure. It is not. It is an extension of what you built for your CECL model and your BSA transaction monitoring system, with three new problems layered on top.

Problem one: the model is often somebody else's. When your core provider embeds a fraud score or your marketing platform adds a propensity engine, you inherit a model you cannot inspect. Problem two: generative systems produce different output for identical input, which breaks the validation playbook built for logistic regression. Problem three: adoption is decentralized. A commercial lender pasting a credit memo into a public chatbot is a governance event, and nothing in your existing model inventory captures it.

A working AI governance framework handles all three without doubling your control environment.

SR 11-7 and OCC 2011-12 Are Still the Rulebook

Examiners are not applying a new AI standard. They are applying the one that already exists. SR 11-7 and the parallel OCC Bulletin 2011-12 define a model as a quantitative method that processes input data to produce output used in business decisions. A large language model summarizing loan documents that a credit officer relies on fits that definition. So does a vendor's AI-driven alert scoring in your AML platform.

Three requirements from SR 11-7 drive most of the AI governance work:

  • Effective challenge. Someone with authority, competence, and independence from the model developer has to push back. For AI, that means your validation function needs people who understand prompt sensitivity and retrieval behavior, not just backtesting.
  • Comprehensive inventory. The guidance is explicit that banks should maintain an inventory of all models in use, under development, and recently retired. Most AI inventory failures are omissions, not errors.
  • Ongoing monitoring. Performance drift for a generative system is harder to measure than for a scorecard, but silence is not an acceptable answer.

Pair those with the NIST AI Risk Management Framework, which gives you the Govern, Map, Measure, Manage structure that regulators recognize as a reasonable control taxonomy. Do not adopt NIST instead of SR 11-7. Map one to the other in a single crosswalk table and put it in the policy appendix.

Committee Structure That Survives an Exam

Do not create a new standalone AI committee if you can avoid it. Extend the model risk committee, add the right seats, and give it a clear escalation path to the board risk committee. A brand-new body with unclear authority is worse than an existing one with a broader charter.

Who Needs a Seat

  • Chief Risk Officer or model risk officer, as chair
  • CIO or head of technology
  • Chief Information Security Officer
  • Chief Compliance Officer, with fair lending and UDAAP coverage
  • Data privacy or legal counsel
  • One business line sponsor, rotating by use case
  • Internal audit, as a non-voting observer

Seven to nine voting members. Meet monthly. Publish a decision log with dates, attendees, use case, risk tier assigned, and conditions of approval. That log is often the single most persuasive document in an exam.

Model Risk Officer Versus AI Product Owner

These two roles get blended, and it causes real problems. Keep them separate in the policy.

The model risk officer owns the standard, the inventory, and validation independence. They do not build, and they do not pick tools. They decide whether a use case is a model, assign the risk tier, define the validation depth, and hold approval authority to block deployment. They report through risk, never through technology.

The AI product owner sits in the business or IT and owns the outcome. They write the use case definition, document data inputs and intended use, run pilots, gather user feedback, and own remediation when performance degrades. They are accountable for benefit realization and for keeping the inventory record current.

One rule that prevents most conflicts: the AI product owner can never approve their own risk tier, and the model risk officer can never be the named owner of a production AI system.

The Model Inventory Fields That Hold Up Under Scrutiny

A spreadsheet is acceptable at this asset size. A spreadsheet nobody updates is not. Minimum fields for each entry:

  • Model ID and version, with a stable naming convention
  • Business owner and technical owner, by name
  • Use case and intended use statement, including what the output must not be used for
  • Model type: statistical, machine learning, generative, or vendor black box
  • Underlying model or platform, named specifically. Claude Sonnet via Amazon Bedrock, GPT-4o via Azure OpenAI, an open-weight Llama or Mistral deployment, or an embedded vendor engine
  • Data inputs and classification, including whether NPI or customer PII touches the model
  • Data residency and retention terms, especially for API-based models
  • Risk tier and the rationale for it
  • Human-in-the-loop control: advisory, reviewed, or autonomous
  • Validation status, validator name, date, and next review date
  • Known limitations and compensating controls
  • Approval date and committee minute reference
  • Decommission date if retired

Use three tiers, not five. Tier 1 covers anything influencing credit, pricing, adverse action, BSA/AML disposition, or capital. Tier 2 covers customer-facing output and operational decisions with financial impact. Tier 3 covers internal productivity where a human reviews everything, such as meeting summaries in Microsoft 365 Copilot or a first-draft policy edit. Tier 1 gets full independent validation. Tier 3 gets an intake record and an annual attestation. If you tier everything as high risk, the program collapses under its own weight within two quarters.

Vendor AI Assessment Criteria for AI Governance

Most AI reaching your bank arrives through vendors, not internal development. Your third-party risk questionnaire probably does not ask the right questions yet. Add these:

  • Which foundation models are used, from which providers, and in which regions?
  • Is customer data used for model training or fine-tuning? Get the contractual no in writing, not a link to a marketing page.
  • What documentation supports conceptual soundness, and will they share validation evidence or a model card under NDA?
  • What are the tested error and hallucination rates for the specific task, not a generic benchmark score?
  • What logging is available to you, and can you export prompts and outputs for your own review?
  • How and when are model versions changed, and do you get advance notice?
  • Can the AI feature be disabled at the tenant level?
  • What are the fair lending and disparate impact testing provisions if the output touches consumers?

If a vendor cannot answer the training data question or the version change question, that is your finding. Document it, tier the use case accordingly, and add a compensating control.

A Practical AI Governance Policy Outline

Twelve to eighteen pages. Anything longer will not be read by the people who need it.

  • 1. Purpose and scope, including the definition of an AI system and how it relates to your existing model definition
  • 2. Regulatory basis, with the SR 11-7, OCC 2011-12, and NIST AI RMF crosswalk
  • 3. Roles and responsibilities, covering board, committee, model risk officer, AI product owner, validators, and end users
  • 4. Prohibited and restricted uses, such as fully autonomous adverse action or NPI in unapproved public tools
  • 5. Intake and approval workflow, with service levels by tier
  • 6. Risk tiering criteria with worked examples from your own bank
  • 7. Validation standards by tier
  • 8. Data governance, covering classification, retention, and permitted training use
  • 9. Third-party AI requirements
  • 10. Ongoing monitoring and drift triggers
  • 11. Incident response for AI failures, including customer harm and disclosure paths
  • 12. Training and attestation requirements
  • 13. Inventory maintenance and audit rights

Where AI Governance Programs Break

Three failure patterns show up repeatedly. The first is shadow adoption discovered during the exam rather than during intake. Before you write policy, run a discovery pass: pull your SaaS spend, check tenant-level usage reports in Microsoft Purview or your CASB, and interview five business leaders. You will find systems nobody registered.

The second is validation capacity. If you tier fifteen use cases as Tier 1 and have one validator, the queue becomes the control failure. Budget external validation support for the first cycle.

The third is a policy written for a technology stack you have not chosen. Write it model-agnostic. A bank running Azure AI Foundry for internal copilots, Claude Opus for long-document analysis in credit review, and an open-weight model on-premises for document classification needs one framework, not three. The Microsoft responsible AI documentation is useful for control patterns, but do not let any single vendor's terminology become your policy language.

Done properly, AI governance is not a brake. It is the thing that lets you say yes to the fourth use case in six weeks instead of six months, because the pattern is already approved and the evidence is already filed.

Talk to CollabPoint

Want a second set of eyes?

Our team works with mid-market IT leaders to capture the upside of AI and the Microsoft cloud without the compounding risk. Start with a focused conversation.

Frequently asked questions

Does SR 11-7 apply to generative AI and large language models?

Yes, in most banking use cases. SR 11-7 defines a model as a quantitative method that processes input data to produce output used in business decisions. An LLM summarizing credit documents or drafting customer communications that staff rely on meets that definition. Examiners are applying existing model risk guidance to AI rather than waiting for a new standard.

Should we create a separate AI governance committee?

Usually no. For banks under roughly $20B in assets, extend the existing model risk committee charter and add seats for the CISO, privacy counsel, and a rotating business sponsor. A new committee with ambiguous authority creates gaps. What matters is a documented decision log with attendees, risk tier assigned, and approval conditions.

What is the difference between a model risk officer and an AI product owner?

The model risk officer owns the standard, the inventory, and validation independence, and holds authority to block deployment. The AI product owner sits in the business or IT and owns the use case, data documentation, pilot results, and remediation. The product owner can never assign their own risk tier, and the model risk officer can never own a production system.

What fields belong in an AI model inventory?

At minimum: model ID and version, business and technical owner, intended use statement, model type, the specific underlying model or platform, data inputs and classification, data residency, risk tier and rationale, human-in-the-loop control level, validation status and next review date, known limitations, approval date with committee minute reference, and decommission date.

How many risk tiers should an AI governance framework use?

Three. Tier 1 for anything touching credit, pricing, adverse action, BSA/AML disposition, or capital. Tier 2 for customer-facing output and operational decisions with financial impact. Tier 3 for internal productivity with full human review. Five-tier schemes slow intake without improving control quality at this asset size.

How do we govern AI that arrives embedded in a vendor product?

Add AI-specific questions to third-party risk reviews: which foundation models are used, whether your data trains them, what validation evidence the vendor will share, what logging you can export, how version changes are communicated, and whether the feature can be disabled at the tenant level. Unanswered questions become documented findings with compensating controls.

How long does it take to stand up an AI governance framework?

A first usable version typically takes six to ten weeks: two weeks of discovery to surface shadow adoption, two to three weeks to draft policy and tiering criteria, and the remainder to populate the inventory and run the first committee cycle. Full validation capacity for Tier 1 models usually takes another quarter.

We use cookies for analytics and to measure our ads. You can accept or decline.