Start with the current authority
On April 17, 2026, the Federal Reserve, OCC and FDIC issued revised model-risk management guidance. Federal Reserve letter SR 26-2 expressly supersedes and replaces SR 11-7 from 2011 and SR 21-8 from 2021, the latter addressing BSA/AML systems. A policy that still identifies SR 11-7 as the current interagency framework needs an authority review. The new document is supervisory guidance, not a statute or a prescriptive regulation. [1][2]
That distinction has practical consequences. The guidance states that it does not establish enforceable standards and that noncompliance with the guidance itself will not produce supervisory criticism. Violations of law or unsafe or unsound practices arising from inadequate management of model risk remain separate matters. Institutions should identify the legal or risk basis for a control rather than describe every internal preference as a regulatory mandate. [2]
Materiality changes the allocation of effort
The revised framework is expected to be most relevant above $30 billion in assets, while recognizing circumstances in which smaller organizations have significant model exposure or complexity. Size is not a sound substitute for understanding use. A small bank relying heavily on a complex third-party underwriting system may have a more consequential model dependency than a larger bank's low-impact internal forecast. [1][2]
Analytical implication: rank review work by the harm a wrong or misused output can cause, the affected exposure, uncertainty and available safeguards. A model used only to prioritize an analyst's reading is not equivalent to a model automatically assigning consumer credit limits. Equally, apparently modest models can become material when embedded in many products or used to feed capital and liquidity decisions.
The definition focuses on complex quantitative methods using statistical, economic or financial theories to produce quantitative estimates. Simple arithmetic and deterministic processes without those theoretical foundations are excluded. This narrows the formal inventory boundary but does not make excluded tools harmless. A spreadsheet that miscalculates refunds still needs accuracy and change controls, even if it is not a model for this guidance. [2]
Separate predictive models from generative agents
The attachment excludes generative and agentic AI from its scope because those technologies are evolving rapidly; traditional quantitative models and non-generative, non-agentic AI remain covered. The agencies point to broader risk management for tools outside scope. The Fed's May 1, 2026 AI speech provides further policy context, but a speech should not be elevated into binding requirements. [2][3]
Recommended architecture: inventory a credit workflow's components separately. A default-probability model, deterministic eligibility rules, document-extraction assistant and agent that drafts a case narrative have different failure modes. Link them into a common decision record so a narrow model definition does not leave the surrounding workflow unowned.
For the predictive component, assess development data, target definition, calibration and use limits. For the generative component, test fabricated facts, prompt injection, source attribution and unauthorized actions. For rules, test precedence, completeness and deployment correctness. These are analytical recommendations for different mechanisms, not a claim that SR 26-2 prescribes a single AI control checklist.
Worked example: the same model can have different risk
Hypothetical comparison: Bank A uses a loss model on a $20 million pilot and requires independent review before each credit decision. Bank B applies the same model automatically to a $2 billion portfolio. If a model error understates expected loss by one percentage point across the relevant exposure, the illustrative error is $200,000 at A and $20 million at B. Actual losses would depend on subsequent behavior, use and portfolio dynamics; the arithmetic is a sensitivity, not a forecast.
The model's code can be identical while exposure and reliance differ substantially. Bank B has a stronger economic reason for rigorous testing, independent challenge, monitoring and rollback capacity. Bank A still needs evidence that manual review actually changes decisions when warranted. A nominal human approval step that simply accepts every output may provide little reduction in risk.
A second comparison concerns shared inputs. Five individually modest models may all depend on the same income feed. A single classification error could distort approvals, line management, collections and loss forecasts together. An inventory that counts models without mapping common dependencies can miss that concentration.
A workable transition plan
Recommended first step: map existing policy provisions to current authority, business risk and internal choice. Preserve useful validation and monitoring practices while removing obsolete claims that a withdrawn letter mandates them. Record changes in governance rationale so later reviewers understand why a control was retained, scaled down or replaced.
Next, reassess materiality and ownership using actual production use. Require an accountable business owner to describe what decisions depend on each model, where overrides occur and how failure would be detected. Establish escalation thresholds linked to loss, consumer outcomes or financial reporting significance rather than a uniform calendar exercise.
For vendor models, negotiate enough evidence and access to test intended use. A proprietary algorithm can still be evaluated through benchmark data, stability, error analysis and outcomes. Where transparency remains limited, narrower use, lower limits or additional review may be more defensible than an unsupported declaration of validation. Outsourcing the model does not outsource the decision to rely on it.
What would change the conclusion
The principal opportunity is more proportionate oversight, with scarce review capacity focused on consequential risks. The principal failure mode is interpreting proportionality as permission to remove controls before understanding exposure. Both overclassification and underclassification have costs: unnecessary documentation can delay beneficial changes, while missing a material dependency can produce widespread errors.
Evidence of successful implementation would include fewer low-value review tasks alongside better identification and correction of consequential weaknesses. A future interagency statement bringing generative AI into scope, revised statutory duties or changed production use would require reassessment. As of this review, the correct baseline is SR 26-2, with broader governance applied explicitly to the components outside its formal boundary.