An examiner is going to ask what testing you did on the AI.
Right now there is almost nowhere to get it.
We are an independent testing lab. We measure how often the language model your institution hosts makes things up, using your own question types, and hand you a signed, dated report for the model-risk file.
SR 26-2 left generative AI out of scope. Your examiner still has questions.
The revised guidance scopes generative AI out, and nothing in it requires an independent measurement. Examiners apply the same model-risk principles by analogy, and banks would rather have their own testing evidence in the file.
What the instrument does is read the model's internal signals while it writes an answer, rather than grading the answer afterward. Point it at the question types your staff actually use, on a model you host or designate, and it maps where the model fabricates and how often. Every flagged answer stays in the findings file so your own people can check it. Detection performance so far: 0.852 AUC on models it has never seen, fabricated-entity discrimination, length-controlled, leave-one-model-out. Method paper DOI: 10.5281/zenodo.21365654.
No customer or member data. Nothing installed. Five business days.
It never touches customer or member data. It needs a sample of your real question types and a model endpoint you designate: a self-hosted open-weight model, a vLLM-class private-cloud deployment, or a fine-tune your institution controls. Private Azure OpenAI exposes probabilities on output tokens only, so that takes behavioral testing instead.
The first engagement gives you the baseline. Quarterly re-measurement after that builds the trend record your governance file points to as models change.
The price is published and it never depends on what I find.
The LLM Validation Report is $9,500. Continuous Assurance is $6,500 per quarter for one model. The pre-deployment Model Selection Study is $4,500 for two candidates, up to $7,500 for four. Full terms are on the pricing page. No work begins without a signed engagement letter, and no fee is ever success-based.
Questions bank risk teams ask first
Does SR 26-2 require this measurement?
What does the test need access to?
We only use vendor-embedded AI. Is this for us?
What about confidential supervisory information?
Read the report before you decide anything.
Email us and the sample validation report comes back the same day. It is the same document your file would hold, so you can see exactly what you would be buying.
Email the lab