Spectralgraph
Insurers · regulatory crosswalk

The AI Risk Evaluation Supplement, item by item.

The Supplement is the successor to what NAIC called the AI Systems Evaluation Tool, and version 5.0 is out for public comment. A few of its items ask for something a measurement can put a number behind. Several ask for facts only the company holds. This crosswalk separates them, and says plainly which items stay outside a fabrication measurement's scope.

VERSION 5.0 IS A DRAFT EXPOSED FOR COMMENT · REFERENCE NUMBERS MOVED BETWEEN 4.0 AND 5.0 · QUOTE THE TEXT, NOT THE NUMBER
Scope statement. The Spectralgraph LLM Validation Report is an independent screening measurement of fabrication risk in a designated, insurer-hosted generative AI (large language model) system. It is not an audit, attestation, examination, or opinion, and it does not certify compliance with the Supplement, with any state's examination process, or with any law or regulation. This crosswalk identifies, item by item, where the Report provides responsive documentation ("Addressed"), where it can serve as one input among others ("Supporting"), and where the item is outside the Report's scope entirely ("Outside scope"). Items marked Outside scope remain the insurer's responsibility through its own AIS Program, and Spectralgraph makes no representation regarding them.
Reference numbers moved, so quote the text. Between version 4.0 and version 5.0 the Exhibit C testing-date field moved from C-9 to C-12, and C-9 now means model limitations. A file citing the old number cites a different line while claiming precision. Version 5.0 is itself a draft, and a further revision is anticipated before adoption is considered, so the item text is the durable reference and the numbering is not.

Exhibit C: model-level items

Supplement itemStatusWhere this appears in the Report / basis
C-12: "Last date of model testing" Addressed A date field, and only testing that actually happened fills it. The engagement date, or the most recent quarter where Continuous Assurance is engaged.
C-9: "Model limitations (i.e., limitations on uses cases or data inputs outside the domain of the training data, or other guardrails to manage model related risk)" Addressed Per-domain results naming where the deployed model holds up and where it does not, measured on the insurer's own question types. The parenthetical describes the condition the instrument reads: inputs outside the domain of the training data.
C-11: "Discuss testing model outputs (e.g. model drift, accuracy, unfair trade practices, unfair discrimination, performance degradation…) and how the model was validated prior to being deployed as well as how its performance is monitored on an ongoing basis" Supporting Pre-deployment measurement, and drift and performance degradation across quarters where Continuous Assurance is engaged. Unfair discrimination and unfair trade practices sit in this same item and are outside scope; they need a different instrument and a different vendor. This item is never covered whole.
C-8: "Model risk(s) (i.e., describe the potential for adverse consumer impact)" Supporting The findings register names where the deployed model fabricates, which is one route to an adverse consumer outcome. Adverse consumer impact is broader than fabrication and includes harms this Report does not measure.
C-1: AI model name and version number
C-5: model development, internal, "owned" third party or true third party
C-6: whether the company is able to modify the model
C-7: model risk classification
Outside scope These are the company's own facts and no outside party should be answering them. The Report records the exact model and weights measured, so its evidence and the insurer's answer agree, but the insurer states them. C-7 in particular: the Supplement puts the classification criteria with the company, and a measurement is evidence going into that judgment rather than the judgment itself.

Exhibit B: AIS Program narrative and checklist

Supplement itemStatusWhere this appears in the Report / basis
Narrative: "the validation and testing procedures performed on internally-developed AI Systems, including whether validation and testing are performed by someone that is independent from development" Addressed The whole of an engagement on a self-hosted or fine-tuned model. The independence clause asks who performed the testing; an outside laboratory with no role in building the system is a direct answer, and internal QA answers only half the question.
Narrative: "the testing and verification that has occurred including frequency, scope and methodology" Addressed All three named parts in writing. Frequency: the engagement date, and the quarterly cadence where engaged. Scope: the domains measured and responses per domain. Methodology: published and citable rather than described (DOI: 10.5281/zenodo.21365654).
Checklist: "Evaluates whether AI Systems are suitable for their intended use and should continue to be used as designed" Addressed Per-domain measurement against the insurer's own question types. The second clause asks a question that can only be answered by measuring again later, which is the quarterly trend record.
Checklist: "Quantifies AI System risk levels" Addressed A number per domain from an instrument whose discrimination is published: 0.852 AUC on fabricated-entity separation, length-controlled, on models the instrument had never seen. A measured figure rather than a workshop rating.
Checklist: "Governs, monitors, tests and ensures transparency of AI Systems developed by vendors" Supporting The tests half, wherever the insurer can designate an endpoint that exposes prompt-side token probabilities. Governance is the insurer's. Transparency here means explainability, and a fabrication index does not explain a decision, so that half is outside scope.
Narrative: validation and testing procedures performed on third-party vendor-supplied AI Systems Supporting Available where the insurer can designate an endpoint we can measure. Pure vendor-hosted SaaS the vendor attests to is outside scope, and whether a given deployment can be measured is a question with a factual answer worth settling before anyone signs anything.
Checklist: "Evaluates the risk of Adverse Consumer Outcomes"
Checklist: "Provides standards and guidance for procuring and engaging AI System vendors"
Supporting Fabrication is one input to the first, not the whole of it. For the second, pre-deployment measurement across candidate models is evidence for a procurement decision; the procurement standard itself is the insurer's.
Exhibit A in full, including the model inventory
Unfair discrimination and consumer-facing disclosure
Explainability and reason codes
The AIS Program itself
Outside scope Exhibit A is an inventory exercise. Discrimination, disclosure and explainability each want a different instrument and a different vendor. The governance sections describe a program; a measurement is evidence that one control operated, never the program.

Source: NAIC AI Risk Evaluation Supplement version 5.0 and its accompanying summary of changes, both read from NAIC's own server on 2026-09-10, alongside the Big Data and Artificial Intelligence (H) Working Group's published materials. Version 5.0 is exposed for public comment and subject to revision; this page will be re-verified against later versions. Nothing here states or implies that NAIC endorses this laboratory, or that any bulletin, supplement or examination requires an outside measurement. See also the model bulletin crosswalk, the state adoption table, and the NIST AI RMF MEASURE crosswalk.

Want this mapping against your own program?

Send your question types and the model you host, and the sample report shows exactly what the evidence looks like before you spend a dime.

Email the lab
REPLIES SAME DAY · EVERYTHING ANSWERED IN WRITING, ON THE RECORD