Someone will read your AI systems file and ask what the model itself does.
That answer has to be measured.
We are an independent testing lab. We measure how often the language model your company hosts makes things up, using your own question types, and hand you dated evidence for the program file.
Where the program file usually runs thin.
Programs built out of policy documents tend to share one weakness. When someone asks how the model itself behaves, the file has adjectives where it needs numbers.
What the instrument does is read the model's own token probabilities while it works, rather than reading the text it hands back. Run it against your real claims, coverage and service questions, on a model your company hosts or designates, and it maps where the model fabricates and how often, down to the specific flagged answers your team can check. What lands in the file is a signed, dated screening measurement. Detection performance so far: 0.852 AUC on models the instrument had never seen, fabricated-entity discrimination, length-controlled, leave-one-model-out. Method paper DOI: 10.5281/zenodo.21365654.
No policyholder data. Nothing installed. Five business days.
It never touches policyholder data. It needs a sample of your real question types, the way your teams actually phrase them, and a model endpoint you designate: a self-hosted open-weight model, a vLLM-class private-cloud deployment, or a fine-tune your company controls. Private Azure OpenAI exposes probabilities on output tokens only, so that takes behavioral testing instead. Vendor-attested SaaS AI is out of scope, and we will say so.
The first measurement is your baseline. Quarterly re-measurement after that builds the trend record as models change. Provision-by-provision detail is in the crosswalk to Bulletin 2024-11, including what stays outside scope. Outside Oklahoma, the same mapping runs against the model text your state adopted: the NAIC model bulletin crosswalk (state adoption table). Connecticut domestics have a date on the calendar, the September 1 annual AI certification. For the examination side, the AI Risk Evaluation Supplement crosswalk maps the Supplement line by line, including the items no outside party should be answering.
The price is published and it never depends on what I find.
The LLM Validation Report is $9,500. Continuous Assurance is $6,500 per quarter for one model. The pre-deployment Model Selection Study is $4,500 for two candidates, up to $7,500 for four. Full terms are on the pricing page. No work begins without a signed engagement letter, and no fee is ever success-based.
Questions insurance compliance teams ask first
Does this satisfy OID Bulletin 2024-11?
Does the test touch policyholder data?
Our AI is inside vendor software. In scope?
What lands in the governance file?
Read the report before you commit to anything.
Email us and the sample validation report comes back the same day, so your compliance team can judge the document itself.
Email the lab