Here is the whole report. Read it before you decide anything.
A complete LLM Validation Report for Meridian Mutual Insurance, a client we invented, because real client reports stay confidential. The measurement behind it came off the same instrument a paid engagement uses.
There is a second sample, framed for your buyers' reviewers.
The report above is written for a company that runs a model under a regulator. If you sell an AI product, the people asking you hard questions are your customers' reviewers, so the pack is framed for them instead. To be plain about it, this is the same measurement as the report above, presented the way a vendor's buyers read it. The instrument doesn't change with the audience. Section 9 does, and it maps the report to what those reviewers actually ask rather than to an insurance bulletin.
Five sections, in the order a reviewer reads them.
Results summary. What the model did, against the baselines it was measured on.
Per-domain reliability. Which question domains held up and which ones slipped.
Flagged answers. Every response the instrument flagged, quoted, with its risk index and percentile, so your team can check each one.
Machine-readable findings. A findings file ships with every engagement, so the evidence goes straight into your own tooling.
Methods and limitations. What was measured, how, with what error rates, and what the measurement can't see.
A screening measurement of the model itself.
Every number in the sample came off the production instrument, which reads the model's internal signals rather than grading its text. Published performance is 0.852 AUC on models it had never seen: fabricated-entity discrimination, length-controlled, leave-one-model-out, in the lab's method paper (DOI: 10.5281/zenodo.21365654). The hash-verified findings.csv ships beside every report, and this one is downloadable above. The engagement packet page lists what else a paid engagement includes.
Questions people ask about the sample
Is Meridian Mutual a real client?
Will my report be published like this one?
What would be different in my report?
Why does the sample admit weaknesses?
Want this file with your model's name on it?
Send your question types and an endpoint you designate. The signed report is in your governance file five business days later. The price is published, and it never depends on what the measurement finds.
Email the lab