Spectralgraph
The sample report

Here is the whole report. Read it before you decide anything.

A complete LLM Validation Report for Meridian Mutual Insurance, a client we invented, because real client reports stay confidential. The measurement behind it came off the same instrument a paid engagement uses.

REAL INSTRUMENT OUTPUT · FICTIONAL CLIENT · CLIENT REPORTS ARE NEVER PUBLISHED
If you sell AI rather than run it

There is a second sample, framed for your buyers' reviewers.

The report above is written for a company that runs a model under a regulator. If you sell an AI product, the people asking you hard questions are your customers' reviewers, so the pack is framed for them instead. To be plain about it, this is the same measurement as the report above, presented the way a vendor's buyers read it. The instrument doesn't change with the audience. Section 9 does, and it maps the report to what those reviewers actually ask rather than to an insurance bulletin.

REAL INSTRUMENT OUTPUT · FICTIONAL VENDOR · CLIENT REPORTS ARE NEVER PUBLISHED
How to read it

Five sections, in the order a reviewer reads them.

Results summary. What the model did, against the baselines it was measured on.

Per-domain reliability. Which question domains held up and which ones slipped.

Flagged answers. Every response the instrument flagged, quoted, with its risk index and percentile, so your team can check each one.

Machine-readable findings. A findings file ships with every engagement, so the evidence goes straight into your own tooling.

Methods and limitations. What was measured, how, with what error rates, and what the measurement can't see.

The instrument behind it

A screening measurement of the model itself.

Every number in the sample came off the production instrument, which reads the model's internal signals rather than grading its text. Published performance is 0.852 AUC on models it had never seen: fabricated-entity discrimination, length-controlled, leave-one-model-out, in the lab's method paper (DOI: 10.5281/zenodo.21365654). The hash-verified findings.csv ships beside every report, and this one is downloadable above. The engagement packet page lists what else a paid engagement includes.

Straight answers

Questions people ask about the sample

Is Meridian Mutual a real client?
No. Meridian Mutual Insurance was invented for this demonstration. Every table in the document came off the real production pipeline. Client reports are confidential and are never published.
Will my report be published like this one?
No. Your report is confidential and it belongs to you. Spectralgraph never publishes a report or names a client without written permission, which is why this sample uses a client who doesn't exist.
What would be different in my report?
Your question domains, your model configuration, your flagged answers, your baselines. The structure is the same one you see here.
Why does the sample admit weaknesses?
Because a reviewer needs to know where the measurement stops. Every report states its scope, its error rates, and what stays outside them.

Want this file with your model's name on it?

Send your question types and an endpoint you designate. The signed report is in your governance file five business days later. The price is published, and it never depends on what the measurement finds.

Email the lab
REPLIES SAME DAY · SEE PUBLISHED PRICING