Spectralgraph is an independent AI-testing laboratory. We measure the fabrication risk of the language model your organization hosts — using your real question types, reading the model's own internal signals — and deliver a signed, dated report built for your governance file.
No customer data touched · Nothing installed · Fixed fee · Five business days
When that question arrives, what's in the file?
When a language model retrieves a fact, its token probabilities concentrate. When it fabricates, they scatter across interchangeable guesses. That difference is measurable at every token — without ground truth, without your data, in a single pass. It's not a checklist or a vendor questionnaire. It's a measurement. In testing so far it catches fabricated answers at 0.852 AUC on models it has never seen before — fabricated-entity discrimination, length-controlled, leave-one-model-out — published in the lab's method paper (DOI: 10.5281/zenodo.21365655).
"Celtrazine" does not exist. Watch the model invent an approval year, a dosage, and a loading dose — and watch the signal catch each one. Simulated for the web; the real instrument runs against your model, on your questions.
A sample of the kinds of questions your model faces in production. Not customer records, not member data, not PHI — just the question types.
COVERED BY WRITTEN CONFIDENTIALITY AGREEMENTPoint us at a model endpoint you control. The measurement runs against your deployment, inside systems you govern. Your data never leaves your hands.
NO-PHI VENDOR BY DESIGNA signed, dated report plus the complete findings file, so your team can verify every flagged answer with their own eyes. Baseline established.
DELIVERED ≤5 BUSINESS DAYS FROM MATERIALS-COMPLETE
SIGNED & DATED
VERIFIABLE FINDINGS
Delivered by email, same day. No call required.
Documentation for OID Bulletin 2024-11's testing-and-validation expectations, before and after deployment, on the models you run.
A dated baseline and quarterly trend record: the answer to the examiner's model-risk question, ready before it's asked.
Independent validation evidence aligned with the Joint Commission's Responsible Use of AI certification. No PHI by design.
Confidentiality-safe measurement for self-hosted models, with ABA Formal Opinion 512 duties in view. Client files never touched.
The measurement runs entirely inside systems you control. Your data never leaves your hands — a design built for organizations where that principle is not negotiable.
PROOF OF METHOD
We ran this instrument across 18 open-weight models and 15 regulated domains under identical conditions, and published every score. Judge the method in the open before you put it on your model. Anthropic’s published circuit-tracing work describes a “known entity” feature behind hallucinations; the lab’s pre-generation signal measures the corresponding event from the outside — complementary, not competing (see The lab).
AI Reliability Index →
You'll never have to ask a testing lab what it costs. And a lab whose fee depends on the result isn't independent — ours never will be.
Two published lines. Below: measurement for organizations that run AI. Further down: evidence for companies that sell it.
The founding cohort, by selection: three organizations will found this practice. Founding terms are the validation report at $4,500 and Continuous Assurance locked at $4,500/quarter for two years, in exchange for permission to describe the engagement in a short case study you approve in writing — named preferred, anonymized accepted at the same rate. The founding invoice is issued only on delivery, with a written guarantee: if the report gives your governance file nothing you can use, say so in writing within fourteen days and the invoice is cancelled. The guarantee is about the deliverable's usefulness, never about what the measurement finds. One per client, three ever, then list. Selection favors organizations whose examiners are already asking: pilot-state insurers, Connecticut domestics facing the September 1 certification, health systems pursuing Joint Commission RUAIH.
EVERY REPORT SIGNED BY NAME · ERRORS CORRECTED AND RE-ISSUED AT NO COST · E&O BOUND BEFORE WORK BEGINSOne model, your question types. Signed report, complete findings file, five business days. Your fabrication-risk baseline.
Quarterly re-measurement, a trend record your governance file can point to, change gates when you swap or fine-tune models, and an incident allowance.
One production model: $6,500. A second model or heavy change cadence: $8,000. Founding clients lock $4,500 and $6,000 for two years. Your exact number is in the engagement letter before you sign.
Comparative measurement across candidate models before you commit — chosen on evidence instead of vendor claims.
Two candidates: $4,500. Each additional candidate: $1,500, up to four for $7,500.
Enterprise questionnaires grew an AI section. Hospital committees ask for validation evidence. The Vendor Assurance line is priced separately from everything above, with its own founding slots.
One product, one model configuration: the signed evidence report, a shareable summary letter, a questionnaire-ready evidence brief, and twelve months of written answers to your buyers' reviewers.
Founding cohort: three vendors, by selection, at $6,500 invoiced only on delivery under the same usefulness guarantee as the founding client program, in exchange for case-study permission (named preferred, anonymized accepted). One post-remediation re-measurement within 60 days included.
Keeps the evidence inside the freshness window enterprise reviews ask about, re-measured as your model changes.
Founding vendors lock $4,000 per quarter for two years.
A ratio-based screening measurement of whether a fine-tune absorbed a specified corpus, with positive and negative controls and its limits stated plainly.
$4,500 alongside an Evidence Pack. CHAI Applied Model Card supplement: $2,500. Full line on the AI vendors page.
No work begins without a signed engagement letter. Everything you share is held confidential under a written agreement first.
One email gets you the complete sample report. Put it in front of your risk committee and judge the deliverable itself. Worst case, you've spent five minutes and you know exactly what independent AI testing evidence looks like.
Request the sample report