We measure how often the language model you run makes things up, using your own questions, and hand you a signed report your governance file can use.
No customer data touched · Nothing installed · Fixed fee · Five business days
When that question arrives, what's in the file?
When a language model knows something, its token probabilities concentrate. When it makes something up, they scatter across interchangeable guesses. That difference is measurable at every token, without ground truth and without your data. The method grew out of ultrasonic non-destructive testing, which is how you inspect a weld without cutting it open.
In testing so far it catches fabricated answers at 0.852 AUC on models it has never seen before: fabricated-entity discrimination, length-controlled, leave-one-model-out. Published in the lab's method paper (DOI: 10.5281/zenodo.21365655).
"Celtrazine" does not exist. Watch the model invent an approval year, a dosage, and a loading dose, and watch the signal catch each one. Simulated for the web; the real instrument runs against your model, on your questions.
The kinds of questions your model faces in production. No customer records, no member data, no PHI.
COVERED BY WRITTEN CONFIDENTIALITY AGREEMENTWe measure against a deployment you control, inside systems you govern. Your data never leaves your hands.
NO-PHI VENDOR BY DESIGNSigned and dated, with the complete findings file, so your team can check every flagged answer.
DELIVERED ≤5 BUSINESS DAYS FROM MATERIALS-COMPLETE
SIGNED & DATED
VERIFIABLE FINDINGS
Delivered by email, same day. No call required.
Evidence for OID Bulletin 2024-11, on the models you run.
A dated baseline and quarterly trend record, ready before the examiner asks.
Validation evidence for Joint Commission RUAIH. No PHI by design.
Confidentiality-safe measurement for self-hosted models. Client files never touched.
The measurement runs entirely inside systems you control, for organizations where that is not negotiable.
PROOF OF METHOD
We ran this instrument across 18 open-weight models and 15 regulated domains under identical conditions, and published every score. Judge the method in the open before you put it on your model. More on how it compares to published interpretability research on The lab.
AI Reliability Index →
Every report is signed by name and dated.
Any factual or methodological error your reviewers find is corrected and the report re-issued at no cost.
Errors and omissions coverage is bound before any client work begins.
Every technical question your examiners or reviewers raise gets a written answer, on the record.
Below: measurement for organizations that run AI. Further down: evidence for companies that sell it.
The founding cohort, by selection: three organizations will found this practice at $4,500, with Continuous Assurance locked at $4,500/quarter for two years, in exchange for a short case study you approve in writing, named or anonymized at the same rate. Invoiced only on delivery. If the report gives your governance file nothing you can use, say so in writing within fourteen days and the invoice is cancelled. That guarantee is about the deliverable, never about what the measurement finds. One per client, three ever, then list.
One model, your question types. Signed report, complete findings file, five business days. Your fabrication-risk baseline.
Quarterly re-measurement, a trend record your governance file can point to, change gates when you swap or fine-tune models, and an incident allowance.
One production model: $6,500. A second model or heavy change cadence: $8,000. Founding clients lock $4,500 and $6,000 for two years. Your exact number is in the engagement letter before you sign.
Comparative measurement across candidate models before you commit, chosen on evidence instead of vendor claims.
Two candidates: $4,500. Each additional candidate: $1,500, up to four for $7,500.
Enterprise questionnaires grew an AI section. Hospital committees ask for validation evidence. The Vendor Assurance line is priced separately from everything above, with its own founding slots.
One product, one model configuration: the signed evidence report, a shareable summary letter, a questionnaire-ready evidence brief, and twelve months of written answers to your buyers' reviewers.
Founding cohort: three vendors, by selection, at $6,500 invoiced only on delivery under the same usefulness guarantee as the founding client program, in exchange for case-study permission (named preferred, anonymized accepted). One post-remediation re-measurement within 60 days included.
Keeps the evidence inside the freshness window enterprise reviews ask about, re-measured as your model changes.
Founding vendors lock $4,000 per quarter for two years.
A ratio-based screening measurement of whether a fine-tune absorbed a specified corpus, with positive and negative controls and its limits stated plainly.
$4,500 alongside an Evidence Pack. CHAI Applied Model Card supplement: $2,500. Full line on the AI vendors page.
No work begins without a signed engagement letter. Everything you share is held confidential under a written agreement first.
More questions answered on the FAQ: hospital BAAs, confidential supervisory information, the founding terms in full, and what the lab stands behind.
One email gets you the complete sample report. Put it in front of your risk committee and judge the deliverable itself. Worst case, you've spent five minutes and you know exactly what independent AI testing evidence looks like.
Read the sample report