Spectralgraph

Your AI answers confidently.
We measure when it shouldn't.

Spectralgraph is an independent AI-testing laboratory. We measure the fabrication risk of the language model your organization hosts — using your real question types, reading the model's own internal signals — and deliver a signed, dated report built for your governance file.

No customer data touched  ·  Nothing installed  ·  Fixed fee  ·  Five business days

US Patent Pending App. No. 19/724,790 Published methodology · DOI 10.5281/zenodo.21365655 Proof of method · 18 models measured, 15 domains Spectralgraph LLC · Enid, Oklahoma
The question that is coming

Someone is going to ask for your AI testing evidence. The only question is when.

Insurance examiners
Oklahoma adopted the NAIC's AI bulletin as OID Bulletin 2024-11: insurers should document testing and validation of AI systems, before and after deployment.
Bank & CU examiners
The Fed, OCC and FDIC’s revised model-risk guidance (April 2026) replaced SR 11-7 and placed generative AI expressly outside its scope — while stating a bank’s own risk management should govern the tools it does not cover. Nobody is going to hand you the test.
Joint Commission
The Responsible Use of AI in Healthcare certification (June 2026) asks health systems to demonstrate monitoring, evaluating, and validating of AI safety performance.
Courts & the ABA
Courts have sanctioned lawyers over AI-fabricated citations, and ABA Formal Opinion 512 put the duties of competence and confidentiality around generative AI in writing.

When that question arrives, what's in the file?

The instrument

We don't grade the model's answers. We read its internal signals while it writes them.

When a language model retrieves a fact, its token probabilities concentrate. When it fabricates, they scatter across interchangeable guesses. That difference is measurable at every token — without ground truth, without your data, in a single pass. It's not a checklist or a vendor questionnaire. It's a measurement. In testing so far it catches fabricated answers at 0.852 AUC on models it has never seen before — fabricated-entity discrimination, length-controlled, leave-one-model-out — published in the lab's method paper (DOI: 10.5281/zenodo.21365655).

Fabrication detection, token by token
What is the recommended dosage of Celtrazine for pediatric patients with acute bronchial inflammation?
Token DR signal

"Celtrazine" does not exist. Watch the model invent an approval year, a dosage, and a loading dose — and watch the signal catch each one. Simulated for the web; the real instrument runs against your model, on your questions.

The engagement

Five business days. Nothing installed. No meetings required.

01

Send your question types

A sample of the kinds of questions your model faces in production. Not customer records, not member data, not PHI — just the question types.

COVERED BY WRITTEN CONFIDENTIALITY AGREEMENT
02

Designate an endpoint

Point us at a model endpoint you control. The measurement runs against your deployment, inside systems you govern. Your data never leaves your hands.

NO-PHI VENDOR BY DESIGN
03

Receive the signed report

A signed, dated report plus the complete findings file, so your team can verify every flagged answer with their own eyes. Baseline established.

DELIVERED ≤5 BUSINESS DAYS FROM MATERIALS-COMPLETE
The deliverable

Judge the work before you spend a dime.

Page one of the sample LLM Validation Report: document control, scope of measurement, and results at a glance with per-domain reliability bars. SIGNED & DATED VERIFIABLE FINDINGS

The sample report is real measurement on a fictional client. Read it first.

  • Results at a glance: domains within baseline, at watch, elevated
  • Per-domain reliability, mapped to your question types, the ones in production
  • The specific flagged answers, so your team can verify each one
  • Plain-scope language your examiners can read without translation
  • Findings file included — every number in the report is checkable
Send me the sample report

Delivered by email, same day. No call required.

Who this is for

Built for organizations that answer to someone.

Insurers

Documentation for OID Bulletin 2024-11's testing-and-validation expectations, before and after deployment, on the models you run.

Banks & credit unions

A dated baseline and quarterly trend record: the answer to the examiner's model-risk question, ready before it's asked.

Health systems

Independent validation evidence aligned with the Joint Commission's Responsible Use of AI certification. No PHI by design.

Law firms

Confidentiality-safe measurement for self-hosted models, with ABA Formal Opinion 512 duties in view. Client files never touched.

Sovereign & data-sovereignty organizations

The measurement runs entirely inside systems you control. Your data never leaves your hands — a design built for organizations where that principle is not negotiable.

PROOF OF METHOD
We ran this instrument across 18 open-weight models and 15 regulated domains under identical conditions, and published every score. Judge the method in the open before you put it on your model. Anthropic’s published circuit-tracing work describes a “known entity” feature behind hallucinations; the lab’s pre-generation signal measures the corresponding event from the outside — complementary, not competing (see The lab). AI Reliability Index →

Engagements & pricing

Fixed. Published. Never contingent on results.

You'll never have to ask a testing lab what it costs. And a lab whose fee depends on the result isn't independent — ours never will be.

Two published lines. Below: measurement for organizations that run AI. Further down: evidence for companies that sell it.

The founding cohort, by selection: three organizations will found this practice. Founding terms are the validation report at $4,500 and Continuous Assurance locked at $4,500/quarter for two years, in exchange for permission to describe the engagement in a short case study you approve in writing — named preferred, anonymized accepted at the same rate. The founding invoice is issued only on delivery, with a written guarantee: if the report gives your governance file nothing you can use, say so in writing within fourteen days and the invoice is cancelled. The guarantee is about the deliverable's usefulness, never about what the measurement finds. One per client, three ever, then list. Selection favors organizations whose examiners are already asking: pilot-state insurers, Connecticut domestics facing the September 1 certification, health systems pursuing Joint Commission RUAIH.

EVERY REPORT SIGNED BY NAME · ERRORS CORRECTED AND RE-ISSUED AT NO COST · E&O BOUND BEFORE WORK BEGINS

LLM Validation Report

$9,500
one-time · list

One model, your question types. Signed report, complete findings file, five business days. Your fabrication-risk baseline.

Where the value compounds

Continuous Assurance

$6,500
per quarter · list

Quarterly re-measurement, a trend record your governance file can point to, change gates when you swap or fine-tune models, and an incident allowance.

One production model: $6,500. A second model or heavy change cadence: $8,000. Founding clients lock $4,500 and $6,000 for two years. Your exact number is in the engagement letter before you sign.

Model Selection Study

$4,500–7,500
one-time

Comparative measurement across candidate models before you commit — chosen on evidence instead of vendor claims.

Two candidates: $4,500. Each additional candidate: $1,500, up to four for $7,500.

For companies that sell AI

Independent evidence your buyers will actually accept.

Enterprise questionnaires grew an AI section. Hospital committees ask for validation evidence. The Vendor Assurance line is priced separately from everything above, with its own founding slots.

Vendor Evidence Pack

$12,500
one-time · list

One product, one model configuration: the signed evidence report, a shareable summary letter, a questionnaire-ready evidence brief, and twelve months of written answers to your buyers' reviewers.

Founding cohort: three vendors, by selection, at $6,500 invoiced only on delivery under the same usefulness guarantee as the founding client program, in exchange for case-study permission (named preferred, anonymized accepted). One post-remediation re-measurement within 60 days included.

Evidence Refresh

$5,000
per quarter · list

Keeps the evidence inside the freshness window enterprise reviews ask about, re-measured as your model changes.

Founding vendors lock $4,000 per quarter for two years.

Ingestion Screen

$8,500
per training event

A ratio-based screening measurement of whether a fine-tune absorbed a specified corpus, with positive and negative controls and its limits stated plainly.

$4,500 alongside an Evidence Pack. CHAI Applied Model Card supplement: $2,500. Full line on the AI vendors page.

No work begins without a signed engagement letter. Everything you share is held confidential under a written agreement first.

Straight answers

The questions a careful buyer should ask.

What access does this require?
A sample of your question types and a model endpoint you designate. That's the whole list. No customer, member, patient, policyholder, or client data. Nothing installed. We'll put the no-PHI design in writing for your vendor file.
We're a hospital. Will you sign a BAA?
You won't need one. No PHI ever reaches us, so there is nothing for a BAA to cover. We provide a written no-PHI design statement instead — the document your privacy office wants.
We're a bank. What about confidential supervisory information?
We never ask for it, accept it, or handle it. The measurement needs your question types and an endpoint. Your supervisory file stays yours.
Is this an audit or a certification?
No. It's an independent measurement. Audits and attestations are licensed work; measurement is ours. The report states what was measured, how, and what was found — in language your examiners read without translation.
Who does the work?
The person who built the instrument. Spectralgraph was founded by Joseph Stephens in Enid, Oklahoma; the method grew out of ultrasonic non-destructive testing and is patent pending (US App. No. 19/724,790). Engage the lab and the founder runs your measurement — not an analyst three layers down.
What does the lab actually stand behind?
Four commitments, in writing. Every report is signed by name and dated. Any factual or methodological error your reviewers find is corrected and the report re-issued at no cost. Errors and omissions coverage is bound before any client work begins. And every technical question your examiners or reviewers raise gets a written answer, on the record. Fees are fixed and never depend on what the measurement finds, so the independence holds on paper.
Why is the founding rate half of list?
Because the first three engagements buy something specific: permission to describe the engagement in a short case study, approved by you in writing before a word is published — named if your policies allow it, anonymized at the same rate if they don't. In return you get the $4,500 report, a two-year lock at $4,500 a quarter on Continuous Assurance, and a founder with everything to prove. The founding invoice arrives only with the delivered report, and it carries a written guarantee: if the report gives your governance file nothing usable, say so in writing within fourteen days and the invoice is cancelled — a guarantee about the deliverable, never about what the measurement finds. Same instrument, same report as list. The method itself is published with a DOI you can check before spending a dime, and the report arrives signed, with errors corrected and re-issued at no cost. After three, the program closes for good.
What does the measurement not do?
It doesn't decide truth, and it doesn't replace your people. It measures the model's fabrication signals and flags exactly what deserves human eyes. Every report says so on page one.

The next step costs nothing and commits you to nothing.

One email gets you the complete sample report. Put it in front of your risk committee and judge the deliverable itself. Worst case, you've spent five minutes and you know exactly what independent AI testing evidence looks like.

Request the sample report
CONTACT@SPECTRALGRAPH.AI · REPLIES SAME BUSINESS DAY