Spectralgraph
Frequently asked questions

Straight answers, on the record.

Everything prospects ask us most, answered the way we answer email: plainly, in writing, with the limits stated as clearly as the strengths.

FIXED PRICE · FIVE BUSINESS DAYS · NO MEETINGS, EVERYTHING IN WRITING · NO CUSTOMER DATA
Straight answers

The questions, answered

What exactly does Spectralgraph measure?
Fabrication, the confidently wrong answers language models produce. What it does is measure the model itself, reading the model's own internal signals rather than grading its response output, and it maps where a model fabricates, how often, and on which topics, using your real question types on a model you host. Flagged answers are individually verifiable by your own team.
What access does the test need?
A sample of your real question types and a model endpoint you designate. It never touches customer, member, policyholder, or patient data, and nothing is installed in your environment. Everything you share is held confidential under a written agreement before any work begins.
Which models are in scope?
Models your organization hosts or controls whose deployment exposes prompt-side token probabilities: self-hosted open-weight models, vLLM-class private-cloud deployments, and your own fine-tunes. Private Azure OpenAI is the honest exception: Azure exposes token probabilities on output tokens only, not the prompt side the instrument reads, so those deployments are served with the behavioral protocol rather than the instrument-grade measurement. Vendor-hosted SaaS AI that the vendor attests to is out of scope, and we will tell you so rather than sell you something with nothing to measure. One deliberate exception: AI product companies selling on closed APIs are served through the vendor line's behavioral evidence protocol: details on the AI vendors page.
Is this an audit or a certification?
No. The deliverable is a screening measurement: a signed, dated report of measured model behavior. We are not an audit firm or a certification body, and we do not issue opinions or attestations. That distinction is deliberate and we keep it everywhere.
How accurate is the instrument?
It catches fabricated answers at 0.852 AUC on models it had never seen before: fabricated-entity discrimination, length-controlled, under leave-one-model-out validation, a result published in the lab's method paper (DOI: 10.5281/zenodo.21365654). A public demonstration of the method, the AI Reliability Index, covers 18 open-weight models across 15 regulated domains.
Why no meetings?
The lab is written-first and async on purpose: every question answered in writing, on the record. Compliance buyers tend to prefer having the answers in the file over having had a phone call about them.
What does it cost?
Fixed fees, never success-based: LLM Validation Report $9,500; Continuous Assurance $6,500 per quarter for one model, $8,000 for a second model or a heavy change cadence; Model Selection Study $4,500 for two candidates, plus $1,500 per additional candidate, to a maximum of four for $7,500. A founding program of three engagements in total, across both price lines, is described on the pricing page.
How fast is delivery?
Five business days from materials-complete: your question types received and your endpoint reachable. Replies to email come the same day.
What arrives at the end?
A signed, dated validation report plus the complete findings file: fabrication frequency by topic, the specific flagged answers, and the measured baseline your governance file keeps. Ask for the sample report by email and judge the deliverable before spending a dime.
Does the EU AI Act require this kind of testing?
No, and we won't pretend otherwise. The EU deferred its high-risk AI obligations to December 2, 2027 (adopted June 2026), and even then conformity is provider self-assessment; nothing in EU law requires an American insurer, bank, or health system to buy independent testing. One honest footnote: the EU's code of practice for frontier-model makers now expects qualified independent external evaluators, the clearest sign anywhere in law that independent AI testing is becoming the norm. The obligations with living dates for our clients are the US ones: NAIC model bulletin adoptions, the Joint Commission certification, and ABA guidance.
What is Connecticut's September 1 AI certification?
Connecticut adopted the NAIC model bulletin as Bulletin MC-25 and asks every Connecticut domestic insurer to certify annually, on or before September 1, either that it does not use AI systems or that its use is substantially consistent with the bulletin. The certification is the insurer's own statement; nothing requires independent testing. What a measurement adds is a dated, independent record to stand behind that signature. The NAIC model bulletin crosswalk on this site maps the report to the exact text Connecticut adopted. The Connecticut page lays out the timeline.
Can a model be prompted or tuned to pass the screening?
Not by coaching it. Context injection changes what a model says, not the probability geometry the measurement reads: in the published injection experiments, prompt-side coaching improved the text while the geometry barely moved, with architecture-specific boundary conditions documented alongside (DOI: 10.5281/zenodo.21365654). The score reads whether trained pathways exist in the weights, so the reliable way to raise it is to train the knowledge in, at which point the model has genuinely improved.
Can fabrication be caught while the model is writing, not just measured afterward?
In the lab, yes. The published method reads three stages: a pre-generation phase that reacts while the model is still processing the prompt, before any answer exists; per-token signals during generation; and the aggregate measurement afterward (DOI: 10.5281/zenodo.21365654). The Beyond the Scorecard page shows demos of the first two, built from real measurement data. The difference from the detection work now appearing in research is structural: our instrument is not a trained classifier. There's no labeled dataset behind it and it never touches the model's internal states; it reads the output probabilities any served endpoint can expose, and the published validation was run on models it had never seen. What we sell today is the measurement: the signed report and the quarterly trend record. A live version that acts on a detection the way your policy chooses, block it or flag it and let it pass, is in development on the same patent-pending instrument. If in-stream handling is what your deployment needs, write us and say so; that interest shapes the build order.
Is fabrication risk measured before or after the model answers?
Both. The reading phase measures whether the model recognizes a topic while it is still processing the prompt, before any answer exists, and the writing phase plus three temporal signals monitor the generated response. The pre-generation signal reaches comparable effect sizes from roughly ten times fewer tokens, and to our knowledge no other evaluation framework measures that phase at all. Per-model results for all three stages are published (DOI: 10.5281/zenodo.21365654).
Is a bigger model safer?
Not by default. In the published 18-model index an 8B model outscores a 70B from the same family, because small models fabricate visibly while large models fabricate with the same fluent geometry as genuine retrieval, which makes visible uncertainty a safety feature. Aggregate scores also hide failure shape: two models a dozen points apart can differ in kind, one growing more cautious on harder questions while the other collapses at the frontier tier. The tier-by-tier curves are in the method paper (DOI: 10.5281/zenodo.21365654).
We're a hospital. Will you sign a BAA?
You won't need one. No PHI ever reaches us, so there is nothing for a BAA to cover. Spectralgraph is a no-PHI vendor by design. We provide a written no-PHI design statement instead, which is the document your privacy office actually wants.
We're a bank. What about confidential supervisory information?
We never ask for it, accept it, or handle it. The measurement needs your question types and an endpoint you designate. Your supervisory file stays yours.
What does the lab stand behind?
Four commitments, in writing. Every report is signed by name and dated. Any factual or methodological error your reviewers find is corrected and the report re-issued at no cost. Errors and omissions coverage is bound before any client work begins. And every technical question your examiners or reviewers raise gets a written answer, on the record. Fees are fixed and never depend on what the measurement finds.
Why is the founding rate half of list?
Because the first three engagements buy something specific: permission to describe the engagement in a short case study, approved by you in writing before a word is published, named if your policies allow it and anonymized at the same rate if they don't. In return you get the $4,500 report, a two-year lock at $4,500 a quarter on Continuous Assurance, and a founder with everything to prove. The founding invoice arrives only with the delivered report, and it carries a written guarantee: if the report gives your governance file nothing usable, say in writing what is missing within fourteen days and the invoice is cancelled. Same instrument, same report as list. One per client, three ever across both price lines, and after three the program closes for good.
What does the measurement not do?
It doesn't decide truth, and it doesn't replace your people. It measures the model's fabrication signals and flags exactly what deserves human eyes. Every report says so on page one.
Who we serve

Segment-specific answers live on their own pages.

Banks and credit unions: independent AI model testing for the model-risk file. Insurers: evidence for your AIS program under OID Bulletin 2024-11. Health systems: validation evidence for Joint Commission AI certification. Law firms: confidential testing built for Opinion 512 duties. Standards programs: model-level evidence for ISO/IEC 42001.

A question we didn't answer?

Email it. Replies come the same day, in writing, and the answer goes in your file, not just your memory.

Email the lab
REPLIES SAME DAY · EVERYTHING ANSWERED IN WRITING, ON THE RECORD