Spectralgraph

About

Origin

This came out of ultrasonic non-destructive testing. You fire sound into a piece of steel and read the echoes. Good material sends back a predictable back-wall echo. Anything that deviates from that wall is a flaw.

The same geometry shows up in language models. Token probabilities running through a model's weights produce a wall of their own, and the deviations from it are where the model is making things up. The weights are the material under inspection, the tokens are the sound wave, and the residual is the flaw indication.

“The shape tells you what's inside without cutting it open.”

THE NDT PRINCIPLE THE LAB IS BUILT ON

Joseph Stephens carried that principle from steel to language models, and he signs every report the lab produces. Meet the founder.

What would a perfect model look like?

A model that genuinely knew the answer to everything would sit at the top of this index. Its confidence would track its correctness question after question, and the measurement would see that same pattern in every domain.

Confidence on its own tells you nothing. What matters is whether it moves with correctness. GPT-OSS 120B is one of the largest models we have tested, and it still lands on the ordinary scale. Size does not set the score.

What the scores measure

Every token a model processes produces a probability distribution across its vocabulary, and that distribution has a shape. Retrieval concentrates it, so one or two tokens dominate. Fabrication spreads it, because the model is picking between interchangeable guesses instead of recalling a specific answer. What we measure is that shape. No ground truth, no repeated generations, no access to weights or training data, one pass per prompt.

Each model gets two measurements. Recognition is the signal from the model reading your prompt, before it writes anything. Accuracy is the signal from the tokens it writes back. They are independent, and recognition carries about ten to fifteen times more signal per token, which is why it gets the larger weight.

Reading the scorecard

Confidence

How much probability a model puts on its top choice at each token. Ninety-nine percent on one word is confident. Probability spread across many competing words is not. The measurement compares how that behaves when a model knows the answer and when it doesn't.

Recognition (0–100)

The model's probability patterns while it reads your prompt, before it generates a word. On topics it knows, its predictions across your question form coherent domain vocabulary. On topics it will fabricate, they scatter into unrelated associations. Higher means the model consistently recognizes what it is being asked about.

Accuracy (0–100)

The same patterns while the model writes its answer. Retrieving a known fact, one or two words dominate. Fabricating, many plausible words compete. Higher means the model generates with knowledge geometry.

Two models, DeepSeek V3 and DeepSeek V3.1, show no accuracy score. That is a result in itself. Most models look measurably different when they fabricate. These two answer with the same confidence and fluency whether they are right or wrong, so we cannot separate the signals, because the model doesn't separate them either. Independent benchmarks and user reports say the same thing: DeepSeek's hallucinations are hard to catch because they sound exactly like its factual answers.

Reliability (0–100)

A composite: 60% recognition, 40% accuracy. Where only one measurement is available, reliability equals that one. 50 is the geometric boundary, the point where the probability geometry turns from retrieval to fabrication. Above 50 trends toward factual retrieval. Strong models land between 55 and 80.

Consistency (0–100%)

How stable a model's reliability is inside a domain. A model at 72 with 90% consistency is safer than one at 75 with 50%, because with the second you cannot predict which prompt comes back fabricated.

Tier Breakdown

Reliability broken out by question difficulty: Foundational, Undergraduate, Professional, Expert, Frontier. The shape of the curve matters more than any single number. A model that scores higher on Frontier than Professional gets cautious at the edge of what it knows. One that declines steeply from Foundational to Frontier fabricates more as the questions get harder.

Why smaller models can score higher

This scorecard does not measure how smart a model is. It measures how catchable it is when it makes something up.

A smaller model that hedges or stumbles when it doesn't know something is easy to catch, because the gap between its confident answers and its uncertain ones is wide. A bigger, more polished model that sounds the same either way is harder to catch, and scores lower here for it. Llama 3.1 8B scores above Llama 3.3 70B for exactly that reason: the 70B has been trained to sound confident about everything, which hides its mistakes from this measurement without extended precision.

If factual reliability matters where you are deploying, you want the model whose mistakes stay visible.

The Name

Spectral: the measurement is spectral analysis of probability distributions. Every token carries a geometric signature in its probability spectrum, and we read that spectrum the way an inspector reads an ultrasonic waveform.

Graph: the output is a map of what the model knows and where it fabricates, across domains, difficulty tiers, and time.

Limitations

The index covers 18 open-weight models, all run on the same inference engine, the same probe library, and the same parameters. That matters. Our own testing shows different inference stacks running the same model can produce measurements 8 to 41 times apart in absolute value. Rankings hold, but mixing instruments on one scorecard would mislead you, so we hold conditions constant.

Closed-weight models like GPT-4o, Claude, and Gemini aren't here yet. The method works on APIs; the scorecard doesn't mix instruments. When those runs can happen on the same standardized infrastructure as everything else here, they'll be added.

Each score is the model as it stood on the date printed beside it. Models change: providers push modifications, fine-tunes shift behavior, infrastructure alters output distributions. That is why dated, repeatable measurement is worth anything at all. The index shows the method at scale; tracking one model over time is what client engagements are for.

The Lab

Spectralgraph LLC is an independent AI-testing laboratory in Enid, Oklahoma. If you host your own model, the lab runs private validation engagements: the same instrument behind this index, pointed at your deployment, with signed evidence for your governance file.

Contact

Compliance and risk teams: validation engagements and the sample report

Model providers and AI companies: the vendor evidence line is published, with scope and pricing

Regulators: full methodology disclosure under NDA

contact@spectralgraph.ai

Put this instrument on your model.

One email gets you the complete sample report. Same-day reply, no call required.

Request the sample report
SPECTRALGRAPH · EST. 2026 · ENID, OKLAHOMA
← AI Reliability Index Beyond the Scorecard →