Spectralgraph

THE INSTRUMENT ROOM

Beyond the Scorecard

A reliability score is one number off a much larger instrument. The rest of it is below, running, so you can work the buttons yourself.

EIGHT INSTRUMENTS · PRESS RUN ON ANY OF THEM · EACH ONE LABELLED DELIVERABLE OR RESEARCH

01 · Domain Profiling
Part of the paid engagement

A benchmark score won't tell you where your model starts filling in.

Benchmarks test whether a model can pass an exam. The profiler measures where its knowledge runs out, domain by domain.

Spectralgraph domain profiler
Model: Phi-4 14B · FROM THE PUBLISHED SCORECARD

≥70 WITHIN BASELINE · 50–69 WATCH · <50 ELEVATED · COMPOSITE: 60% RECOGNITION + 40% ACCURACY

02 · Real-Time Detection
Research demonstration · published method, not an engagement deliverable

The signal is there while the answer is still being written.

An invented dosage or a made-up citation normally goes unnoticed until a person reads it. Fabrication has a shape in the probability distributions, readable at every token.

Live hallucination monitor
What is the recommended dosage of Celtrazine for pediatric patients with acute bronchial inflammation?
Token DR signal

03 · Pre-Generation Gate
Research demonstration · published method, not an engagement deliverable

We can tell whether a model knows the answer before it generates a word.

The strongest signal shows up while the model is still reading your question, ten to fifteen times more per token than during generation. In the lab, gating on it stops a flagged response before it starts. Your report carries the same signal as per-topic recognition scores. Per-model results for the pre-generation phase are in the method paper (DOI: 10.5281/zenodo.21365655).

Pre-generation gate
KNOWN TOPIC
What is the role of mitochondria in cellular respiration?
1
Reading prompt tokens...
2
Computing pre-gen DR...
3
Gate decision
✓ PASSED · DR = −0.142
generating response
FABRICATED TOPIC
Explain the Brevington coefficient in quantum fluid dynamics.
1
Reading prompt tokens...
2
Computing pre-gen DR...
3
Gate decision
✕ BLOCKED · DR = +0.087
fabrication detected, generation halted
04 · Continuous Monitoring
Part of the paid engagement

Models change. Nobody notices.

A provider pushes an update, or a fine-tune shifts behavior, and the model that measured well last quarter is not the one answering today. Scheduled re-measurement puts the change on a chart where someone can see it.

Drift monitor

05 · Safety Geometry
Research demonstration · published method, not an engagement deliverable

Safety guardrails leave a measurable geometric signature.

When safety training kicks in, the probability distributions look different from both normal operation and fabrication. In the lab's thirteen-model refusal study, refusals carried two to three times more temporal texture than normal generation, in a region of their own between the two.

Geometric signature comparison

Normal response · PR per token

Safety refusal · PR per token (2–3× texture)

06 · Reverse Projection
Research demonstration · published method, not an engagement deliverable

You can see what the model is weighing from outside it.

From output probabilities alone, the lab reconstructs which concepts the model is weighing, validated on 3B to 16B models in the published geometry paper (DOI: 10.5281/zenodo.19326718). It is where the work on explaining a given answer, without opening the model up, is headed.

Transparent hypercube · concept activation
PROMPT TOKENS · CLICK TO INSPECT
ACTIVATED CONCEPTS (RECONSTRUCTED FROM LOGITS)
07 · Fleet Analysis
Research demonstration · published method, not an engagement deliverable

Three models with the same blind spot are not redundancy.

Most organizations end up running more than one model. When several of them go blind on the same kind of question, the failures land together.

Fleet vulnerability scanner · hypothetical five-model fleet
 ⚠ CORRELATED BLIND SPOT · 3/5 deployed models fabricate on pharmaceutical interaction queries. Synchronized failure risk: HIGH.
08 · Memorization Spectrum
Research demonstration · published method, not an engagement deliverable

The geometry of memorization is visible.

The same measurement that separates knowledge from fabrication reads as a continuous spectrum, running from pure fabrication to verbatim reproduction.

Content relationship analyzer
GeneratedFamiliarMemorizedVerbatim

Put this instrument on your model.

One email gets you the complete sample report. Same-day reply, no call required.

Request the sample report
SPECTRALGRAPH · EST. 2026 · ENID, OKLAHOMA
← AI Reliability Index Model Validation → About & Methodology →