Eighteen open-weight models, run through the same instrument on the same questions, across 15 domains from foundational to frontier difficulty.
PROOF OF METHOD · 18 OPEN-WEIGHT MODELS · ONE INSTRUMENT, ONE SET OF CONDITIONS · EVERY SCORE PUBLISHED
FOR REGULATED ORGANIZATIONS
We run this same measurement on the model you host and hand you signed, dated evidence for your governance file. No customer data, nothing installed.
Model Validation →
Reliability: 0–100, composite of recognition and accuracy scores · Recognition: pre-generation signal · Accuracy: generation signal
AI Reliability Index: 18 open-weight models, 15 regulated domains, one instrument. The interactive view loads with JavaScript; the scores below are the same data.
#
Model
Provider
Params
Reliability
Recognition
Accuracy
1
Phi-4 14B
Microsoft
14B
75.8
65.4
91.2
2
Qwen 3 8B
Alibaba
8B
72.2
64.2
84.2
3
Llama 3.1 8B
Meta
8B
71.0
61.5
85.1
4
GPT-OSS 20B
OpenAI
20B
70.5
56.9
91.0
5
Mixtral 8x7B
Mistral AI
8x7B
70.2
63.2
80.6
6
DeepSeek V3.1 †
DeepSeek
671B
69.5
69.5
—
7
Mistral 7B
Mistral AI
7B
69.2
63.7
77.6
8
DeepSeek V3 †
DeepSeek
671B
68.4
68.4
—
9
Qwen 2.5 7B
Alibaba
7B
67.8
62.3
76.0
10
Command R 35B
Cohere
35B
67.6
56.6
84.1
11
GPT-OSS 120B
OpenAI
120B
66.8
58.8
78.9
12
Qwen 2.5 14B
Alibaba
14B
66.6
61.8
73.7
13
Qwen 2.5 32B
Alibaba
32B
65.8
61.4
72.3
14
Gemma 3N E4B
Google
E4B
65.5
63.7
68.3
15
Yi 1.5 34B
01.AI
34B
65.1
58.1
75.6
16
Llama 3.3 70B
Meta
70B
64.9
62.5
68.6
17
Mistral Small 24B
Mistral AI
24B
63.7
58.4
71.8
18
Gemma 2 27B
Google
27B
58.1
51.4
68.2
† Recognition-only: no writing-phase (accuracy) score is reported for this model; its reliability is the recognition score alone. Methodology and the reason are on the About page.
Scoring note: DeepSeek V3 and DeepSeek V3.1 carry no writing-phase (accuracy) score; their reliability is the recognition score alone, not the 60/40 composite the other sixteen models use. The reason is on the About page.
How to cite this index
Spectralgraph LLC. AI Reliability Index: independent fabrication-risk measurements of 18 open-weight language models across 15 regulated domains. Enid, Oklahoma, 2026. https://spectralgraph.ai/scorecard.html (include your access date; scores reflect the measurement run shown on the page). Free to cite with attribution; questions to contact@spectralgraph.ai. The measurement method behind the index is published: Logits Don't Lie (DOI: 10.5281/zenodo.21365654).
Put this instrument on your model.
One email gets you the complete sample report. Same-day reply, no call required.