Spectralgraph

AI Reliability Index

Eighteen open-weight models, run through the same instrument on the same questions, across 15 domains from foundational to frontier difficulty.

PROOF OF METHOD · 18 OPEN-WEIGHT MODELS · ONE INSTRUMENT, ONE SET OF CONDITIONS · EVERY SCORE PUBLISHED

FOR REGULATED ORGANIZATIONS
We run this same measurement on the model you host and hand you signed, dated evidence for your governance file. No customer data, nothing installed. Model Validation →

Reliability: 0–100, composite of recognition and accuracy scores · Recognition: pre-generation signal · Accuracy: generation signal

AI Reliability Index: 18 open-weight models, 15 regulated domains, one instrument. The interactive view loads with JavaScript; the scores below are the same data.
#ModelProviderParamsReliabilityRecognitionAccuracy
1Phi-4 14BMicrosoft14B75.865.491.2
2Qwen 3 8BAlibaba8B72.264.284.2
3Llama 3.1 8BMeta8B71.061.585.1
4GPT-OSS 20BOpenAI20B70.556.991.0
5Mixtral 8x7BMistral AI8x7B70.263.280.6
6DeepSeek V3.1 DeepSeek671B69.569.5
7Mistral 7BMistral AI7B69.263.777.6
8DeepSeek V3 DeepSeek671B68.468.4
9Qwen 2.5 7BAlibaba7B67.862.376.0
10Command R 35BCohere35B67.656.684.1
11GPT-OSS 120BOpenAI120B66.858.878.9
12Qwen 2.5 14BAlibaba14B66.661.873.7
13Qwen 2.5 32BAlibaba32B65.861.472.3
14Gemma 3N E4BGoogleE4B65.563.768.3
15Yi 1.5 34B01.AI34B65.158.175.6
16Llama 3.3 70BMeta70B64.962.568.6
17Mistral Small 24BMistral AI24B63.758.471.8
18Gemma 2 27BGoogle27B58.151.468.2

† Recognition-only: no writing-phase (accuracy) score is reported for this model; its reliability is the recognition score alone. Methodology and the reason are on the About page.

Scoring note: DeepSeek V3 and DeepSeek V3.1 carry no writing-phase (accuracy) score; their reliability is the recognition score alone, not the 60/40 composite the other sixteen models use. The reason is on the About page.

How to cite this index

Spectralgraph LLC. AI Reliability Index: independent fabrication-risk measurements of 18 open-weight language models across 15 regulated domains. Enid, Oklahoma, 2026. https://spectralgraph.ai/scorecard.html (include your access date; scores reflect the measurement run shown on the page). Free to cite with attribution; questions to contact@spectralgraph.ai. The measurement method behind the index is published: Logits Don't Lie (DOI: 10.5281/zenodo.21365654).

Put this instrument on your model.

One email gets you the complete sample report. Same-day reply, no call required.

Request the sample report
Model Validation → Beyond the Scorecard → About & Methodology →