Spectralgraph
AI management · regulatory crosswalk

The NIST AI RMF MEASURE function, subcategory by subcategory.

The MEASURE function asks an organization to evaluate an AI system for trustworthy characteristics. Some subcategories want a measurement of the model itself. Most want something else, and a fabrication instrument has no business near them. This crosswalk walks MEASURE in full and says which is which, including the seven it does not touch.

AI RMF 1.0 IS UNDER REVISION · QUOTE THE TEXT, NOT THE IDENTIFIER · PLAYBOOK WORDING USED WHERE IT IS FULLER
Scope statement. The Spectralgraph LLM Validation Report is an independent screening measurement of fabrication risk in a designated, organization-hosted generative AI (large language model) system. It is not an audit, attestation, examination, or opinion, and it does not certify conformity with the AI RMF or any other framework. This crosswalk identifies, subcategory by subcategory, where the Report provides responsive documentation ("Addressed"), where it can serve as one input among others ("Supporting"), and where the subcategory is outside the Report's scope entirely ("Outside scope"). Subcategory text is quoted from NIST's own AI Resource Center; where the Playbook prints fuller wording than the Core document, the Playbook is quoted and marked.

MEASURE 1: appropriate methods and metrics

MEASURE subcategoryStatusWhere this appears in the Report / basis
1.3: "Internal experts who did not serve as front-line developers for the system and/or independent assessors are involved in regular assessments and updates." Addressed An outside laboratory with no role in building the system, engaged on a stated cadence. Both halves the subcategory names, independence and regularity, are the arrangement itself rather than a byproduct of it.
1.1: approaches and metrics selected for implementation, starting with the most significant risks
1.2: "Appropriateness of AI metrics and effectiveness of existing controls are regularly assessed and updated"
Supporting One selectable approach, with published discrimination behind the choice, and regular re-assessment of that one metric. The selection process, the risk ranking and the control environment belong to the organization.

MEASURE 2: evaluation for trustworthy characteristics

MEASURE subcategoryStatusWhere this appears in the Report / basis
2.3: "AI system performance or assurance criteria are measured qualitatively or quantitatively and demonstrated for conditions similar to deployment setting(s)." Addressed Measurement runs against the endpoint the organization designates, using its own question types. Absolute results do not carry across serving stacks, so a number taken anywhere else describes a different deployment; measuring in the deployment setting is the point rather than a convenience.
2.5 (second sentence): "Limitations of the generalizability beyond the conditions under which the technology was developed are documented." Addressed Per-domain results naming where the deployed model holds up and where it does not, on the organization's own subject matter. The first sentence of 2.5 is broader and appears below.
2.1: "Test sets, metrics, and details about the tools used during TEVV are documented." Addressed The probe library, the metric, the instrument and the method, documented and published with a DOI so a reader can reproduce it rather than take it on description.
2.13: "Effectiveness of the employed TEVV metrics and processes in the measure function are evaluated and documented." Addressed This subcategory asks whether the measurement method is any good, and few can answer it. Published discrimination (0.852 AUC, length-controlled, on models the instrument had never seen), published negative results, and a pre-registered experiment whose thresholds were not met and were reported in full.
2.5 (first sentence): the system "is demonstrated to be valid and reliable"
2.4: functionality and behavior "monitored when in production"
2.6: evaluated regularly for safety risks
2.9: the model is "explained, validated, and documented"
Supporting Fabrication is one dimension of reliability and one safety risk, measured on a cadence, and the validated and documented halves of 2.9 are responsive. A cadence is not production monitoring and is not described as one. "Explained" is outside scope: reading a model's internal signals is not explainability in the regulatory sense, which means reason codes an assessor can follow for a single decision.
2.2: human-subject evaluation
2.7: security and resilience
2.8: transparency and accountability risks
2.10: privacy risk of the AI system
2.11: fairness and bias
2.12: environmental impact
Outside scope Six subcategories that want a different instrument and, in most cases, a different vendor. Fairness and bias in particular is protected-class outcome analysis, which this laboratory does not perform and does not claim, in any framework. On 2.10, our own no-data posture is a fact about us as a vendor and is not an answer to the organization's privacy-risk subcategory.

MEASURE 3 and 4: tracking over time, and feedback

MEASURE subcategoryStatusWhere this appears in the Report / basis
4.3: "Measurable performance improvements or declines based on consultations with relevant AI actors including affected communities, and field data about context-relevant risks and trustworthiness characteristics, are identified and documented." (Playbook wording) Addressed The Playbook assigns this to TEVV actors, asking for baseline quantitative measures and for sensitivity analysis characterizing variance in performance after system or procedural updates. A re-measurement after a change produces that before-and-after from outside the organization. The consultation half is outside scope, and identifying a change is not the same as causing one: measured intervention to improve model accuracy is something this lab tested and could not demonstrate, and it is not sold.
3.1: tracking risks based on "intended and actual performance in deployed contexts"
3.2: settings where risks are "difficult to assess using currently available measurement techniques or where metrics are not yet available"
Supporting A trend record across quarters on the fabrication dimension, in the deployed context. On 3.2, the subcategory contemplates exactly the case where a metric does not yet exist; fabrication measurement is one such metric, with its discrimination published. It is one approach, not the category of them, and emergent risk of every other kind stays with the organization.
4.1: measurement approaches "connected to deployment context(s)"
4.2: results "informed by input from domain experts and relevant AI actors"
Supporting Measuring on the organization's own question types against its own endpoint is what connects a measurement to the deployment context, and flagged answers are verifiable by the organization's own reviewers, whose questions get written answers. The domain-expert consultation is the organization's to run, and we are not its domain experts.
3.3: end-user feedback and appeal processes Outside scope A process the organization operates with its users and impacted communities.

GOVERN, MAP and MANAGE

FunctionStatusBasis
GOVERN · MAP · MANAGE Outside scope GOVERN is policy, roles and accountability, and a measurement is evidence that a control operated rather than the control. MAP is context and risk identification, which the organization does. MANAGE is risk response and mitigation: a findings register feeds it, and this laboratory does not remediate.

Source: NIST AI RMF 1.0 (NIST AI 100-1) and the AI RMF Playbook, read from NIST's AI Resource Center on 2026-09-10. Where the Core document and the Playbook print a subcategory differently, the Playbook wording is used and marked; MEASURE 4.3 is one such case. AI RMF 1.0 is under revision, so subcategory wording and numbering may change and this page will be re-verified against the revised text. Nothing here states or implies that NIST endorses this laboratory, or that any framework requires an outside measurement. See also the ISO/IEC 42001 page and the NAIC AI Risk Evaluation Supplement crosswalk.

Want this mapping against your own program?

Send your question types and the model you host, and the sample report shows exactly what the evidence looks like before you spend a dime.

Email the lab
REPLIES SAME DAY · EVERYTHING ANSWERED IN WRITING, ON THE RECORD