Protocols and standards

NIST AI RMF

The NIST AI Risk Management Framework (AI RMF 1.0) is voluntary guidance published by the United States National Institute of Standards and Technology in January 2023 for identifying, measuring and managing the risks of AI systems across their life cycle. Its core organises that work into four functions — GOVERN, MAP, MEASURE and MANAGE — and, unlike a management-system standard, it carries no conformity or certification scheme, so an organisation adopts it and evidences its own adoption rather than being certified against it.

also called NIST AI Risk Management Framework · AI RMF 1.0 · NIST AI 100-1 · GOVERN MAP MEASURE MANAGE

The framework was directed by the National Artificial Intelligence Initiative Act of 2020, developed through open consultation with published drafts and comment periods, and released as NIST AI 100-1 on 26 January 2023. It is deliberately voluntary, sector-neutral and use-case agnostic, and it contains no requirements clauses. Its influence comes from two places instead: it is the vocabulary US federal agencies, insurers and boards already have, and it is named in legislation and procurement language as an acceptable risk-management framework — Colorado’s AI legislation is the example usually cited. That status has a consequence worth stating early, because vendors blur it: there is no such thing as being certified to the AI RMF, and ‘NIST compliant’ is not a status any NIST document confers.

Part 1 sets out seven characteristics of trustworthy AI: valid and reliable, safe, secure and resilient, accountable and transparent, explainable and interpretable, privacy-enhanced, and fair with harmful bias managed. Validity and reliability are treated as the foundation — a system that does not work does not become trustworthy by being fair about it — and NIST is explicit that the characteristics trade off against one another, that resolving those trade-offs is a contextual human judgement rather than an optimisation, and that a system can be technically sound and still unacceptable in a particular use.

Part 2 is the Core. GOVERN is cross-cutting rather than sequential: culture, policy, accountability, workforce competence and third-party risk, running through the other three functions continuously instead of preceding them once. MAP establishes context — intended purpose, who is affected, what the system touches, which risks the context implies, and where the boundaries of the deployment actually sit. MEASURE analyses, benchmarks and tracks the risks MAP identified, with an explicit instruction to document what could not be measured rather than to leave it out. MANAGE allocates resources against the mapped and measured risks — treat, transfer, avoid or accept — and monitors after the action is taken. The functions are broken into categories and subcategories (19 and 72 in version 1.0), and the companion Playbook suggests actions for each. Profiles are the mechanism for applying the framework: a use-case profile describes the framework as applied to a specific application, and a current-and-target profile pair turns the framework into a gap analysis.

The companion most agent teams actually need is the Generative AI Profile, NIST AI 600-1, published on 26 July 2024. It names twelve risks unique to or amplified by generative AI — confabulation, information integrity, information security, data privacy, harmful bias and homogenisation, human-AI configuration, intellectual property, value chain and component integration, environmental impact and others — and maps suggested actions for each onto the four functions, which makes it far more directly usable for an agent estate than the base framework’s deliberately abstract categories.

Mapped against a runtime control, three of the four functions land squarely and one thins out, and knowing which is which prevents an overclaiming mapping table. GOVERN, MAP and MANAGE are about accountability, context and action: a registry, versioned policies with named owners, budgets, approval gates and a stop control all produce artefacts against them. MEASURE is largely about evaluating models and outcomes, and anything sitting in the request path sees traffic rather than model behaviour — spend, policy match and block rates, and shadow-mode findings are measurable there, while validity, bias and explanation are not, and come from an evaluation stack instead. The framework’s own answer to that is better than a tick: record the risk as mapped, the measurement as not currently feasible, the reason, and the person who accepted it. That instruction — document what you could not measure — is the discipline most control mappings abandon first, and it is the one that makes the rest of the mapping credible.

in practice

A MEASURE row that says nothing, honestly

A team maps its agent estate to the four functions. GOVERN, MAP and MANAGE fill in from the registry, the versioned policy set and the kill switch. Under MEASURE they can produce month-to-date spend by agent, policy match and block rates, and shadow-mode findings for every rule not yet enforcing. They cannot produce anything about whether the models’ answers were correct, because nothing in the request path evaluates them and no evaluation suite is wired up. The framework’s answer is not to leave the row blank and it is certainly not to fill it with a proxy metric that measures something easier. It is to record the risk as mapped, the measurement as not currently feasible, the reason it is not, and the named person who accepted that. A row reading ‘not measured, here is why, here is who owns it’ is a working risk register. A row with a tick in it is a document that will not survive its first audit.

not the same as

What nist ai rmf is routinely confused with

ISO/IEC 42001
The AI RMF is voluntary guidance with no audit and no certificate; ISO/IEC 42001 specifies a management system an accredited body certifies against. They crosswalk cleanly enough that many organisations use both — the RMF to structure the thinking, 42001 to produce the artefact a customer will accept — but only one of them produces something you can put on a supplier questionnaire.
NIST Cybersecurity Framework (CSF 2.0)
Same institute, similar function-based shape — CSF 2.0 runs GOVERN, IDENTIFY, PROTECT, DETECT, RESPOND, RECOVER — but a different subject. An AI RMF profile is not a CSF profile, they are separate documents with a published crosswalk between them, and a CSF programme does not cover AI-specific risk simply because both start with GOVERN.
NIST SP 800-53
800-53 is a control catalogue that is mandatory for US federal information systems under FISMA, with a defined assessment process. The AI RMF mandates nothing and assesses nothing. Conflating them produces the phrase ‘NIST compliant’, which describes neither document accurately.
next

Related terms

ISO/IEC 42001

ISO/IEC 42001:2023 is the international standard specifying requirements for an artificial intelligence management system — the governance structure, processes and records an organisation puts in place to develop or use AI responsibly. It is the first AI standard an organisation can be certified against by an accredited certification body, and, like ISO/IEC 27001, it certifies a management system within a declared scope rather than any product, model or software.

EU AI Act

The EU AI Act is Regulation (EU) 2024/1689, which regulates AI systems placed on the market or used in the European Union in proportion to the risk they present, and which places materially different duties on the organisation that builds a system (the provider) and the organisation that uses it under its own authority (the deployer). It entered into force on 1 August 2024 and applies in stages: the prohibited practices from 2 February 2025, the general-purpose AI model obligations from 2 August 2025, and most remaining obligations, including those on high-risk systems listed in Annex III, from 2 August 2026.

OWASP Top 10 for LLM Applications

The OWASP Top 10 for Large Language Model Applications is a community-maintained list of the ten most significant security risks in applications built on large language models, published by the OWASP GenAI Security Project. It is an awareness and prioritisation document rather than a standard: nothing certifies against it, and its ranking comes from consensus among contributing practitioners rather than from measured incident data.

Protocols and standards

The terms next to this one

The substrate everything here conforms to, and the four regulatory instruments that decide what evidence an operator has to be able to produce.

get in touch

Definitions are the easy part.

The glossary is written to be useful whether or not you ever buy anything. If you have got to the point of deciding how to implement one of these in your own estate, say what your agents do and you will get a straight answer about what it would actually take.

no form · no qualification step · no sales desk · the other three ways in