For AI assurance panels and model risk reviewers

A model name is not a provenance answer, and neither is a licence page.

What you need on file is where the weights came from, what can and cannot be substantiated about their training material, whether the model can be replaced without rebuilding your process, and what happens the day it changes underneath something you approved. This page answers those four, including where the honest answer is that nobody can substantiate it.

You are asked to record an origin that nobody in the chain can evidence

An assurance panel is asked to record where a model came from. What arrives is a name, a version string, a marketing page and a licence. None of those is an origin, and the reviewer is left recording a name in a field that asked for a lineage.

Training material is the part where honesty gets scarce. Weights published by a laboratory rarely come with an enumerable corpus, and where a description exists it is a summary written by the party with the most to lose from precision. A reviewer who writes anything definitive about what a general-purpose model absorbed is recording somebody’s marketing as a finding.

Then the ground moves. A hosted model is updated by its provider, the version identifier may or may not change, and behaviour that was evaluated in March is not the behaviour running in September. A panel that approved an assessed behaviour has approved something that no longer exists, and frequently has no mechanism that would tell them.

Licensing sits underneath all of it and is examined least. Weights carry terms, terms carry use restrictions, restrictions carry indemnity consequences, and the organisation carrying the risk of a violation is the one that deployed it rather than the one that published it. Very few review packs contain the actual terms.

And the largest omission is usually what happens when the model is wrong or unavailable. A panel spends its attention on which model was chosen and almost none on what the process does when the model returns nothing, returns nonsense, or is unreachable — which is the case that determines whether an outage becomes an incident.

Replaceability is a stronger answer than a provenance claim

No party in the chain can substantiate the complete training material of a general-purpose model — so an assurance position built on a provenance claim rests on somebody’s marketing, and the durable position has to be built on substitutability and evidence instead.

The structural answer here is that the model is a replaceable component behind a stable interface rather than a dependency the process is built around. Changing which model runs a step does not change what the step is, what it is permitted to do, or what evidence it leaves. That is what makes a provenance concern actionable: if a model becomes unacceptable to your panel, it is substituted rather than argued about.

You may bring your own provider, using credentials your organisation holds, at no markup. That converts the provenance question into one you already answered internally — if your organisation has approved a provider, that approval carries here, and the model relationship is between you and them rather than mediated by us. For many panels this is the whole resolution, because the review already exists.

Beneath all of it is a deterministic floor that uses no model at all. The operating decisions that matter — routing, scheduling, ordering, thresholds, the rules that govern what happens next — are computed rather than generated, and they continue to work when no model is available. That is why the failure case has an answer: a model that is unreachable or returns nothing degrades the output rather than stopping the operation, and no decision that matters waits on a generated response.

What is recorded, per output, is which model produced it and when. That is the artefact a panel actually needs and rarely gets: not a claim about what a model absorbed years ago, but a record that lets you go back afterwards and say precisely which component produced a particular result. A provenance claim you cannot verify is worth less than an attribution record you can read.

On our own learning, the standing rule is narrow and worth stating because the question is asked. What improves the system is the structure of our own work — patterns, classes, counts — and a customer’s content is not training material for a general model, is not pooled across organisations, and a provider’s outputs are never harvested to imitate that provider. Where a customer wants none of it, that is a setting rather than a negotiation.

What a panel should be able to put in the file

Which model produced a specific output — measured by the attribution record, readable for any past result rather than reconstructed.

Whether a model can be substituted without rebuilding — measured by substituting one in a trial and confirming the step and its permissions are unchanged.

Whether your own approved provider can be used — measured by connecting your credentials and confirming no markup is applied to their charges.

Behaviour when no model is available — measured by disabling the provider in a trial and confirming the operation continues on computed rules.

What of yours is used to improve anything — measured by the written statement of what is captured, and the setting that disables it entirely.

Whether a change would be visible — measured by comparing attribution records across periods, from your own copies.

no claim to enumerate the complete training material of any third-party model, because no party in the chain can. No claim that a model will not change under its provider. No accuracy figure is quoted, because none has been measured for your operation. No claim that generated output is correct — the deterministic floor is what the operating decisions rest on. An independent SOC 2 Type II attestation is in progress and no report exists yet.

Bringing your own provider, and what that changes

A provider your organisation has already approved can be connected using credentials you hold and can revoke. Charges go to your account at their rates with nothing added, which matters to a panel for a reason beyond cost: it means the relationship, the terms and the data handling are the ones your organisation already reviewed, rather than a second arrangement to review.

Where you use ours instead, the model in use is named and can be changed. Nothing in the process is written against a particular provider’s behaviour, which is the property that makes substitution real rather than theoretical — a system tuned to one model’s quirks is a system that cannot actually replace it.

The attribution record goes to your log stream along with everything else, so the question of which model produced which output is answerable from artefacts you hold. A panel relying on a supplier to produce a report about their own model usage is in the weakest available position when that question is asked seriously.

What cannot be substantiated, said plainly

Nobody in this chain can enumerate the complete training material of a general-purpose model, and a supplier claiming to has described a summary produced by an interested party. That is stated here because a panel writing anything definitive in that field is recording marketing as a finding, and the honest entry is that the origin is not substantiable at that level.

What is substantiable is everything downstream of it: which model ran, when, against what, producing what, under which permission, with what recorded afterwards. A panel that shifts its assurance weight from an unverifiable origin claim to a verifiable attribution record has a stronger file, not a weaker one.

On learning from your material, the rule is narrow and it is a rule rather than a preference. What improves the system is the structure of our own work — patterns and classes rather than content. A customer’s records are not training material for a general model and are not pooled between organisations. A provider’s outputs are never collected in order to imitate that provider, which is both a terms question and a position we hold independently of terms.

No AI certification, seal or assessment is held or claimed. A SOC 2 Type II attestation is in progress and no report exists yet; it is an attestation with a scope and a period rather than a certification. Where an assurance framework requires an independent model assessment as a condition, that condition is not satisfiable today.

What to put in the file, and what to stop asking for

Stop requiring a training-material attestation from a supplier who cannot produce one honestly. The field will be filled either way, and the version that gets filled is a summary somebody wrote for a different purpose. Requiring an attribution record instead produces something a reviewer can actually check.

Record the substitution property explicitly, because it is what converts a future provenance objection from a crisis into a change. A panel that has confirmed a model can be replaced has bounded the consequence of every subsequent concern about that model.

Test the no-model case in a trial. Disable the provider and watch what the operation does. That single test tells you whether generated output is decorating a computed process or whether it is the process, and the two carry entirely different risk.

And decide the learning setting deliberately rather than by default. It is a setting, the written statement of what is captured is available before you decide, and an organisation that switches it off should see no change in the operating behaviour.

Start with the substitution test and the failure test

A trial in which your team substitutes one model for another and then disables the provider entirely, observing what changes and what does not, before any model is formally reviewed.

Those two tests settle more of a model risk file than a provenance questionnaire does, because they establish the bound on every future concern. If a model can be replaced and the operation survives its absence, no single model becomes a dependency your organisation cannot exit.

Where your organisation has already approved a provider, connecting it is the shortest path and it reuses a review you have already paid for. That is worth doing before anything else, because it may remove the model question from this procurement entirely.

The attribution record is the artefact to ask for next, and to check against a past period rather than a demonstration. A record that only exists going forward answers the easy version of the question.

Questions buyers actually ask

What was the model trained on?

For any third-party general-purpose model, nobody in the chain can enumerate that honestly, including us. What exists is a description written by the party with the most to lose from precision, and a reviewer who records it as a finding has recorded marketing. We will not supply one. What we will supply is everything downstream that is checkable — which model ran, when, on what, producing what, under which permission — and the ability to replace the model entirely if your panel finds it unacceptable. An attribution record you can verify is worth more to your file than an origin claim nobody can.

The provider changes the model underneath us and we would never know.

That is a real risk with every hosted model and the mitigation is structural rather than contractual. Each generated output carries a record of which model produced it and when, and that record lands in your log stream, so a comparison across periods is something your own team can perform from artefacts you hold. And because the model is a replaceable component behind a stable interface, discovering that behaviour moved is a substitution rather than a rebuild. The alternative — a promise that nothing will change — is one no supplier can keep.

Can we use the provider we have already approved?

Yes, with credentials your organisation holds and can revoke, billed to your account at their rates with nothing added by us. That is frequently the fastest resolution to this whole section, because the review already exists inside your organisation and carries over. It also changes the shape of the risk: the terms, the data handling and the relationship are the ones you assessed, rather than a second arrangement you would have to assess.

What happens when the model is wrong or unavailable?

The operating decisions do not rest on it. Routing, scheduling, ordering, thresholds and the rules governing what happens next are computed rather than generated, so an unreachable or empty model degrades the output rather than stopping the operation. Test that directly in a trial by disabling the provider and watching what continues — it is the single most informative test on this subject, because it distinguishes a process that uses generated text from one that depends on it, and those carry entirely different risk.

Are you training on our data?

Your records are not training material for a general model and are not pooled between organisations. What improves the system is the structure of our own work — patterns, classes and counts rather than content — and a provider’s outputs are never collected in order to imitate that provider, which is a position held independently of anybody’s terms. Where your organisation wants none of it, that is a setting rather than a negotiation, the written statement of what is captured is available before you decide, and switching it off should produce no change in operating behaviour.