For innovation leads and programme sponsors

A pilot authority is for work whose outcome is genuinely uncertain. Not for work you would rather not compete.

These authorities exist because some questions cannot be specified precisely enough to compete, and answering them is worth money. That is a real category and we would like to be in it. It is not the same category as a capability you have already decided to buy — using an innovation authority for that produces a finding, and the finding lands on your programme rather than on the vendor.

A budget for trying things, and a system built to buy things you already specified

The ordinary acquisition machinery assumes you can describe what you want precisely enough to compete it. That assumption holds for most purchases and fails completely for the work an innovation office exists to do, where the honest requirement is a question rather than a specification.

So special authorities exist, and using one correctly is genuinely hard. The scope has to be uncertain enough to justify the route, bounded enough to be evaluable, and structured so that a result — including a negative one — is worth what it cost.

Meanwhile the pressure is to produce successes. An innovation office judged on how many pilots became programmes will select for pilots that were going to work, which are precisely the ones that did not need the authority. The best possible outcome of a real experiment is sometimes a clear negative, and almost no reporting structure treats it that way.

And pilots go where pilots go: a demonstration, a favourable write-up, a stakeholder briefing, and then a cliff, because nobody planned the acquisition path that would follow success. The capability works and cannot be bought, and eighteen months later the office runs another pilot.

The vendors do not help. An innovation authority looks to a seller like a fast door with less competition, so the applications are dominated by companies proposing capabilities they already sell, described as research.

The question that justifies the authority is usually never written down

A pilot is only legitimately a pilot if there is a stated question whose answer is genuinely unknown, and most pilots begin with a capability rather than a question — which makes the authority the wrong instrument and the result unevaluable.

That is the failure that produces both problems at once. Without a written question there is no criterion for success, so a negative result cannot be recognised as a result, and there is no basis on which to shape the follow-on acquisition.

So the first act is to write the question and the criterion, before choosing an instrument at all. What is unknown. What would count as an answer. What would count as a clear negative. What decision the answer changes. If those four cannot be written in a paragraph, the work is probably not a pilot and the honest route is an ordinary procurement.

That test disqualifies most of what gets proposed to innovation offices, including things we would like to sell. A bounded observation scope that measures where a flow actually breaks IS a genuine question in most organisations — nobody knows the answer and the answer changes a decision. A deployment of a capability somebody has already chosen is not, whatever it is called.

And the follow-on path belongs in the design rather than after it. Before a pilot starts, the agency should be able to say what instrument would carry a positive result, on what timeline, with what money. A pilot with no plausible follow-on is a demonstration.

What an innovation office would expect, and how it would check

Pilots that begin with a written question — measured by count of pilots whose stated question and success criterion existed before the instrument was chosen.

Negative results recognised as results — measured by count of pilots concluded as clear negatives, which in most offices is zero and should not be.

Pilots with a stated follow-on path — measured by count of pilots that named the instrument, timeline and money for a positive result before starting.

Time from result to a decision — measured by elapsed days from the criterion being met or missed to a documented sponsor decision.

Proposals that are products described as research — measured by count screened out by the written-question test, before evaluation effort is spent.

Cost of learning a specific answer — measured by money spent per stated question answered, counting clear negatives as answers.

that a pilot authority is a route around competition, that a successful pilot creates any entitlement to a follow-on award, or that we should receive one. A pilot answers a question; the acquisition that follows is a separate decision made competitively where the rules require it, and a vendor treating a pilot as a foot in the door is proposing your finding.

What a pilot has to produce to be worth its authority

A measurement the agency keeps. Whatever the result, the observation and the data belong to your programme in an exportable form, so a negative result still leaves the office holding something — which is what makes a negative acceptable to report.

A written criterion applied honestly. The criterion is set before the work and is not revised afterwards to accommodate what happened. A moved criterion is how a negative becomes a success in a briefing, and it destroys the value of every future pilot the office runs.

A stated follow-on instrument. Before starting, the agency should name what would carry a positive result. If nothing plausible exists, that is worth knowing before spending the money rather than after the demonstration.

And no entitlement. A pilot creates no claim on a follow-on award, we do not treat it as one, and where the follow-on must be competed we expect to compete for it like anybody else.

The uncomfortable half of this page

The uncomfortable half is that the written-question test disqualifies work we would like to be paid for. A capability you have already decided you want is not a pilot, and proposing it as one would be helping you create a finding in exchange for a contract.

The second uncomfortable half is the criterion. Setting it in advance means we can lose — the measurement can come back saying the constraint is somewhere we cannot help with, and that is a legitimate and reportable outcome rather than a failure to be reframed.

Everything produced belongs to the agency and is exportable, including a negative result and the data behind it. A pilot whose data leaves with the vendor has produced a story rather than a finding.

And on assurance: an independent SOC 2 Type II attestation is in progress and no report exists yet. We hold neither a FedRAMP authorization nor a FedRAMP certification and claim no equivalency, and a pilot authority does not change that — an authority governs how work may be bought, not what security position a vendor holds.

The contracting officer, the sponsor, and whoever will read this later

The contracting officer’s question is whether the scope genuinely belongs inside the authority. The written question and criterion are the evidence for that, and they should exist before the instrument is chosen rather than be drafted to justify one already selected.

The sponsor’s question is what decision the result changes. If no decision changes on either outcome, the pilot is a demonstration and the money is better spent elsewhere.

An inspector general or an auditor reading this in two years will ask why this route rather than a competition. The answer needs to be a document written at the start, not a recollection — and a pilot with a stated negative criterion is far easier to defend than one where success was the only defined outcome.

And the assurance position is unchanged by the instrument: no authorization held, no equivalency claimed, deployment inside a boundary the agency already authorized where that route applies.

Write the question first; choose the instrument second

One written question with a success criterion and a negative criterion, authored by your sponsor before any instrument is selected — and a bounded observation scope that would answer it.

Writing the question costs nothing and eliminates most candidate pilots, including some of ours. If the four sentences cannot be written — what is unknown, what counts as an answer, what counts as a clear negative, what decision changes — the work is not a pilot.

Where they can be written, the bounded observation scope this architecture recommends everywhere is usually the shape of the answer: measure where the flow actually breaks, which is genuinely unknown in most organisations and changes a real decision.

And name the follow-on instrument before starting. If nothing plausible would carry a positive result, that is the finding, and it is cheaper to have it now.

Questions buyers actually ask

We want to use this authority because a competition would take too long.

That is the reason not to use it. An innovation authority applied to a capability you have already chosen produces a finding, and the finding lands on your programme rather than on the vendor who encouraged it. If the requirement is specifiable and you have decided what you want, the honest route is a procurement — slower, and defensible in two years when somebody asks. We would rather compete for that than win a pilot that becomes your problem.

Our office is measured on pilots that convert. A defined negative criterion is career risk.

It is, and that is a real problem with how innovation offices are measured rather than a reason to avoid the criterion. The practical mitigation is to define the negative in terms of what the agency LEARNS rather than what the vendor fails to deliver: "we will learn where the constraint actually sits" is true on either outcome, and a negative on the capability is still a positive on the question. If your reporting structure cannot accommodate that, it is worth naming as a constraint on the office rather than absorbing it into scope design.

What if the pilot succeeds and we cannot buy the thing afterwards?

Then the pilot was a demonstration and the money bought a briefing, which is why the follow-on instrument belongs in the design rather than after it. Before starting, your acquisition office should be able to name what would carry a positive result, on what timeline, with what money. If nothing plausible exists, that is worth discovering before the work rather than eighteen months later — and it is a legitimate reason not to run the pilot at all.

Every vendor tells us their product is research.

Which is why the test is a document rather than a conversation: what is unknown, what counts as an answer, what counts as a clear negative, what decision changes. Apply it to every proposal including ours. A capability with a known outcome cannot produce four honest sentences, and the failure is visible on the page rather than in a meeting. It will disqualify a lot of what reaches your office, and some of what we would like to propose.