When AI writes the first draft, public bodies still own the record

The next serious public sector AI failure may not look like a machine making a decision, but a helpful summary - quietly corrected once or twice, accepted under pressure, filed, and relied on by everyone who comes after.

An AI-generated summary can be entirely plausible and still change the meaning of the whole record, says UK public sector practitioner, Karl Hopkins.

Most public sector artificial intelligence (AI) governance looks at the end of the process: did the system influence a decision, produce a risk score, or recommend an outcome?

 

At the stage that anyone asks that question, much of the real work may already be done.

 

Public bodies exercise power through the records they create, not only the decisions they announce. A case summary, a chronology or a decision rationale determines what the next person sees as relevant, what looks settled, and what can later be challenged.

 

The record sets the practical boundaries of what can later be decided, reviewed or challenged.

 

That matters because generative AI is likely to enter most public services initially as a drafting tool, not a decisionmaker.

 

It may summarise calls, structure case notes or produce the first version of a report or letter. This usually gets treated as administrative support: a way of saving time on paperwork.

 

In practice, it can become part of how an organisation constructs its own account of events.

The impact of AI-edited information

 

This is not a future scenario.

 

Ambient AI tools that turn clinical consultations into draft notes are already in use. A 2025 physician study found largely negative views on accuracy and style, particularly note length and editing requirements.

 

Karl Hopkins is UK public sector practitioner with more than 17 years' experience across policing, investigations, safeguarding, operational leadership and risk. Image: Hopkins' LinkedIn

Across investigations, safeguarding and operational work, one lesson recurs: decisions are rarely made from every available fact; they are made from what has already been selected, ordered and made usable.

 

Once a summary is accepted into a file, the next reviewer tends to treat it as the starting point, rather than returning to the source.

 

Generative AI sharpens this concern considerably. Hallucination gets most of the attention, but it isn't the only risk.

 

A generated summary can be entirely plausible and still change the meaning of the whole record - smoothing over hesitation or turning genuine uncertainty into confident administrative language.

 

Nothing must be factually wrong for the record to become misleading.

 

A polished draft also looks like it has already been thought through, so under pressure the instinct is to ask whether it reads right - not whether it has dropped context, or merged two accounts into one.

A deeper problem

 

There is a deeper version of this problem that is easy to miss.

 

Even where a mistake is caught and corrected - a reviewer spots it, fixes it, moves on - the correction itself usually leaves no trace. The file looks the same either way: clean, considered, unremarkable.

 

An organisation's own records may struggle to tell the difference between a tool that is genuinely reliable and one that only looks reliable because someone is quietly catching and fixing its mistakes, case after case, without anyone recording that this is happening at all.

 

The Post Office Horizon scandal and Australia's Robodebt programme were not generative AI failures, but both show how a system-produced account can become embedded in an official process and acquire authority before it has been properly interrogated.

 

The person affected is then left trying to disprove an account they did not create, often without being able to reconstruct how it was produced.

 

None of this is an argument against AI-assisted drafting.

 

Administrative overload is real, and well-designed tools can reduce delay and improve consistency.

 

It is an argument for governing high-impact drafting as record-making, rather than ordinary office automation.

Four safeguards to help public agencies to be accountable

 

Four safeguards would make a real difference, and each answers the same underlying question: if something were quietly going wrong, would the organisation actually be able to see it?

 

First, preserve the source. A generated summary should never replace the underlying interview, call, document or assessment, so the result can always be checked against it.

 

Second, preserve the draft. Where a record may affect rights, services, benefits, care or legal status, the organisation should be able to show what the AI produced before it was edited.

 

Logging that AI was used is not enough if nobody can tell afterwards which wording was the machine's and which was the reviewer's.

 

Third, make corrections visible, not just the fact that a review took place. Version history should show what was accepted, changed or rejected, and by whom. However, that only covers one file.

 

Substantive corrections; where a draft omits important context, distorts meaning or introduces unsupported content, should also be logged as AI near misses.

 

This allows organisations to identify recurring problems across cases. Otherwise, careful staff may quietly prop up a weak system without the organisation recognising it.

 

Fourth, link key statements back to their source. Timestamps or short references let a reviewer test whether the wording fairly reflects the original material, especially where it draws conclusions about someone's behaviour.

 

These controls cannot sit with the individual user alone. If an organisation authorises AI-assisted drafting, it also owns the risk the workflow creates.

 

A supplier cannot carry a public duty, and a frontline practitioner should not be left to absorb institutional risk personally.

 

Ownership must sit with senior leaders: which records may be AI-assisted, what must be retained, what checking is proportionate, and when use must be disclosed.

 

Assurance functions have a role too. Not just testing average performance, but measuring correction burden: how often staff restore missing context, rewrite conclusions, or reject the output altogether.

 

A low recorded error count says little if quiet corrections are not being counted.

 

The most useful thing a public body could do tomorrow is simple: pick one AI-assisted drafting use case and map the record - from source to final approval.

 

Ask what is retained, what could be reconstructed if challenged, how often staff materially correct the output, and whether the person affected could contest the result.

 

The next serious public sector AI failure may not look like a machine making a decision. It may look like a helpful summary, quietly corrected once or twice, accepted under pressure, filed, and relied on by everyone who comes after.

 

By the time a formal decision is made, the version of events that shaped the outcome may already have been written - and nobody may be able to say how safe that process really was.

 

Karl Hopkins writes in a personal capacity and his writing does not reflect the views of his employer or GovInsider.

 

---------------------------------------

 

The author is a UK public sector practitioner with more than 17 years' experience across policing, investigations, safeguarding, operational leadership and risk. He holds a First-Class BSc (Hons) in Applied Criminal Justice and writes on AI governance, organisational learning and the integrity of public-sector decision-making.