The human-review queue is the next digital divide

When AI sends an ambiguous case to a public official for review, governments should measure who waits, how long, and what happens while citizens wait for a response.

The next digital divide is machine-readable versus exceptions requiring human review and intervention, causing citizens to wait longer to receive a service. Image: Canva

When public agencies say an artificial intelligence (AI) service includes human review, the phrase sounds reassuring. 


The machine handles routine cases, while a person takes over when confidence is low. That is often the right design. 


But it can create a second system hidden behind the first: a fast lane for cases the model recognises and a slow lane for everyone else.


A citizen with variable income, mixed documentation, a non-standard address or a request written in an unfamiliar language variety may not receive a wrong automated answer.


They may receive no answer yet.


Their case is sent to a human queue. If that queue is understaffed or allowed to grow, the citizen simply waits longer.


This is the next digital divide: not online versus offline, but machine-readable versus exception.

A good system can improve the average and still worsen the exception


GovInsider recently reported on ERREKA, an AI system used by the Provincial Council of Gipuzkoa in Spain, to classify incoming administrative documents. 


The system sorts cases into high, medium and low-confidence bands. High-confidence cases are routed automatically, medium-confidence cases go to an officer for one-click approval, and low-confidence cases follow the usual manual process.


Roney Lima do Nascimento, mathematics educator, AI specialist and doctoral candidate in Pure Mathematics at the University of São Paulo. Image: do Nascimento

That is a sensible architecture. It keeps people involved where uncertainty is greater. 


Yet the same design raises a management question that accuracy and automation rates cannot answer: what happens to the cases that leave the fast path?


Did their waiting time fall or rise? Which types of cases accumulate? How many are reassigned? How often does a reviewer overturn the machine’s suggested route? 


A system can make 80 per cent of cases faster and still make the remaining 20 per cent slower, less predictable and more frustrating.


If low confidence is correlated with unusual documentation, language variation, complex household circumstances or cases spanning several agencies, the queue may fill with exactly the citizens who already find government hardest to navigate.

What a human-review dashboard should show


Every public sector AI pilot that refers cases to people should include a human review queue dashboard from day one. 


It does not require a new platform. It requires agencies to measure the part of the workflow that averages usually hide.


The dashboard should show at least six things:


  1. Referral rate: the share of cases sent to a person, broken down by service, case type, language, channel and geography, and by relevant equality characteristics where collection is lawful and appropriate.
  2. Time to first meaningful human action: not only the median, but also the 90th percentile and the age of the oldest unresolved case. A small tail of extreme delays can disappear inside a respectable average.
  3. Reviewer capacity: open cases per qualified reviewer, expected inflow, actual throughput and the proportion of cases waiting for a specialist rather than a generalist.
  4. What happens while the citizen waits: whether a benefit stops, a deadline continues to run or an application can expire. A case described internally as “pending review” may feel externally like a denial.
  5. Reversal rate: frequent changes to the automated route or decision may indicate weak confidence thresholds, poor data, a badly designed form or a model that does not understand the service context.
  6. Abandonment and repeat referral: citizens who withdraw, fail to respond or return with the same unresolved issue should not disappear from the success statistics.

The building blocks already exist


Singapore’s ServiceSG uses ServiceConnect, a system that combines queue management, case management, a statistical dashboard and citizen feedback across hundreds of services.


New Zealand Police used centralised case management and automated routing to reduce a backlog of 5,000 non-emergency reports and cut response times from as long as two weeks to four hours.


These are not identical to AI review queues. But they show that public bodies can make waiting visible, connect it to staffing and improve outcomes when leaders can see where work is accumulating.


The wider governance gap is now clear. The OECD’s 2026 Digital Government Outlook found that 35 of 36 OECD countries use AI somewhere in government. 


Yet only 10 report measuring the financial or non-financial impact of AI use cases, and only eight have citizen feedback or complaint mechanisms for AI in government.


Governments are scaling the technology faster than they are instrumenting the citizen experience around it.

Add a safe-hold rule


One rule matters more than any chart: when an AI system refers a consequential case to human review, adverse consequences should normally pause until that review occurs.


A citizen should not lose a benefit, miss a statutory deadline or be treated as non-compliant simply because the system was uncertain. 


There may be exceptional public safety situations where immediate action is necessary, but the default should be that institutional uncertainty is carried by the institution, not transferred to the citizen.

Start with a 30-day exception audit


Before an AI-assisted service goes live, agencies can run a parallel 30-day exception audit using historical or shadow cases. 


Measure daily referral volume, reviewer capacity, median and 90th-percentile waiting time, reversal rate, and the oldest case. 


Then set a maximum queue age and a clear trigger for adding staff, changing the threshold or pausing the automation.


Public sector leaders do not need another abstract principle. They need a service-level objective for the human loop.


“Human review” is not a complete safeguard unless the queue is visible, staffed, time-bound and connected to a remedy. The next digital divide will not always look like a wrong decision. 


It may look like one citizen receiving an answer in seconds while another is told to wait for a human who never has enough time to arrive.


----------------------------


Roney Lima do Nascimento is a mathematics educator, AI specialist and doctoral candidate in Pure Mathematics at the University of São Paulo. His work focuses on model evaluation, education, institutional capacity and responsible AI governance. He writes in a personal capacity.