Government AI pilots need a graduation test before they scale

Indonesia is building a bridge from public sector AI pilots to real adoption. The next step is to define, before the demo, what evidence a system must produce before public money follows it into scale.

Every public sector AI pilot should begin with a graduation test: a short, prewritten statement of what the system must demonstrate, what would count as failure, and what evidence must exist before the agency can expand its use. Image: Canva

Public sector artificial intelligence (AI) has a demo problem. A pilot can look impressive without proving that it deserves a budget line. 

 

GovInsider's recent reporting from Indonesia makes the opportunity clear. Through the AI Incubation for Public Sector programme, ministries acted as problem owners while five innovators tested solutions against real government workflows.  

 

The programme is designed to create a clearer route from experimentation to institutional adoption.

 

And this bridge matters. A separate GovInsider’s report on the UK-Indonesia initiative captured the recurring failure mode: pilots often end as demonstrations because ownership, adoption and procurement do not survive the proof-of-concept stage. 

 

The better approach is to make pilots easier to judge, with approval tied to declared evidence rather than presentation quality. 

 

Every public sector AI pilot should begin with a graduation test: a short, prewritten statement of what the system must demonstrate, what would count as failure, and what evidence must exist before the agency can expand its use. 

Define success before the vendor arrives 

 

A pilot should start with the public service baseline. How long does the current process take? How many officers does it require? Where do errors occur? What does a citizen experience when the process fails?  

Burak Oktenli is an independent researcher working at the intersection of AI governance, national security, and critical infrastructure resilience.
 

Without that baseline, a faster AI demo can look transformative even when it moves the work somewhere else. 

 

The graduation test should name the outcome that matters.  

 

For a document-processing tool, that may be turnaround time plus error rates and officer workload.  


For a citizen-facing assistant, it may be successful case completion, escalation quality and accessibility.  

 

For a fraud-screening system, it may include both detection and the cost imposed on legitimate users. 

 

Singapore's Ministry of Digital Development and Information (MDDI) has already said that public service AI should be measured for productivity gains, service improvements and risks, while acknowledging that maturity remains uneven across agencies.  

 

That is the right direction. A pilot should be designed around those measures rather than adding them after a favourable demonstration. 

Test the conditions that could change the decision 

 

Model accuracy is rarely enough. Government systems operate through data feeds, permissions, interfaces, staff procedures and external services.  

 

A pilot can succeed while all of those conditions are unusually clean. 

 

The graduation test should therefore include representative failure conditions.  

 

What happens when data arrive late, a source is missing, the model sees a case outside its development set, an officer disagrees with the recommendation, or an external service is unavailable? If the system can act, what stops it from exceeding its assigned authority? 

 

A successful result should identify the exact configuration that produced it. Software version, model version, material data dependencies, system prompts or policies, and relevant external services should be recorded.  

 

A result for one configuration should not silently become evidence for a materially different one after an update. 

 

This is also where independent challenge helps.  

 

The team building the system should not be the only team choosing the difficult cases.  

 

A small challenge set prepared by the agency, another technical team or an external evaluator can reveal whether the pilot survives conditions it was not optimised to showcase. 

Let pilots fail usefully 

 

Governments often create the wrong incentive when every pilot is expected to lead to deployment.

A pilot that discovers a serious limitation before procurement has done valuable work. 

 

GovInsider's Lagos coverage offers a useful principle from another context: test ideas, document the evidence, and improve them at lower cost before committing significant public resources. The same logic applies to AI. 

 

A graduation decision should have at least three legitimate outcomes: scale, extend the pilot within defined limits, or stop.  

 

The second option matters. An AI tool may be useful for a narrow class of cases while remaining unready for the full service. 

 

A stopped pilot should leave behind a record of what failed and why.  

 

That evidence can improve the next procurement, prevent another agency from repeating the same mistake, and show suppliers which capability is genuinely missing. 

Make the evidence portable into procurement 

 

The pilot should end with a compact evidence dossier that procurement teams can use.  

 

It should include the baseline, tested configuration, representative failure cases, measured service outcomes, human workload, unresolved limitations, operating costs, monitoring needs, and the conditions under which the agency would reconsider deployment. 

 

It should also answer an uncomfortable question early: what happens if the agency decides not to continue with this supplier?  

 

Government should know whether it can export its data, preserve decision records, maintain the essential service and migrate without rebuilding the entire workflow from scratch. 

 

Singapore's Innovative Procurement Partnership already provides a useful structure: agencies can procure a pilot with an option to scale if testing succeeds.  

 

The word 'succeeds' should carry operational content. The graduation test supplies it. 

 

Shared test infrastructure can lower the cost of producing evidence.  

 

Public sector AI needs the institutional equivalent: reusable test methods, challenge environments and evidence templates that let smaller suppliers prove capability without every agency inventing a new process. 

The bridge from pilot to scale should be evidence 

 

Governments can keep experimenting, but each pilot should finish with a decision. 

 

Indonesia's current work is important because ministries are testing solutions against real problems before scale. The next step is to make the exit from the pilot as disciplined as the entry into it. 

 

Before the demo begins, write down what evidence would justify procurement, what evidence would narrow the deployment, and what result would stop it.  

 

Then preserve the failed cases as carefully as the successful ones. 

 

Public sector procurement should follow demonstrated service improvement under declared conditions, rather than the persuasiveness of the demo. 

 

----------------------------- 

 

Burak Oktenli is an independent researcher based in Washington, DC, focusing on assurance, verification and the governance of AI-enabled systems.  

 

He holds a bachelor's degree in computer science and engineering from the University of South Florida and an MBA, and is completing a Master of Professional Studies in Applied Intelligence at Georgetown University.  

 

His analysis has appeared in Eurasia Review, RealClearDefense, the Modern War Institute at West Point, RUSI, Articles of War and The Space Review.