Connecting government data with APIs is the easy part. Understanding it is not
Connecting government systems is not the same as making their data interoperable. The missing piece is often a common language for government data.
-1788496518734.jpg)
Connecting systems allow data to be moved. But it does not guarantee that the receiving system understands what the data means. Image: Canva
Imagine three government information systems recording the same administrative district.
In one system, it is 07. In another, BISH-01. In a third, it is 3482.
The application programming interfaces (APIs) work. The integration platform works. The data moves exactly as designed.
But the systems still do not understand each other.
This is where many digital government programmes discover that technical interoperability is only half the problem.
Connecting systems allow data to be moved. But it does not guarantee that the receiving system understands what the data means.
The mapping problem
When two systems use different codes for the same concept, developers normally solve the problem with a mapping table.
System A says 07. System B says BISH-01. The integration layer knows that both values represent the same object.
For two systems, this is manageable. For a whole government, it becomes an architectural problem.
If every information system maintains its own classifications and reference data, mappings have to be created between an increasing number of systems.
With five systems, there may be up to 10 pairwise relationships. With 10 systems, 45. With 20 systems, 190.
And each connection may require mappings for administrative units, economic activities, organisation types, public services and dozens of other reference datasets.
This does not scale.
What begins as a few transformation rules gradually becomes a mapping mesh.
Every new system adds more mappings. Every change to a classification can require multiple mappings to be updated.
Diagnosing errors becomes harder because the meaning of data can change somewhere along the integration chain.
Eventually, the integration layer starts doing work that should have been solved at the data layer.
Governments need a semantic layer
The alternative is to reduce the need for translation in the first place.
Administrative units. Economic activities. Organisation types. Forms of ownership. Units of measurement. Categories of public services.
These datasets may look mundane compared to digital ID, cloud infrastructure or interoperability platforms. But they perform an essential architectural function: they give government systems a shared vocabulary.
Governments therefore need more than an interoperability layer for transporting data. They also need a shared semantic layer that helps systems interpret it.
One way to build this is through a National Reference Data Layer (NRDL): a common mechanism for publishing and distributing authoritative classifications, code lists and reference datasets across government.
The principle is straightforward: when one government system sends a code to another, both should already know what that code means.
Not another giant government database
A NRDL does not mean moving every classification into one centrally managed database.
The ministry or agency responsible for a particular domain can remain as the authoritative owner of its reference data.
What becomes common is the mechanism through which official versions are published, discovered, and consumed by government information systems.
That common infrastructure can provide machine-readable publication, APIs, versioning and change management.
Ownership can remain distributed. Access can be standardised.
The institutional model will differ between countries. What matters is that critical reference datasets have clearly identified authoritative owners, while government systems have a common way to access and use their official versions.
This is an important distinction. Centralising infrastructure does not necessarily mean centralising responsibility for the data itself.
Start with the “boring” data
Governments do not need to redesign their entire digital architecture to start.
Start with the “boring” data.
Identify a small set of classifications and reference datasets that are most frequently exchanged between government systems.
Determine which institution is authoritative for each one.
Publish their official versions through a common catalogue and repository. Make them machine-readable. Introduce versioning and change management. Require new government information systems to use national classifications where they exist.
Legacy systems do not have to be rebuilt overnight. Their reference data can be harmonised gradually as systems are modernised.
Over time, the architecture will begin to change.
Instead of System A translating its codes into System B's codes, and System B maintaining another mapping for System C, systems increasingly refer to the same authoritative reference data.
Eventually, we will see the mapping mesh start to disappear.
The invisible infrastructure of digital government
Reference data is unlikely to attract the same attention as a new digital identity platform, government superapp or AI service.
Citizens may never know that a national reference data repository exists.
But mature digital governments increasingly depend on precisely this kind of invisible infrastructure.
The more systems governments connect, the more important it becomes to distinguish between moving data and understanding data.
APIs allow government systems to talk. Reference data allows them to understand what they are saying. Digital government needs both.
---------
The author is an IT consultant working on digital transformation in the public sector, with a focus on data governance and interoperability. His recent work includes national-level projects in the Kyrgyz Republic related to metadata repositories and government data exchange systems.