Rethinking observability: Three Singapore organisations on turning alerts into action

At Dynatrace's Innovate Roadshow in Singapore, speakers from the Defence Science and Technology Agency (DSTA), Ministry of Education (MOE) and Changi Airport Group (CAG) highlighted that observability isn't about watching more dashboards, but deciding faster.

Speakers at Dynatrace Innovate Roadshow in Singapore highlighted the shift from monitoring, which simply tells you that a system is broken, to observability, which tells you why, and what to do about it. Image: Canva

Downtime on a defence platform is a question of national security, while an outage on a school system ripples into classrooms nationwide, and an airport app that stumbles become the public's lasting impression of Singapore itself.

 

At the Dynatrace Innovate Roadshow in Singapore on July 22, engineering leaders from the Defence Science and Technology Agency (DSTA), Ministry of Education (MOE) and Changi Airport Group (CAG), described a similar mindset shift happening inside their organisations.

 

This is the shift from monitoring, which simply tells you that a system is broken, to observability, which tells you why, and what to do about it.

 

Particularly in the public sector, the bill of disrupted services goes beyond financial to also affect public trust and continuity of essential services.

 

Amidst ageing IT estates and evolving cyber risks, the lesson from the three speakers was to look beyond dashboards and alerts, turning context into decisions

DSTA: From cracked glass to a single pane

 

DSTA's Head of Engineering (Application Platform Services), Ng Boon Wee, opened his presentation with "we cannot manage what we cannot see."

 

For an agency whose mission is equipping Singapore's soldiers with the systems that defend the

DSTA's Head of Engineering (Application Platform Services), Ng Boon Wee. Image: Dynatrace

country, the shift from traditional monitoring to modern observability is key to be able to quickly act when something goes wrong.

 

According to him, the distinction between monitoring and observability is that monitoring shows where a symptom exists, like a CPU spike or a memory leak, but observability goes further than that.

 

The latter allows you to investigate and act on an unpredictable, system-wide failure ripples across an interconnected architecture.

 

But he recognised that observability is harder in a defence environment, where workloads are segmented across isolated networks for security which traps telemetry in silos.

 

He likened this to "looking through a cracked glass", where engineers manually stitch together logs from disparate teams.

 

Since 2022, DSTA's answer has been a self-hosted, multi-tenanted architecture managed by DSTA's Cloud and Platform Engineering (CAPE) team which Ng is part of.

 

This allows each tenant to get an isolated, self-managed workspace, while DSTA retains oversight across the whole environment.

 

The rollout was phased, as it started with real user experience on the frontend and later with services and dependencies on the backend.

 

His team eventually extended that single view to one of DSTA's largest platforms, its national service (NS) digital system.

 

Looking ahead, Ng's team is pushing to extend observability into automated vulnerability and threat management.

 

This meant moving beyond simply answering "why" to potentially automate the detection of threats and triggering the responses to tackle these threats.

 

"Observability is a resilience capability, not just a tool," he says, adding that success depends on breaking down organisational silos and not just fixing dashboards.

MOE: When alerts aren't the problem, but the context

 

 "Cyber safety is no longer just about getting alerts. It's about reducing uncertainty, understanding the impact, and making faster decisions during an incident," says MOE's Observability Evangelist, Kenneth Yeo.

 

This framing matters considering MOE's scale. According to him, the Schools Standard ICT Operating Environment (SSOE) spans some 380 schools, 35,000 teachers and 450,000 students.

 
MOE's Observability Evangelist, Kenneth Yeo. Image: Dynatrace

An incident is therefore measured by how many schools and services are hit, and how fast recovery happens, he explained.

 

"Teams are often overloaded with alerts, vulnerabilities and findings. The question is not whether risk exists. The question is which risks matter more now," he added.

 

While monitoring detects symptoms, observability explains impact.

 

He noted that having a correlated view of the context allows cyber defence team to scope down the risks to focus on, instead of investigating in silos and losing valuable time piecing together the whole picture.

 

"The question is not how many dashboards we have. The better question is: are we reducing decision time and cyber exposure? The goal is not more data. The goal is better cyber decisions under pressure," he summarised.

CAG: From "logs and luck" to a 100MB gif lesson

 

Reflecting on CAG's journey, Saurabh Dutta, General Manager, recalls a time when the mobile app ran on 'logs and luck', which was a starting point that set the stage for the transformation observability would bring.

 
CAG's General Manager, Saurabh Dutta. Image: Dynatrace

As the airport's digital front door, the Changi app handles payments, shopping, parking and

rewards for roughly 40,000 daily members and 160,000 transactions a month.

 

Every incident meant digging through logs after it happened, then investigating whose layer was at fault.

 

After implementing observability, Dutta highlighted that his team was able to translate technical faults into business impact.

 

Weeks after the app went live, he shared an incident where senior leadership on an overseas retreat feedbacked that the app was draining their roaming data.

 

It took his team five minutes to trace the cause to an oversized, 100MB animated GIF loading on every screen. The gif was then compressed and fixed within ten minutes.

 

That precision enabled by observability allowed his team to scope for real incidents tightly instead of treating every alert "as a building on fire."

 

Since the implementation, he highlighted mean time to detect has dropped to around five minutes, most incidents resolve within an hour. Most importantly, his team now catches over half of issues before the travellers notice.

 

"Observability stopped being an IT cost. It became a way to protect the travellers' experience," he said.

 

Whether the mission is national defence, public education or keeping an airport running, the three speakers converged on the takeaway that context matters more than isolated alerts, and turning that context into decisions is key to close the downtime between service incident and recovery.