Home Services Industries Stories Insights About Sovereign AI Contact us
Home / Stories / Seven Servers, Seventy-Two Hours
Customer story

Seven Servers, Seventy-Two Hours.

Aboriginal Community Controlled Health Organisation · Northern Territory · 7 min read

Why a healthcare organisation needed more than a helpdesk — it needed a team that could scale

Not if, but when

Healthcare IT in the Northern Territory does not resemble healthcare IT anywhere else in the country. Sites sit hundreds of kilometres apart on links never engineered for the load now crossing them. Hardware ages faster in the heat and the dust. And the people at the end of it are not processing invoices — they are seeing patients.

Our client is a large Aboriginal community controlled health service working across multiple remote communities in the Territory. More than 150 staff, spread across sites that can be a day's travel apart, all working from a centralised virtual desktop environment. Clinical records, email, core systems — everything arrives through the same door.

Which means that when the door closes, it closes for everyone at once. A nurse in a remote clinic and an administrator in town lose access in the same second, and neither has a local copy to fall back on.

For this organisation, a server outage isn't an inconvenience. It's a clinical event.

Seven years before the phone rang

We had been managing this environment since 2018. That matters more than it sounds, and it is the reason the rest of this story reads the way it does.

2018–2019. The on-premise servers were ageing and the backup systems were near the end of their working lives. Our Projects Director made the case that the organisation had outgrown what could responsibly be maintained in a comms room, and recommended a move to professionally managed colocation.

2019–2021. A staged migration into a Telstra enterprise data centre — three phases across eighteen months. Assessment and initial migration, then full server relocation, then decommissioning of the legacy on-site hardware. Every stage planned, signed off and executed alongside the organisation's own IT manager.

In February 2020, midway through that arrangement, Telstra's own infrastructure failed — 0638 to 1800 hours. Our client's systems went down with it. Nothing we could have done would have prevented it; the fault sat entirely outside our control. What sat inside our control was the response, and because the run books were documented and the team already knew the environment, the day was managed rather than improvised.

2023–2026. A further migration onto cloud-hosted virtual infrastructure with Zettagrid: seven RDS servers carrying the organisation's entire desktop environment, with Veeam backup running against every workload.

Seven years of institutional knowledge. That's what was already in the room when the call came in.

The warning signs

The failure did not arrive unannounced.

For weeks beforehand, our Projects Director had been escalating a resourcing problem. Analysis of the virtual environment showed all seven RDS servers running at or near capacity — no headroom, no slack, nothing held in reserve for a difficult day.

His note read, in part:

"Am I reading this correctly — there's no spare RAM to play with and they need more immediately?"

A formal quote for the upgrade had already gone to the client. The problem was identified, the recommendation made, the paperwork raised.

This is what a managed service looks like on an ordinary Tuesday. Not waiting for the phone. Watching an environment closely enough to notice it drifting toward a limit, and saying so while there is still time to act.

The fire hadn't started yet. But we could already smell smoke.

The incident

On 29 June 2026, routine maintenance ran across all seven RDS servers: temporary file cleanup and a scheduled reboot. Standard procedure, performed many times before without incident.

What no one could see was that a Windows Component Cleanup process, triggered six days earlier, was still working through a backlog in the background. It was the first time that process had run since the servers were built years before, so it had a great deal to get through. It surfaced no progress indicator and appeared nowhere in the maintenance workflow. The reboot interrupted it mid-execution.

All seven servers stopped in the same moment — each one hanging before it could complete startup, locked on a "Preparing Windows" screen. Every remote desktop in the organisation went dark simultaneously.

We began within hours, working the conventional recovery paths first: safe mode, update rollback, system restore, file system integrity scans. None of them moved it. The interrupted cleanup had left the servers in a state no repair route could reach.

So we stopped trying to repair them and started rebuilding — and we did it in parallel rather than in sequence.

At 6:23pm that evening, full disk backup restoration began on the primary server. By 9:12pm, a clean reinstallation had started on a second. Through the night and into the next day, our engineers worked across the remaining five, bringing servers back into service one at a time as each was rebuilt and verified rather than holding the environment until all were finished. The last server came back on 1 July.

Seven servers rebuilt from scratch, progressively restored across roughly 72 hours.

While that ran, we opened a separate fault with the infrastructure provider, having identified that backup restoration was running slower than it should — a second investigation underway before the first was finished.

This was not one engineer working overnight. It was a coordinated response across several disciplines at once, and it did not stop until the organisation was running.

What a managed service provider actually is

Most organisations judge their IT partner on tickets. How fast did someone answer? How quickly was the laptop replaced?

That is not the wrong question, and any provider who cannot answer the phone should be replaced. But it is not the question that determines whether your organisation gets through a bad week.

The question that matters is what happens when something business-critical fails in a way no single technician can solve — when the problem spans infrastructure, backup systems and vendor relationships simultaneously, and the clock is running because clinical staff cannot work.

What you need in that hour is:

  • Institutional knowledge. People who have run this environment for years and do not have to learn it while it is failing.
  • Proactive monitoring. A team that raised the resource risk before the incident, not in the report afterwards.
  • Scalable response. Enough senior engineers to work one problem in parallel, around the clock, instead of queuing it behind everything else.
  • Vendor coordination. Existing relationships with infrastructure providers, so a second fault gets raised and investigated at speed rather than sitting in a web form.
  • A plan for when the plan fails. Recognising early that restoration will not get there, making the call to rebuild, and executing without hesitating over it.

Any IT company can close a ticket. Not every IT company can rebuild seven servers and keep a health service running while it does.

The outcome, and what came after

The organisation was restored in full. A mitigation went in immediately: Windows Component Cleanup now runs weekly, so the backlog that caused the failure cannot build again. The slow restoration was investigated with Zettagrid and a fix identified.

And the resource upgrade we had recommended before the incident remained exactly where it was — on the table, still needed, still ours to keep pushing. The work that began before the crisis carried on after it, because that work is the job. It is not something you take up once the emergency has passed.

For organisations at the edge of the coverage map, where an IT failure means something more serious than a slow morning, the question was never whether your provider can close a ticket quickly. It is whether they can carry you through the week your systems do not come back on their own.

CategoryCustomer story
SectorAboriginal Community Controlled Health Organisation
Read time7 min

Emerge IT has delivered managed IT, cybersecurity and infrastructure services across regional and remote Australia for more than 25 years, from head office in Darwin. If your organisation depends on systems that cannot be allowed to fail quietly, we should talk.

Keep reading

Related reading.

Customer stories

More stories from the field

Read more
Insight

Speech-to-text, compared

Read more
Insight

AI in health services

Read more
All stories

Turn remote complexity into reliable advantage.

Tell us what you're working towards, and we'll help you build a hyper-efficient, connected organisation — supported end to end, wherever your people are.

Get in touch · 1300 EIT AUS