Platform engineering ·

What highly available means for customer-facing hosting services

Highly available, for a hosting service, means a customer can keep using the product when a single machine, process, or dependency fails. It is not a slogan about uptime percentages. It is a design choice about what is allowed to be a single point of failure.

On customer-facing hosting products that choice shows up in ordinary places. A site has to keep resolving and serving. A control panel has to keep accepting a change. An API used by provisioning has to fail in a way that can be retried, not in a way that leaves an account half-created. Internal tools get the same treatment when support and engineering depend on them during an incident.

Getting there is mostly unglamorous. Health checks, clear ownership, a deployment path that can roll forward or back, and interfaces that do not share a fate with every neighbor. Third-party integrations need the same attention, because a hosting product is often a set of platforms talking to each other.

That is the availability work I have been doing at Newfold Digital, on brands that serve small and medium businesses, and in the hosting roles before it. The lead responsibility is to keep that bar in the design, not only in the incident review.

Like this:

Back to the blog