Glossary · System coordination, integration and orchestration
High availability cluster
Also known as: HA cluster, failover cluster
German: Hochverfügbarkeitscluster
In IT and OT infrastructure, a high availability cluster is a group of servers that monitor each other and run services redundantly, so that a service moves to another node or continues on remaining nodes when one node fails.
- System integration
In one sentence
A high availability cluster is a group of servers that monitor each other and keep services running when one node fails.
Example
The MES database runs on a two-node cluster with a witness; when the active node crashes, the database restarts on the second node within a minute.
How it applies
- Architecture: Clusters combine redundant nodes, shared or replicated storage, heartbeat monitoring and a mechanism to avoid split brain, such as quorum or a witness.
- Operation: High availability depends on operation as much as on design: patching nodes one at a time, monitoring replication and testing Failover regularly.
- Documentation: Operations documentation should state the availability target, which failures the cluster tolerates and which it does not (for example a site-wide power failure), and step-by-step procedures for maintenance and recovery.
High availability vs. fault tolerance
A high availability cluster minimizes downtime but usually allows a short interruption during failover. Fault-tolerant systems continue without interruption, typically with lockstep redundancy, at higher cost. Which approach fits depends on how long the process can tolerate an interruption.