Glossary · Automation software engineering and architecture
Liveness probe
Also known as: Liveness check
German: Lebenszeichenprüfung
In container orchestration and service monitoring, a liveness probe is a periodic check that determines whether an application instance is still running correctly or has become stuck, so the platform can restart it automatically if the check fails repeatedly.
- Software engineering
In one sentence
A liveness probe periodically checks whether an application instance still works, so the platform can restart it if it is stuck.
Example
The orchestration platform calls the data collector's liveness endpoint every 10 seconds and restarts the container after three failed responses.
How it applies
- Engineering: A liveness probe should detect states the application cannot recover from by itself, such as a deadlock or a hung main loop. It should not depend on external systems, or an outage elsewhere will cause needless restarts.
- Operation: Probe intervals, timeouts and failure thresholds need tuning: too strict causes restart loops during slow startups or high load, too lax delays recovery. Startup and readiness probes handle initialization and traffic routing separately.
- Documentation: Document what the liveness check tests, its parameters and what happens on failure. Operators should know that restarts may hide recurring faults and where restart counts are recorded.
Liveness probe vs. watchdog
A liveness probe is checked from outside by an orchestration platform, and the reaction is a restart. A watchdog timer in a controller must be triggered by the running software, and its expiry leads to a defined reaction, such as stopping outputs or resetting the device. Both detect software that no longer runs as expected.