Glossary · Automation software engineering and architecture
Fault handling
Also known as: Fault management
German: Störungsbehandlung
In automation and systems engineering, fault handling is the detection, reporting, reaction to and recovery from faults in hardware, software, communication or the process, so that a machine or system reaches and maintains a defined state and can be returned to operation in a controlled way.
- Software engineering
In one sentence
Fault handling detects, reports and reacts to hardware, software or process faults and brings the system back to operation in a controlled way.
Example
When a drive reports an overcurrent fault, the fault handling stops the conveyor section, shows the drive fault on the HMI and requires acknowledgment and reset before restart.
How it applies
- Engineering: Fault handling covers detection (monitoring, diagnostics, plausibility checks), classification, Fault reaction (stop, degrade, continue with warning) and recovery (acknowledgment, Reset function, restart). It is often structured by fault classes with defined reactions.
- Operation: Operators need to understand what happened and what to do. Alarm texts, troubleshooting steps and clear reset conditions reduce downtime and prevent unsafe workarounds.
- Documentation: Troubleshooting documentation, alarm lists and service procedures are the user-facing side of fault handling. The documentation team should work from the same fault list as developers, so every fault message has a documented cause, effect and remedy.
Fault handling vs. error handling
Error handling concerns errors inside software. Fault handling covers the whole machine, including hardware and process faults. Where faults can lead to hazards, the reaction is often implemented as a safety function under functional safety standards; standard fault handling does not replace it.