Glossary · PLC and control programming
Unicode
Also known as: Unicode Standard, UTF-8, UTF-16
German: Unicode
In computing, Unicode is the character encoding standard that assigns a unique code point to each character of the world's writing systems, stored in encodings such as UTF-8 and UTF-16; in IEC 61131-3, the data types WSTRING and WCHAR hold double-byte characters.
- PLC programming
- Standards
In one sentence
Unicode gives every character a unique code point; in PLCs, WSTRING and WCHAR hold such text for multilingual HMIs and messages.
Example
A machine sold in China, Germany and Poland stores alarm texts in WSTRING variables so that Chinese characters and Polish diacritics display correctly.
How it applies
- Engineering: Classic PLC STRING types hold single-byte characters, which cannot represent all languages. WSTRING and WCHAR store double-byte characters; the exact encoding and conversion functions depend on the controller.
- Integration: Text passes through PLC, HMI, OPC UA, databases and files. Every step must agree on the encoding, otherwise characters turn into question marks or garbage. UTF-8 is the common choice for files and web interfaces.
- Operation: Operators must be able to read alarm and instruction texts in their language. Wrong characters in safety messages reduce readability, see Safety information readability.
- Documentation: Translation workflows deliver texts that end up on HMIs and in PLC messages. Specify the encoding and maximum string lengths in the interface description, because translated texts are often longer than the source.
Unicode vs. encoding
Unicode defines which number represents which character. Encodings such as UTF-8 and UTF-16 define how those numbers are stored as bytes.