Glossary Updates12 new terms added to the glossaries · October 2, 2026, 22:44 CEST
AI TechDocKnowledge

Glossary · 1 · Foundations: how AI works

Transformer

The transformer is the neural network architecture behind today’s large language models. It uses a mechanism called attention to weigh how strongly every token in a sequence relates to every other token, which lets it model long-range context and train efficiently on very large amounts of text.

  • Intermediate
  • Expert
  • Developers

In one sentence

The transformer architecture explained: attention over tokens, the design behind GPT, Claude, Gemini, Llama and nearly every modern LLM.

Example

In “The pump stops when the valve closes because it senses the pressure drop,” attention helps the model link “it” to “the pump” rather than “the valve.”

Why it matters on your learning path

  • Technical writers: Understanding attention explains why clear, unambiguous references in your source content help AI answers.
  • Technical project managers: Transformer-based models dominate the market; architecture is rarely a selection criterion, model quality and terms of use are.
  • Developers: Knowing the architecture helps you reason about context length, cost per token and why output is generated one token at a time.

Background

The architecture was introduced by Google researchers in the 2017 paper “Attention Is All You Need.” The “T” in GPT stands for transformer.

By knowledge.aitechdoc.world · Published September 26, 2026 · Last reviewed

Source: Vaswani et al., “Attention Is All You Need,” NeurIPS 2017

Definitions follow the cited standards and specifications. Where a source is a copyrighted publication, such as an ISO, IEC or EN standard, the definition is a close paraphrase, not a verbatim quotation, so as not to infringe copyright. We recommend reading the original publication. The sections “How it applies” are editorial commentary by AI TechDoc Knowledge and are not part of any standard.

Seen a mistake? Send us a note!