Glossary · 4 · Grounding: content, retrieval and knowledge
Chunking
Also known as: Text splitting, Document chunking
Chunking is the splitting of documents into smaller passages — chunks — before they are embedded and indexed for retrieval, so that a RAG system can retrieve just the relevant part of a document.
- Intermediate
- Technical writers
- Developers
In one sentence
Chunking explained: how documents are split for AI retrieval, and why topic-based content chunks better than long PDFs.
Example
A 300-page PDF manual split every 500 tokens cuts a procedure in half; the same content authored as DITA topics can be chunked topic by topic, each with its metadata.
Why it matters on your learning path
- Technical writers: Topic-based authoring is chunking done by humans, with meaning intact — a strong argument for structured content in AI projects.
- Developers: Chunk along document structure (headings, topics) where possible, keep titles and metadata with each chunk and test chunk sizes with your evaluation set.