Tokenization and Embeddings in Enterprise LLM Pipelines
As of 2026-09-27
What is Tokenization and Embeddings in Enterprise LLM Pipelines?
Tokenization choices silently inflate SAP prompt costs and truncate documents — German SAP terms cost 20-40% more tokens than English, and fixed-token chunking can cut a sentence in half without warning.
What it is
Tokenization and embeddings are the two pieces of invisible infrastructure that sit between raw text and everything a large language model does with it—and they are exactly the pieces that go unnoticed until they cause a cost spike, a silent document truncation, or a retrieval failure that makes Joule look less capable than it is. Understanding both lets you design SAP knowledge pipelines that work on the first attempt rather than the third.
Why it matters
- SAP-specific and German compound terms split into 8-12 tokens vs. 2-3 for English equivalents, directly inflating per-call cost.
- Numbers tokenize digit-by-digit, so material numbers and cost-centre codes bloat token counts and degrade arithmetic reasoning unless pre-processed.
- Chunking SAP PDF documentation at a fixed 512 tokens silently truncates sentences that cross the boundary — the model never sees the missing half.
Key points
- BPE vocabulary size: GPT-4 uses ~100K tokens; SAP compound German terms tokenize into 8-12 tokens vs. 2-3 for equivalent English — budget +20-40% tokens for German/SAP-notation prompts.
- Numbers tokenize digit-by-digit in most LLM vocabularies; SAP material numbers and cost centre codes inflate token counts and degrade arithmetic reasoning — pre-process where possible.
- Truncation is silent: always chunk SAP documents at sentence or paragraph boundaries, never at fixed token counts, to avoid splitting mid-sentence.
- Embeddings (768-4096 dimensions) cluster semantically similar texts; enable synonym-tolerant retrieval ('cost object' = 'controlling object') that keyword search misses.
- HANA Cloud Vector Engine (GA 2024) stores and queries embeddings natively alongside relational data — no separate vector DB needed for most SAP deployments.
- SAP domain-adapted embedding models on AI Core outperform generic OpenAI embeddings by 8-15% recall@5 on SAP Help Portal retrieval tasks.
Terms used on this page
- BPE (Byte-Pair Encoding)
- Subword tokenisation algorithm that iteratively merges the most frequent adjacent character pairs until reaching a target vocabulary size.
- Embedding
- Dense vector representation of text where semantically similar texts cluster nearby in the high-dimensional space; foundation of semantic search and RAG retrieval.
- HANA Cloud Vector Engine
- Native vector storage and approximate nearest-neighbour (HNSW) search capability in SAP HANA Cloud, enabling embedding-based retrieval without a separate vector database.
- Recall@k
- Retrieval quality metric: fraction of relevant documents found in the top-k results; standard benchmark for RAG pipeline evaluation.
- SentencePiece / WordPiece
- Subword tokenisation schemes closely related to BPE (SentencePiece: Kudo & Richardson 2018; WordPiece: used in BERT, Devlin et al. 2018) — differ in merge criteria but share the same digit-by-digit and compound-word overhead behaviour this card describes.
- Chunking
- Splitting a source document into passages small enough to embed and retrieve individually. The chunk boundary must follow sentence or paragraph structure — chunking at a fixed token count silently truncates content mid-sentence.
- Vector index (HNSW)
- Hierarchical Navigable Small World graph (Malkov & Yashunin, 2016): the approximate-nearest-neighbour index structure most production vector stores, including SAP HANA Cloud's Vector Engine, use to make similarity search fast at scale.
- Domain-adapted embedding model
- An embedding model further trained on domain-specific text (e.g. SAP Help Portal, S/4HANA documentation) rather than open web text alone — measurably improves recall@k on domain-specific retrieval versus a strong generic baseline, at the cost of a hosting and re-embedding dependency.
Sources
- SAP HANA Cloud Vector Engine documentation
- SAP AI Core — embedding models on BTP
- Sennrich et al. — Neural Machine Translation of Rare Words with Subword Units (BPE, 2016)
- OpenAI tiktoken tokeniser
- SAP Business AI — official product page
- SAP Joule (work companion) — official product page
- SAP Generative AI — official product page
- SAP Community — Querying RDF Graphs with SAP HANA Cloud Knowledge Graph Engine
- OpenAI — New embedding models and API updates (text-embedding-3-large, 2024)
- Kudo & Richardson — SentencePiece, arXiv:1808.06226 (2018)
- Devlin et al. — BERT: Pre-training of Deep Bidirectional Transformers (WordPiece), arXiv:1810.04805 (2018)
- Malkov & Yashunin — Efficient and Robust Approximate Nearest Neighbor Search Using HNSW, arXiv:1603.09320 (2016)
Full card available to members. What the full card adds: the full decision framework · the SAP vs Snowflake / Databricks / Fabric comparison · the common pitfalls and their fix · the cheat sheet · the architecture schemas · the code blocks · the facts worth quoting.