AI & Analytics Legends The knowledge platform for SAP Analytics
Academy module

LLM Internals for SAP Architects — Attention, Positional Encoding, Long Context and Mixture-of-Experts

LLM Internals for SAP Architects — Attention, Positional Encoding, Long Context and Mixture-of-Experts — Analytics Legends section illustration for the SAP Analytics knowledge base (concepts, studies, Academy)

As of 2026-10-04

Opens the model behind SAP's generative AI hub for architects who already consume it. It teaches scaled dot-product attention and the KV cache, grouped-query attention, rotary position embeddings and the context-extension methods built on them, the gap between advertised and effective context (Lost in the Middle, RULER), sliding-window and sparse attention, and mixture-of-experts routing (Mixtral, DeepSeek-V3), each tied to a consequence: time to first token, memory per user, tokens billed, retirement and rate limits. The lab turns SAP AI Core's documented model-discovery fields and metering example into a sizing sheet for a Datasphere-grounded assistant, with a position test run on your own data.

What you will learn

  • Explain scaled dot-product attention, multi-head attention and the KV cache well enough to predict how prompt length drives time to first token and memory
  • Compute the KV-cache size of a model from its layer count, key-value heads and head dimension, and show what grouped-query attention changes
  • Describe how rotary position embeddings encode relative position, why models fail beyond their trained length, and what position interpolation and YaRN do about it
  • Separate a model's advertised context window from its effective window using Lost in the Middle and RULER, and design a prompt budget accordingly
  • Explain sliding-window and sparse attention and mixture-of-experts routing, and state which cost they cut (compute per token) and which they do not (memory for all experts)
  • Read the SAP AI Core model-discovery fields (contextLength, inputCost, outputCost, retirementDate), the metering page and the rate-limit page, and turn them into a sizing decision for a Datasphere-grounded assistant

Module overview

Who this is for. You already use the SAP generative AI hub: you have picked a model name in an orchestration configuration, seen a 429, and been surprised by a token bill. This module opens the box one level down. It explains the five mechanisms that decide what a large language model can read, how fast it answers and what it costs: attention, positional encoding, the context window, the key-value cache and mixture-of-experts routing. It assumes M333 (AI and LLM fundamentals for SAP consultants) and benefits from M325 (generative AI hub hands-on) and M365 (LLM cost engineering). It teaches you to read a model card, a contextLength field and a latency curve like an architect. Primary sources are the original papers on arXiv and vendor documentation; every mechanism below ends in a consequence you can act on in an SAP project.

Prerequisites

  • Completion of M333 (AI & LLM Fundamentals for SAP Consultants) or equivalent working knowledge of tokens, embeddings and context
  • Basic linear algebra (dot product, matrix multiplication) and comfort reading a model card or a short arXiv abstract
  • Optional: an SAP AI Core tenant with the extended plan (generative AI hub) to run the model-discovery call in the lab; otherwise use the field descriptions in SAP's documentation

Outcomes

  • Predict, from prompt length and model shape, whether a request will be slow to start, slow to stream or constrained by memory, and say which mechanism causes it.
  • Size a KV cache and a monthly token volume for a grounded assistant and defend the numbers with SAP's documented fields.
  • Design a prompt budget and a position test that exposes the gap between advertised and effective context.
  • Explain to a client why a mixture-of-experts model can be cheap per token and expensive to host, without quoting a price that SAP has not published.
  • Choose between aggregating in the data layer, retrieving fewer chunks and moving to a larger window, with evidence.

Full module available to members. The full module adds: the decision framework · the end-to-end scenario walkthrough · the KPI scorecard · the anti-patterns · the code blocks · the knowledge check · the diagrams.

Open in the app →