Inference Economics — Cost, Latency, Quantization, Distillation
As of 2026-09-27
What is Inference Economics?
At 200K daily Joule sessions, a 10% token reduction alone saves an estimated $36K-$252K a year (an order of magnitude on public vendor list prices, not an SAP rate) — which is why prompt engineering is billable consulting work, not a nice-to-have.
What it is
Inference economics is the discipline of running LLMs at the cost, latency, and throughput that enterprise workloads can sustain. For SAP practitioners, this means understanding the per-1K-token math for BTP-hosted Joule, knowing which optimisation levers are available within the SAP trust boundary, and being able to build a business case for AI spend that CFOs will approve.
The cost of inference has three components: compute (GPU time per forward pass), memory bandwidth (KV cache read/write), and I/O (network + storage). In managed API pricing (OpenAI, Anthropic, SAP BTP Generative AI Hub), these are bundled into a per-token rate: input tokens billed at one rate, output tokens at a higher rate (output generation is autoregressive and therefore compute-intensive). On SAP BTP Generative AI Hub, model consumption is metered in AI Core capacity units on SAP's own published rate card, per model and per token direction — not an OpenAI list price passed through with a reseller margin. Read the current rate card and convert capacity units with your own BTP contract before you quote any per-token figure to a client. A Joule S/4HANA help-desk assistant with 200K daily sessions, 1K input + 500 output tokens per session = 200M input + 100M output tokens/day = approximately $1,000-$7,000/day on public vendor list prices (OpenAI/Anthropic, 2026) — an estimate, not an SAP rate; the BTP figure comes out of your capacity-unit conversion. At that scale, every 10% token reduction saves $36K-$252K/year — which is why prompt engineering is a billable consulting skill.
Why it matters
- Output tokens cost 3-4x input tokens because generation is autoregressive, so trimming output length saves more than trimming input.
- TTFT scales with input length — a 10K-token prompt takes roughly 10x longer to first token than a 1K-token one, which breaks sub-2-second interactive Joule assistants unless input is bounded and streamed.
- INT8 quantization halves compute and memory for under 1% quality loss, while INT4 saves 4x but costs 3-8% quality — fine for summaries, not for financial decision support.
Key points
- Inference cost = input tokens × input rate + output tokens × output rate; the BTP Generative AI Hub meters model consumption in AI Core capacity units on SAP's published rate card — quote it from that rate card and your BTP contract, never from a vendor's public per-token price.
- Time-to-first-token (TTFT) scales with input length; inter-token latency (ITL) is roughly constant per step — interactive Joule workflows must bound input length and enable streaming.
- INT8 quantization: ~2× memory/compute reduction, < 1% quality loss — production-safe for most SAP tasks. INT4: ~4× reduction, 3-8% quality loss — acceptable for low-stakes summarisation only.
- Distillation trains a small student model (7B) to mimic a large teacher (GPT-4-class) on a specific task — 10-100× inference cost reduction, 70-90% task-specific accuracy retained.
- Worked estimate, not an SAP price: at 200K daily SAP Joule sessions (1K input + 500 output tokens), public vendor list prices give $1,000-$7,000/day; a 10% token reduction = $36K-$252K/year.
- Batch vs. interactive trade-off: for nightly SAP analytics batch jobs, maximise throughput (large batch size, INT8); for interactive Joule assistants, minimise TTFT (short inputs, streaming, prefix caching).
Terms used on this page
- TTFT (Time-to-first-token)
- Latency from request submission to when the first output token is produced; dominated by the prefill pass over the full input context.
- Quantization
- Reducing the numerical precision of model weights (e.g. FP16 → INT8 → INT4) to lower memory footprint and inference compute, with a controlled quality trade-off (see C243 for strategy detail).
- Distillation
- Training a smaller student model to mimic a larger teacher model's output distribution on a target task; produces a task-specific model 10-100x cheaper at inference, at the cost of a one-time training run.
- Throughput
- Number of tokens generated per second across all concurrent requests; maximised by large batch sizes and quantization; the relevant metric for batch analytics workloads.
- AI Core capacity unit
- The unit SAP's BTP Generative AI Hub actually bills model consumption in, per model and per token direction, on SAP's own published rate card — the number to convert from a client's BTP contract, never a vendor's public per-token list price passed through unchanged.
- Bring Your Own Generative AI (BYOG)
- SAP AI Core's documented path for self-hosting an open-weight or custom-quantized model (Ollama, LocalAI, llama.cpp, vLLM, custom Hugging Face Transformers server) instead of consuming a hosted foundation model through the generative AI hub.
- Compute-bound vs. memory-bound
- Whether a workload's speed is limited by the accelerator's arithmetic throughput (compute-bound, favours batching) or by how fast data can move between memory tiers (memory-bound, favours FlashAttention/KV-cache techniques, see C214/C221).
Sources
- SAP BTP Generative AI Hub — pricing and service plans
- Dettmers et al. — LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale (2022)
- Hinton et al. — Distilling the Knowledge in a Neural Network (2015)
- Kwon et al. — Efficient Memory Management for LLM Serving with PagedAttention (2023)
- SAP Help Portal — Joule
- Google Cloud — 101 real-world generative AI use cases
- GitHub SAP-samples — btp-generative-ai-hub-use-cases: bring-your-own OSS LLM on SAP AI Core (Ollama, LocalAI, llama.cpp, vLLM, custom HF Transformers server; Standard/Extended plan required)
- SAP Help Portal — Choose a Resource Plan for training/inference in SAP AI Core (GPU instance types incl. SAP Note 3660109)
- SAP News Center — How SAP and NVIDIA Advance Enterprise AI Transformation (NIM inference optimisation, SAP-ABAP-1, 2026-03-17)
- SAP — AI pricing and AI Units (product pricing page)
- GitHub SAP-samples — btp-gen-ai-hub-sdk-samples: setting up the Extended AI Core service plan (the plan tier required for self-hosted BYOG serving)
- SAP — Value of AI: Oxford Economics 2026 (report PDF, ROI benchmark context for a cost/benefit comparison)
Full card available to members. What the full card adds: the full decision framework · the SAP vs Snowflake / Databricks / Fabric comparison · the common pitfalls and their fix · the cheat sheet · the architecture schemas · the code blocks · the facts worth quoting.