AI & Analytics Legends The knowledge platform for SAP Analytics
Concept card

AI Cost Attribution in SAP AI Core — Tokens, Capacity Units and Resource Groups

AI Cost Attribution in SAP AI Core — Tokens, Capacity Units and Resource Groups — Analytics Legends section illustration for the SAP Analytics knowledge base (concepts, studies, Academy)

As of 2026-10-06

What is AI Cost Attribution in SAP AI Core?

There is no SAP ‘Model Gateway’ billing product: custom generative AI in SAP AI Core is metered in tokens, converted into GenAI tokens per model (SAP Note 3437766) and billed in BTP capacity units. The resource group is the finest cost grain SAP reports natively; cost per scenario or agent needs your own labels — inference observability or application logs.

Why this card was re-centred

Earlier versions of this card explained AI cost attribution through an ‘SAP AI Core Model Gateway’ that metered every Joule and custom LLM call and billed it per scenario. No product of that name exists in SAP's documentation (SAP AI Core service guide and generative AI hub documentation, September 2026). The underlying question is real and increasingly asked by finance teams: who consumed which AI, and what did it cost? In the SAP stack the answer lives in three places — the metering of the generative AI hub in SAP AI Core, the Costs and Usage pages of the SAP BTP global account, and whatever tagging you add yourself. This card explains each layer, how fine it goes, and where your own design has to take over.

Why it matters

  • Finance will ask which AI feature drove the bill; SAP's native answer stops at the resource group, so the attribution design has to exist before the first production deployment.
  • SAP-delivered AI (AI Units, SAP for Me) and custom AI on AI Core (capacity units, BTP contract) are billed in different currencies and places — a single ‘AI cost’ figure that mixes them is wrong by construction.
  • Budgets alert but never stop consumption, and rate limits cap requests per minute, not monthly spend: spend control is a design task, not a setting.

Key points

  • ‘SAP Model Gateway’ does not exist; metering of custom generative AI happens in SAP AI Core's generative AI hub (extended plan) and billing in SAP BTP.
  • Metering chain: input and output tokens → GenAI tokens at model-specific rates (SAP Note 3437766) → capacity units (CU); output tokens are generally slightly more expensive.
  • Compute, storage and baseline are waived for generative AI; grounding storage and retrieval, content filtering, data masking and inference-observability storage are metered separately.
  • Joule and embedded SAP AI are a different world: Base AI or AI Units, monitored in SAP for Me (C333), invisible in your AI Core tenant.
  • Native cost grain: global account (billable) → directory/subaccount (estimated, labels searchable) → resource group → model and orchestration lines in the AI Core usage export.
  • Per scenario, agent or cost centre: log the usage object (prompt_tokens, completion_tokens) with business context, or record requests with inference observability (metadata mode, up to 16 ext.ai.sap.com labels).
  • Controls: BTP budgets (≤ 30 active, 3 thresholds, alert only), RPM rate limits per model/tenant or resource group, model allow lists, batch consumption for bulk native LLM calls.
  • A label-based estimate reconciles on trends and shares; the monthly balance statement stays legally binding.

Terms used on this page

GenAI token
Virtual unit that converts a model's input and output tokens at model-specific rates (SAP Note 3437766) before conversion into capacity units.
Capacity unit (CU)
Billing metric of SAP BTP; AI Core converts GenAI tokens and module metrics into CU.
Resource group
Isolation unit inside an AI Core tenant; the finest grain at which SAP reports AI Core consumption in the usage export.
Costs and Usage page
Global-account page of the BTP cockpit with Overview, Billing, Usage and Budgets tabs; subaccount and directory figures are estimations.
Inference observability
Opt-in recording of flagged generative AI hub requests (metadata or full payload) with labels and feedback; GA since 8 May 2026.
Inference label
Key-value pair with the prefix ext.ai.sap.com attached to an inference record (max 16, 64 characters each) for later filtering.
AI Units
Pooled SAP Business AI currency for premium SAP-delivered AI such as Joule agents, monitored in SAP for Me — not part of AI Core billing.

Sources

  1. SAP AI Core service guide — metering and pricing, consumption information (PDF, 4 Sep 2026)
  2. SAP AI Core — Metering and Pricing for Generative AI (SAP-docs)
  3. SAP AI Core — Inference Observability (SAP-docs)
  4. SAP AI Core — Record an Inference in Inference Observability (SAP-docs)
  5. SAP AI Core — Rate Limit Management (SAP-docs)
  6. SAP AI Core — Batch Consumption (SAP-docs)
  7. SAP AI Core — Prompt Caching (SAP-docs)
  8. SAP BTP — Monitoring Usage and Consumption Costs in Your Global Account (SAP-docs)
  9. SAP BTP — Managing Budgets in Your Global Account (SAP-docs)
  10. SAP AI Core — What's New (SAP-docs)
  11. SAP Learning — Introducing Joule: Understanding the Commercial Model (Base AI, AI Units)
  12. SAP-docs — Resource Groups (SAP AI Core, the isolation unit used for cost attribution)
  13. SAP — ai-sdk-js release v2.16.0 (model list change, Mistral Large removal dated 30 Sep 2026)
  14. SAP — ai-sdk-js release v2.13.0 (gpt-4.1, o3, o4-mini, claude-4-sonnet retirement dates in October 2026)

Full card available to members. What the full card adds: the full decision framework · the SAP vs Snowflake / Databricks / Fabric comparison · the common pitfalls and their fix · the cheat sheet · the architecture schemas · the code blocks · the facts worth quoting.

Open in the app →