Model Access in the Generative AI Hub (the ‘SAP Model Gateway’ question)
As of 2026-09-27
What is Model Access in the Generative AI Hub (the ‘SAP Model Gateway’ question)?
SAP ships no product called ‘SAP Model Gateway’. The gateway functions architects look for — one entry point, model switching without code change, an approved-model policy, failover, quotas and consumption reporting — are delivered by the orchestration service of the generative AI hub and by SAP AI Core rate limits and metering, each with precise limits worth knowing.
Why this card was re-centred
Earlier versions of this card described an ‘SAP Model Gateway’ with quality tiers, per-skill cost centres and automatic routing, announced at Sapphire 2026. No such product appears in SAP's AI Core and generative AI hub documentation (September 2026), in the Sapphire 2026 platform announcement or in SAP Learning. The need behind the name is real, though: every enterprise running several LLMs wants one controlled path to them. In the SAP stack that path is the generative AI hub, and the useful exercise is to map each gateway expectation to the mechanism SAP actually ships — and to its limits.
Why it matters
- Pitching or budgeting a ‘Model Gateway’ SKU that SAP does not sell undermines credibility in the first architecture review.
- Without a deployment-level allow list, the approved-model policy is advisory; any developer can call any model the tenant offers.
- Rate limits are counted in requests per minute per model, not in tokens per user — spend ceilings per application must be built, not configured.
Key points
- ‘SAP Model Gateway’ is not an SAP product; the gateway role is played by the generative AI hub orchestration service plus AI Core rate limits and metering.
- Single entry point: one orchestration deployment per resource group, POST /v2/completion, harmonized OpenAI-style API.
- Model switch = change model.name/version in prompt_templating or in a stored orchestration config — no new integration.
- Approved-model policy: modelFilterList + modelFilterListType (allow/deny) on the orchestration deployment.
- Failover: ordered fallback configurations; triggers are regional unavailability and, for non-streaming calls, 408, 429 and 5xx.
- Quotas: requests per minute per model and tenant, optional dedicated resource-group limits via the quota-management API; 429 + Retry-After.
- Cost attribution natively down to model, orchestration and resource group in the BTP usage export; finer grain needs labels and logs.
- Joule's own LLM calls run through AI Core under SAP management with no-retention/no-training terms; your controls apply to what you build.
Terms used on this page
- Harmonized API
- Orchestration interface that normalises providers onto the OpenAI API conventions for messages, parameters and responses.
- modelFilterList / modelFilterListType
- Parameter bindings of an orchestration deployment that allow or deny specific models and versions.
- Fallback configuration
- Ordered list of module configurations tried in sequence on defined errors or regional unavailability.
- intermediate_failures
- Response field reporting why higher-preference configurations were skipped.
- RPM rate limit
- Maximum requests per minute per model within a tenant, shared by all versions of that model.
- Quota-management API
- AI Core endpoints to read and change rate limits, including dedicated resource-group limits.
- GenAI token
- Virtual unit converting a model's input and output tokens at model-specific rates before conversion into capacity units.
Sources
- SAP AI Core — Orchestration (SAP-docs)
- SAP AI Core — Harmonized API (SAP-docs)
- SAP AI Core — Model Restriction (SAP-docs)
- SAP AI Core — Orchestration with Fallbacks (SAP-docs)
- SAP AI Core — Rate Limit Management (SAP-docs)
- SAP AI Core — Batch Consumption (SAP-docs)
- SAP AI Core service guide — metering and pricing (PDF, 4 Sep 2026)
- SAP News — SAP Unveils Business AI Platform (Sapphire, May 2026)
- SAP Community — Trust by Design: Security, Guardrails & Governance in Joule (Aug 2026)
- SAP Help — Generative AI hub overview (model access)
- SAP Help — Metering and pricing for generative AI (SAP AI Core service guide)
- SAP Help — Prompt Registry (generative AI hub)
- SAP Developers — Tutorial: generative AI hub prompt registry
Full card available to members. What the full card adds: the full decision framework · the SAP vs Snowflake / Databricks / Fabric comparison · the common pitfalls and their fix · the cheat sheet · the architecture schemas · the code blocks · the facts worth quoting.