SAP generative AI hub consultant: what they configure, deliver and cost
As of 2026-10-07
A SAP generative AI hub consultant sets up and runs the governed route from an application to a language model on SAP BTP: an AI Core tenant on the extended plan, resource groups, an orchestration deployment, grounding sources, masking and filtering rules, and the cost reporting that shows what each use case spends. In the corpus of 7 October 2026 only 4 of 1,657 live postings named AI Core or the hub (3 freelance, all in Germany), against 93 naming generative AI or GenAI (54 freelance, 61 in France). The work is advertised under generic wording far more often than under SAP's product names.
What the hub is, and two names that do not exist
The generative AI hub is a capability of SAP AI Core and SAP AI Launchpad, available only on the extended AI Core plan. Upgrading from standard is supported, downgrading is not. No SAP product is called SAP AI Hub, and none is called SAP Model Gateway. The gateway functions architects look for, one entry point, model switching without code change, an approved-model list, failover and quotas, come from the orchestration service plus AI Core rate limits and metering.
Two scenarios give access to models. Under foundation-models each model version is its own deployment. Under orchestration one deployment exposes a harmonised, OpenAI-style interface, so switching model is a configuration change. Allow and deny lists restrict which models a deployment may call. Pinned versions stop on their deprecation date, while a latest tag upgrades itself; the consultant chooses which risk the client prefers.
AI Core: the object model to set up
AI Core's vocabulary runs scenario, executable, configuration, then execution or deployment. A configuration only runs once it is turned into a deployment, which is why the object chain doubles as a promotion gate. Resource groups isolate executions, deployments, configurations and artefacts; scenarios, executables and registry secrets are shared across the tenant. The default ceiling is 50 resource groups per tenant, and each subaccount holds one tenant.
A proof of concept usually runs in the default resource group under one developer's credentials. Production needs a deliberate resource group per environment and a promotion decision. Our delivery material puts the move from proof of concept to production at 30 to 50 % of the original effort again, to be priced as its own line.
Orchestration: the order of steps, and the deadlines
The orchestration service runs a fixed sequence: grounding, templating, input translation, masking, input filtering, the model call, output filtering, unmasking and output translation. Only templating, which carries the model block, is mandatory. A list of configurations gives fallbacks: the service switches on an unsupported model, and for non-streaming calls on 408, 429 or 5xx errors.
Two dated items matter for inherited projects. The old masking provider property is scheduled for removal on 15 September 2026, and the first-generation completion endpoint is decommissioned on 31 October 2026. On a handed-over proof of concept, read first which version its calls use.
Grounding and RAG: choose the channel first
SAP offers three channels: the grounding module of the orchestration service, document grounding for Joule and its agents, and structured grounding from Business Data Cloud data products and the knowledge graph. Managed pipelines read SharePoint, S3, SFTP, ServiceNow, Google Drive and others, refresh daily and hold at most 8,000 documents each. A retrieval returns at most 100 chunks or 10 documents.
One limit drives design. For Joule, ingested documents are visible to all Joule users of the tenant; there is no per-user permission filtering. Where access rules matter, query the HANA Cloud vector engine directly and keep row-level security. For structured questions, ground on governed data products rather than raw tables. Fine-tuning is justified only when output format is non-negotiable, knowledge is stable and call volume passes roughly a million a day; for SAP content that changes every release, retrieval stays current by re-indexing.
Masking and filtering: where the settings leak
Masking uses one provider and about thirty entity types with uneven reach: person names are detected reliably in English only, and locations and addresses for the United States only. Anonymisation is irreversible and cannot tell two people apart; pseudonymisation is restored in the answer and in tool-call arguments. Masking of the retrieval query is off by default, and a PDF input needs an explicit method. A client working in French or German needs a test on real documents.
Filtering offers two providers: graded severity from Azure Content Safety and fourteen yes-or-no categories from Llama Guard. A rejected input returns HTTP 400 and consumes no model tokens. A blocked output returns HTTP 200 with a content-filter finish reason, so the application must handle it as a normal response. Azure OpenAI models also carry a global filter for severity 4 and 6 that your thresholds cannot lower, and each filter type is billed as a separate request. A consultant on this layer should be able to show a call that failed because a filter fired.
Cost: from tokens to capacity units, with a worked example
Custom generative AI is metered in tokens, converted to GenAI tokens at a model-specific rate and billed in BTP capacity units. Compute, storage and baseline charges are waived; grounding, filtering, masking and inference observability are metered separately. SAP's service-guide example, built on fictitious rates, prices 25,000 RAG requests of 3,500 input and 300 output tokens at 232.25 for the model calls and 186.3 capacity units for grounding storage and retrieval, filtering and masking: the surrounding services cost nearly as much as the model in that example, so budget them.
The resource group is the finest grain SAP reports natively. Per use case or agent, log token usage with business context or use inference observability labels, up to 16. Budgets alert but never suspend consumption. Explicit prompt caching works only for Claude and Nova models; batch consumption is cheaper but covers native calls only, so it bypasses masking, filtering and grounding.
Decision table: which mechanism for which case
One model, no safeguards needed — Layer: foundation-models scenario · Check: deprecation date, rate limit · Risk: provider-shaped calls tie code to one model.
Masking, filtering or grounding required — Layer: orchestration deployment · Check: module order, current version · Risk: rebuilding those steps in application code.
Questions over documents — Channel: grounding pipeline · Check: 8,000-document cap, daily refresh, access model · Risk: documents visible to every Joule user.
Questions over structured business data — Channel: data products or Datasphere models · Check: semantic descriptions · Risk: a fluent answer from an ambiguous label.
Personal data in prompts — Layer: pseudonymisation with allow list · Check: language and region coverage · Risk: names outside the detected scope pass through.
What we cannot assert
No source we hold publishes a day rate for this role or a list price for AI Units. The capacity-unit figures come from SAP's own worked example, which uses fictitious rates. Corpus counts depend on how skills are tagged in postings and mix SAP-specific and generic wording.
Frequently asked
Is there a certification for this role?
C_AIG_2604 is the SAP Certified Associate for the Generative AI Developer: a 3-hour, system-based assessment with a 76 % pass mark, replacing C_AIG_2412.
Does the generative AI hub work on the standard AI Core plan?
No. It needs the extended plan; standard covers custom AI only, and there is no downgrade from extended.
Where is the cost of a single use case visible?
Natively down to model, orchestration line and resource group. Finer attribution needs your own labels or logs.
What this page is built on
- SAP Generative AI Hub
- SAP AI Core
- SAP AI Core resource groups, scenarios and executables
- Generative AI Hub orchestration pipeline — the six modules in order
- Grounding & RAG (Retrieval-Augmented Generation)
- Document Grounding Vector Store (the ‘AI Foundation Vector Store’)
- Model Access in the Generative AI Hub (the ‘SAP Model Gateway’ question)
- Data Masking and Anonymisation in the Orchestration Service
- Content Filtering in the Generative AI Hub
- AI Cost Attribution in SAP AI Core — Tokens, Capacity Units and Resource Groups
- LLM Cost Engineering on SAP — AI Units, Token Budgets, Caching
- From proof of concept to production — the SAP AI delivery lifecycle
- RAG vs Fine-tuning vs In-Context Learning — Decision Frame for SAP Knowledge
- From SAP Analytics consultant to SAP AI specialist — the skills map
- SAP AI Launchpad
- Analytics Legends live contract corpus, 7 October 2026
External sources
- SAP AI Core — Generative AI Hub (SAP-docs)
- SAP AI Core — Orchestration (SAP-docs)
- SAP AI Core — Orchestration Workflow V2 (SAP-docs)
- SAP AI Core — Orchestration with Fallbacks (SAP-docs)
- SAP AI Core — Model Lifecycle (SAP-docs)
- SAP AI Core — Rate Limit Management (SAP-docs)