Foundation Model Selection — OpenAI vs Claude vs Gemini vs Llama
As of 2026-10-06
What is Foundation Model Selection?
Generic leaderboards (HumanEval, MMLU) don't predict SAP task performance — and long context no longer separates the vendors: Claude Opus 5.5 and Sonnet 5 and Gemini 2.5 Pro all offer 1M-token windows, so the choice turns on task-level evals, data residency and cost.
What it is
Foundation model selection is the recurring architectural decision every SAP AI initiative eventually faces: which large language model — the GPT-6 family from OpenAI (Astra, Sol, Luna), Claude from Anthropic, Gemini from Google, or an open-weight model such as Llama or Mistral — should sit behind a given Joule extension, BTP AI API integration, or custom analytics feature. It matters because the decision is expensive to reverse: prompt engineering, evaluation harnesses, fine-tuning data, and cost models are all built around a specific model's behavior, and swapping providers later means re-validating all of it against a new baseline.
The mistake most enterprise architects make is treating this as a leaderboard lookup. General benchmarks — coding tests, broad knowledge quizzes, reasoning puzzles — measure capabilities that correlate weakly with the tasks SAP analytics work actually needs: extracting typed fields from a semi-structured S/4HANA BAPI response without inventing values that were not present, generating ABAP or SQLScript that respects SAP naming and performance conventions, reasoning correctly over a Datasphere Analytic Model's metadata, or drafting ESRS-aligned narrative from structured ESG figures. A model that tops a general leaderboard can still underperform on these narrow, high-stakes extraction and generation tasks, and the only reliable way to know is to run task-specific evaluation against real SAP data before committing.
Why it matters
- Structured-data extraction from S/4 BAPI JSON/XML has to be benchmarked on your own payloads — vendor leaderboards do not predict it, and no public SAP-specific leaderboard exists
- Only Llama 3 on SAP AI Core's EU-region deployment fully satisfies Art. 44 GDPR for high-sensitivity personal data — OpenAI and Gemini EU endpoints cover standard data only
- Code generation has to be evaluated per task — ABAP generation, long-context SQL reasoning and SQLScript rank models differently, and the ranking moves with every model generation
Key points
- Generic benchmarks (MMLU, HumanEval) do not predict SAP-specific task performance — run your own eval on ABAP generation, S/4 extraction, and Datasphere metadata tasks.
- OpenAI GPT-6 (Astra, Sol, Luna; GPT-6.1 Sol added 29 Sep 2026): the deepest Microsoft/Azure integration; evaluate it on ABAP generation and S/4 structured-data extraction.
- Claude (Opus 5.5, Sonnet 5; Haiku 4.5 for volume): 1M-token context on Opus and Sonnet; evaluate it on long-context SQL reasoning and instruction-following.
- Gemini (2.5 Pro and later): 1M-token context; evaluate it on full InfoProvider catalogue or large CSRD disclosure analysis.
- Llama 3 70B on SAP AI Core EU: best for EU data residency compliance and cost-at-scale; requires fine-tuning for SAP-specific tasks.
- Model selection is a 3-5 year architectural commitment; migration on underperformance costs 6-12 months.
Terms used on this page
- Foundation model
- Large pre-trained model (LLM or multimodal) used as a general-purpose base for downstream enterprise tasks via fine-tuning or prompting.
- GPT-6
- OpenAI's current model family (GPT-6: Astra, Sol, Luna; GPT-6.1 Sol added 29 September 2026; the GPT-5 line, released August 2025, is the prior generation); available via Azure OpenAI / Microsoft Foundry, including EU data-zone deployments.
- Claude
- Anthropic's model family — as of September 2026 Claude Opus 5.5 and Sonnet 5 (1M-token context) and Haiku 4.5 (200K); available via the Claude API, AWS Bedrock, Google Cloud Vertex AI and Microsoft Foundry.
- Gemini
- Google's model family; Gemini 2.5 Pro offers a 1M-token context window; available via Vertex AI with EU-region deployments.
- Llama 3
- Meta's open-weight LLM family (2024); self-hostable on EU infrastructure or available via SAP AI Core model catalogue.
- SAP BTP AI API
- SAP's managed API gateway for foundation model access on Business Technology Platform — provides unified endpoint for multiple models via SAP AI Core.
Sources
- Anthropic — Claude models overview (checked 2026-09-22)
- SAP AI Core — Supported Foundation Models Documentation
- Meta Llama 3 Model Card — Meta AI
- Google Vertex AI — Gemini Models Overview
- Azure OpenAI Service — models documentation
- SAP Business AI — official product page
- SAP Joule (work companion) — official product page
- SAP Generative AI — official product page
- Meta AI — Llama model research
- SAP Help Portal — Generative AI hub in SAP AI Core: model access overview
- SAP Help Portal — Orchestration service: content filtering and data masking pipeline modules
- Anthropic (via Claude documentation) — Context windows: current model limits (checked 2026-09-27)
- EUR-Lex — Regulation (EU) 2016/679 (GDPR), Article 44: general principle for data transfers
- The Decoder — GPT-6.1 Sol comes close to Astra at a fifth of the price (GPT-6.1 Sol pricing and OpenAI-reported benchmarks, 29 Sep 2026)
- InfoWorld — OpenAI pulls the plug on GPT 6.1 Astra as agents keep crossing lines (planned Astra release withheld over safety, 29 Sep 2026)
- The Decoder — Google Gemini 4 Argon closes the gap with OpenAI and Anthropic but doesn't take a clear lead (Argon pricing, Artificial Analysis scores, staged rollout)
- Databricks — October 2026 platform release notes (Gemini 2.5 Pro and Flash retired 2 Oct 2026)
- Databricks — September 2026 platform release notes (GPT-6.1 Sol on Unity Gateway, 29 Sep 2026)
- The Decoder — Google's new Gemini tiers cut free users to its weakest model and lock $5/month subscribers out of Pro (consumer plan limits, 4 Oct 2026)
Full card available to members. What the full card adds: the full decision framework · the SAP vs Snowflake / Databricks / Fabric comparison · the common pitfalls and their fix · the cheat sheet · the architecture schemas · the code blocks · the facts worth quoting.