AI & Analytics Legends The knowledge platform for SAP Analytics
Concept card

Databricks

Databricks — Analytics Legends section illustration for the SAP Analytics knowledge base (concepts, studies, Academy)

As of 2026-09-27

What is Databricks?

Databricks is a data and AI platform built around the lakehouse idea: one storage layer holding open table formats, with engines for SQL, streaming, data science and machine learning reading the same tables instead of each keeping a copy.

Databricks is a data and AI platform built around the lakehouse idea: one storage layer holding open table formats, with engines for SQL, streaming, data science and machine learning reading the same tables instead of each keeping a copy. For most of its life it was, from an SAP practice's point of view, a neighbouring stack — something the customer also owned, that someone else administered.

That changed when SAP embedded it in Business Data Cloud. BDC ships a Databricks capability inside the SAP estate, and the significance for a consultant is not the technology: it is that the boundary moved. Questions that used to be answered with "that lives on the data platform side" now land inside an SAP conversation, and the consultant in the room is expected to have a view.

What it is actually good at, stated plainly. Databricks earns its place where the work is large-scale transformation, streaming ingestion, and machine learning on data that is not only SAP — the workloads a semantic modelling layer is not designed to carry. SAP Datasphere earns its place where the work is business semantics, governed reuse and a modelling paradigm SAP data already speaks. Neither replaces the other, and an architecture that treats the choice as either/or usually ends up rebuilding the semantic layer by hand in notebooks.

The joint is the data, not the tool. What makes the pairing workable is open sharing of tables between the two sides rather than another copy of the data: the SAP side keeps the semantics and the governance, the lakehouse side keeps the scale and the ML surface. The failure mode worth naming is duplication — the moment the same business entity is defined once in a semantic model and again in a notebook, the two definitions drift, and the question "which number is right" has no owner.

Where a consultant's value sits. Not in operating clusters. It sits in being the person who can say which questions belong on which side, what crosses the boundary, and who owns the definition when it does. That is a semantics and governance conversation, and it is the one an SAP analytics consultant is better placed to hold than anyone else in the room.

The pieces worth knowing by name. Underneath the marketing the platform is four things a consultant should be able to place. Delta Lake is the open table format that gives files on object storage transactional behaviour — the reason a lakehouse can be read and written safely by several engines at once. Unity Catalog is the governance layer: catalogs, schemas, tables, ownership, access and lineage in one place, and the surface on which any credible "who may see this" conversation happens. Delta Sharing is an open protocol for handing another platform live read access to a table without copying it. Mosaic AI and MLflow are the model side — training, tracking, serving. Everything else in a Databricks estate hangs off those four.

How SAP actually connected it. SAP launched Business Data Cloud in February 2025 as a joint offering with Databricks, with an embedded SAP Databricks inside the SAP estate, and later shipped BDC Connect for customers who already run their own Databricks workspace. Both routes lean on the same idea: share tables rather than copy them. What matters to a consultant is the detail underneath — SAP's semantic metadata can travel with the shared data, so the lakehouse side can read the SAP business context instead of re-deriving it from field names. That is the difference between a share and an export, and it is the whole argument for doing it this way.

The decision, and the trade-off it hides. The question a customer asks is "Datasphere or Databricks"; the question worth answering is which of their actual questions belongs on which side, and who owns each business definition once the boundary is drawn. The trade-off is real: put a definition on the lakehouse side and you gain scale and ML reach but leave the SAP semantic layer; keep it on the SAP side and you gain governed reuse but pay for movement when the ML workload needs it. What is not a trade-off is defining it twice — that is not a choice, it is a defect that shows up in an audit six months later.

The Joule angle, September 2026

The semantics-versus-scale boundary this card names is the same boundary that decides which AI surface reasons over which side of a joint SAP-Databricks estate, and a consultant should draw both boundaries in the same conversation rather than treating AI as a separate workstream. On the Databricks side, the Mosaic AI Agent Framework builds multi-step agents natively against Unity Catalog-governed tables, registered models and Vector Search indexes, and Genie answers single-turn natural-language questions against Databricks-resident data using free-text definitions and trusted SQL examples — both read the lakehouse directly, with no dependency on SAP's own semantic layer. On the SAP side, Joule Agents ground reasoning in the SAP Knowledge Graph, which is built from data modelled as a governed data product inside Datasphere, not from raw lakehouse tables. This means the "who owns the definition" question this card already frames for BI has an exact AI-layer counterpart: a Joule Agent answering a question needs the SAP-side semantic model to exist as a Datasphere data product, and a Mosaic AI agent answering the same underlying business question from the lakehouse side needs no such SAP-side artifact at all — the two AI surfaces do not automatically converge just because BDC lets both sides share the underlying tables.

The practical consequence for scoping AI work on a joint SAP-Databricks estate: route by workload shape, the same logic this card already applies to BI. A question that needs SAP process context — approval status, document lifecycle, organisational hierarchy — belongs with a Joule Agent reading the Knowledge Graph, because Genie and Mosaic AI agents have no equivalent structured model of SAP business processes. A question that is fundamentally an ML-training or large-scale feature-engineering task — a churn model, a demand forecast retrained nightly on terabytes of transactional history — belongs on Mosaic AI, because Joule Agents are not built for model training and serving at that scale. Getting this routing wrong in either direction reproduces, at the AI layer, the exact duplication failure this card already names for semantics: define the same business logic once inside a Joule Agent's grounding and again inside a Mosaic AI agent's tool code, and the two will drift, with nobody owning which one is authoritative.

The honest caveat for anyone quoting Databricks' AI capability set in an SAP conversation: this platform's own fact base does not carry hard GA dates or version numbers for Mosaic AI Agent Framework or Genie the way it does for SAP's own Joule releases, because Databricks' own release cadence is not SAP-authored content — check Databricks' current documentation directly before quoting a specific capability as shipped, rather than treating this card's naming of them as a status confirmation.

Why it matters

  • Since BDC embedded it, Databricks is no longer a stack an SAP consultant can decline to have a view on — the boundary moved inside the SAP estate.
  • It is still a thin skill in the market this platform tracks: 244 of the 9,116 published firm profiles mention Databricks against 4,184 that mention S/4HANA, and 60 of 3,307 live opportunity rows name it (measured 2026-08-29 on public/api/company-profiles.json and public/api/contracts.json).
  • The decision it forces — semantics here, scale there, who owns the definition — is exactly the one that has no owner by default, which is why it is worth being the person in the room who can settle it.

Key points

  • Lakehouse platform: one open-format storage layer, several engines reading the same tables instead of each keeping a copy.
  • Four pieces to place by name: Delta Lake (open table format), Unity Catalog (governance), Delta Sharing (open sharing protocol), Mosaic AI/MLflow (model lifecycle).
  • SAP launched Business Data Cloud with Databricks in February 2025 and embedded an SAP Databricks in the SAP estate; BDC Connect covers customers who already own a workspace.
  • The joint is a SHARE, not a copy — and SAP's semantic metadata can travel with the shared data, so business context is not re-derived from field names.
  • It carries large-scale transformation, streaming and ML on data that is not only SAP; Datasphere carries business semantics and governed reuse.
  • Neither replaces the other; treating the choice as either/or usually means rebuilding the semantic layer by hand in notebooks.
  • The failure mode is DUPLICATION of a business definition across both sides — after which 'which number is right' has no owner.
  • A consultant's value here is boundary work — routing questions, naming what crosses, assigning ownership — not operating clusters.

Common pitfalls

  • Defining the same business entity twice — Signal: "Revenue" exists in a semantic model AND in a notebook Fix: One definition, one owner, one side. The other side consumes it — it does not re-derive it.
  • Copying instead of sharing — Signal: A third copy of the same tables lands in the lakehouse "for performance" Fix: Start from a share. A copy needs a named owner, a refresh contract and a reason that survives review.
  • Framing it as Datasphere versus Databricks — Signal: A slide with two boxes and one arrow between them Fix: Sort the customer's real questions into two columns instead. The sort is the answer.
  • Exporting data without its semantics — Signal: Lakehouse notebooks reconstructing meaning from SAP field names Fix: Share the semantic metadata alongside the tables — that capability is the point of the joint.
  • Buying a lakehouse for a modelling problem — Signal: A programme whose only heavy workload is a monthly aggregation Fix: Scale, streaming and ML justify it. A modelling problem does not, and the running cost will say so.
  • Leaving governance on one side only — Signal: Access rules in Unity Catalog that no one has reconciled with the SAP-side authorisations Fix: Both catalogues describe the same entities. Reconcile them explicitly, or the audit will do it for you.

Decision framework

Decision framework
DecisionOption AChoose A whenOption BChoose B when
Which side answers this question?Lakehouse (Databricks)Large-scale transformation, streaming ingestion, ML, or data that is not only SAPDatasphere / BDC semantic layerBusiness semantics, governed reuse, definitions the business already recognises
Embedded SAP Databricks or your own workspace?SAP Databricks inside BDCNo incumbent lakehouse, and the value is in staying inside one commercial and governance perimeterOwn workspace via BDC ConnectAn established Databricks estate, its own governance and FinOps, and non-SAP data already landed there
How does data cross the boundary?Share the tables (Delta Sharing)The default — one copy, live, with the SAP semantic metadata travelling alongsideReplicate a copyOnly where a hard latency, residency or availability constraint forbids a share — and then as a declared, owned copy
Who owns a business definition that both sides need?One owner on the SAP sideThe definition is a business rule the enterprise already governs — the lakehouse consumes itOne owner on the lakehouse sideThe definition is genuinely derived there (a model output, a computed feature) — then SAP consumes it back
The customer asks 'Datasphere or Databricks?'Answer with a product comparisonAlmost never useful — it turns an ownership question into a feature argumentAnswer with their questions sorted into two columnsThe sort IS the deliverable, and it usually shows the argument was never about the tools

How SAP compares

How SAP compares
CapabilitySAPSnowflakeDatabricksMicrosoft Fabric
Where SAP business semantics live nativelyNative in Datasphere and BDC — the semantic model is the product.Not native. Partner-connected, so SAP context arrives only as far as the connection carries it.Embedded inside BDC as SAP Databricks, so SAP semantics are reachable from the lakehouse rather than rebuilt.Not native. Mirrored into OneLake, with the business meaning modelled on the Fabric side.
Commercial shape with SAPIt is SAP — one contract, one perimeter.Partner-connected architecture with a dual commercial relationship and Snowflake-credit pricing.OEM inside BDC (SAP Databricks), plus BDC Connect for a customer-owned workspace.Connected through BDC Connect for Fabric, using OneLake mirroring.
Large-scale transformation, streaming and MLNot what a semantic modelling layer is designed to carry.Yes, on its own engine and its own governance.Its core ground — this is the workload the platform was built around.Yes, within the Microsoft stack and its tooling.
Open table format and sharingConsumes and publishes shares inside BDC rather than owning a table format.Iceberg tables plus its own sharing model.Delta Lake as the table format, Delta Sharing as an open protocol — the reason the joint can be a share rather than a copy.Delta/Parquet in OneLake, reached through shortcuts and mirroring.
Governance surfaceThe BDC/Datasphere catalogue and the SAP-side authorisations.Its own role-based access model.Unity Catalog — ownership, access and lineage in one place.The Microsoft governance stack around OneLake.
Failure mode when it stands aloneA heavy ML workload squeezed into a modelling tool, then blamed on the tool.SAP semantics re-derived from field names, once per project.The semantic layer rebuilt by hand in notebooks, and a second definition nobody owns.Mirrored data with no owned definitions behind it.

Facts worth quoting

  • 244 of the 9,116 firm profiles measured as of 2026-09-16 the directory held as of 2026-09-16 (2.7%) mention Databricks, against 4,184 (46.0%) that mention S/4HANA — the skill is real in the market but far thinner than the ERP it is being attached to (measured 2026-08-29, source: public/api/company-profiles.json).
  • 60 of the 3,307 live opportunity rows on this platform's radar name Databricks, against 267 that name S/4HANA — a ratio of roughly 1 to 4.5 (measured 2026-08-29, source: public/api/contracts.json).
  • 53 of the 4,130 news rows in this platform's review window mention Databricks (measured 2026-08-29, source: public/api/news.json).
  • 53 of the 330 concept cards in this corpus reference Databricks — including nine dedicated cards on the BDC/Databricks joint, its connect surface and its agent tooling (measured 2026-08-29, source: src/data/concepts-100.json).
  • SAP launched Business Data Cloud as a joint offering with Databricks in February 2025; BDC Connect for Databricks reached general availability afterwards (source: Databricks and Constellation Research announcements cited below — this card states no revenue, seat or capacity figure, none of which could be sourced to a primary document on 2026-08-29).

Sources

  1. Databricks — Lakehouse platform
  2. Databricks Docs — Lakehouse architecture
  3. Databricks Docs — Unity Catalog data governance
  4. Databricks Docs — Data sharing (Delta Sharing in Unity Catalog)
  5. Databricks — Unity Catalog product page
  6. Delta Sharing — open protocol (Linux Foundation)
  7. Databricks — Announcing general availability of SAP Business Data Cloud Connect to Databricks
  8. Databricks — Unlocking SAP Business Context in Databricks with Semantic Metadata Delta Sharing
  9. Databricks Marketplace + Delta Sharing
  10. SAP Help Portal — BDC Connect for Databricks
  11. Constellation Research — SAP launches Business Data Cloud in partnership with Databricks: what it means
  12. SAP Business Data Cloud Architecture: Datasphere, Databricks and the Unified Data Layer — SAP Community
  13. How to provision SAP BDC Connect for Databricks? — SAP Community
  14. Sharing SAP S/4HANA Data with Databricks Using BDC Connect and Delta Sharing — SAP Community
  15. Breaking SAP Data Barriers with Datasphere & Databricks: A Medallion Journey — SAP Community
  16. Get started with SAP Databricks: Introduction — SAP Community
  17. SAP Databricks is now GA — skilling yourself with Mosaic AI — SAP Community
  18. Triggering SAP Databricks Jobs through SAP Datasphere Task Chains — SAP Community
  19. The Added Value of SAP Business Data Cloud with Databricks, Snowflake and MS Fabric — SAP Community
  20. Expose BW 7.5 objects as Data Products in Databricks — SAP Community
  21. SAP Sapphire Orlando 2025: SAP and Databricks open a bold new era of data and AI — SAP Community
Open in the app →