BDC Connect for Databricks
As of 2026-09-27
What is BDC Connect for Databricks?
BDC Connect for Databricks is the zero-copy bridge that lets SAP Business Data Cloud and a Databricks Lakehouse read each other's governed data without either side copying, exporting, or replicating a single row.
How does SAP business data cloud work with Databricks?
SAP Business Data Cloud ships with an embedded Databricks workspace (SAP Databricks) and connects to it through BDC Connect, based on Delta Sharing: SAP data products become queryable Delta tables in Databricks, and Databricks tables can be published back as BDC data products, with no copy pipeline. Datasphere and SAP Analytics Cloud keep the semantic and reporting layers; Databricks takes data science, machine learning and engineering on the same governed data.
BDC Connect for Databricks is the zero-copy bridge that lets SAP Business Data Cloud and a Databricks Lakehouse read each other's governed data without either side copying, exporting, or replicating a single row. For any architect weighing a Databricks-heavy analytics estate against SAP's own semantic layer, it changes the question from "which platform wins" to "how do the two platforms share one governed truth."
What it is and why it matters
Historically, connecting SAP data to a Databricks Lakehouse — or vice versa — meant building and maintaining an extract-transform-load pipeline: a batch job, a staging area, a second copy of the data that could drift out of sync with the source and that needed its own access controls, its own monitoring, its own failure recovery. BDC Connect for Databricks removes that pipeline entirely by using the Delta Sharing open protocol, the same mechanism Databricks already uses to share Delta Lake tables across organisational boundaries, paired with SAP's own data product governance layer inside BDC. The result is not a connector in the traditional sense; it is a shared governance contract laid over two separate storage planes that never actually move data between them.
How it works
On one side, BDC exposes its governed data products — the semantic-layer-enriched views built on Datasphere or BW/4HANA — as Delta Share endpoints. Databricks Unity Catalog, the platform's metastore, registers that share using a share URL and a scoped bearer token, and from that point on a Databricks notebook can query the SAP data product directly with standard Spark SQL, with reads flowing straight from BDC-managed storage to the Databricks compute cluster. The reverse direction works the same way: a Databricks-managed Delta table can be registered as an external data product inside BDC, and then consumed from an SAP Analytics Cloud story or a BDC Analytical Application via a live connection, with the data staying in Databricks-managed cloud storage the entire time. Governance rides along with every read in both directions — a data product carries an owner, a freshness commitment, a sensitivity classification, and a subscriber list, and column-level masking rules defined on the SAP side are enforced even when the data is queried from the Databricks side.
When to use it, and the trade-off
The traditional alternative — batch replication into a staging layer, whether via classic ETL tooling or SAP's own Replication Flows — remains the right choice when the consuming system needs data transformed, reshaped, or heavily aggregated before use, or when the two platforms are not both able to speak the Delta Sharing protocol. BDC Connect for Databricks wins clearly when the requirement is read access to governed, current data with no transformation needed, and when both source and consumer already sit in ecosystems that support Delta Sharing natively — because it eliminates pipeline maintenance, removes replication lag entirely, and avoids the double-storage cost of keeping two copies. The trade-off is architectural lock-in to the Delta Sharing protocol and to Unity Catalog as the governing metastore on the Databricks side; a shop that is not already invested in Databricks gains little by choosing this path over a native SAP-only stack, and a shop with heavy transformation needs will still need a pipeline somewhere, just downstream of the zero-copy read rather than upstream of it. It is also worth weighing against BDC Connect for Snowflake or for Microsoft Fabric, which apply the same governance-contract idea to different lakehouse platforms — the right choice tracks wherever the organisation's actual compute estate already lives, not a platform preference in the abstract.
Pitfalls and anti-patterns
Teams sometimes treat BDC Connect as a substitute for data modelling discipline, assuming that because no data moves, no governance work is needed on the Databricks side — in practice, Unity Catalog permissions still have to be actively maintained, and a BDC data product with weak lineage metadata will surface as an equally weak, ungoverned table on the Databricks side. Another common mistake is underestimating the token and share-registration lifecycle: OAuth-scoped bearer tokens expire and need rotation, and a share that quietly stops refreshing produces silent staleness that looks, at first glance, like a live connection. A third pitfall is assuming zero-copy means zero-cost — every query against a shared Delta table still consumes compute on the reading side, and a heavy analytical workload against a BDC-exposed data product can generate real Databricks compute charges that were not budgeted for when the pipeline elimination was pitched as a pure saving.
The Joule angle, September 2026
BDC Connect for Databricks is a data bridge, not a model bridge — that distinction is the first thing to get right when a client asks how it relates to Joule or SAP AI Core. The connector moves governed rows between Datasphere/BW4HANA and Databricks Delta tables; it does not route inference calls, and SAP's own generative AI hub does not call a Databricks-hosted model through it. What it does enable is the federation pattern this platform's own review of Databricks Mosaic AI Agent Framework (C173) describes: a Joule Agent stays the process-aware orchestrator a business user actually talks to, and calls out to a Mosaic AI agent as a specialist tool when the reasoning genuinely needs Delta tables, a registered ML model, or a Vector Search index that lives natively on Databricks. This connector is the plumbing that lets that specialist tool read SAP actuals — say, S/4 general-ledger data — without an export step, and with SAP's Data Access Controls still governing which rows the Databricks-side reader can see.
The reverse direction matters just as much for AI work: when a Databricks-trained model (a demand-sensing score, a churn prediction) needs to reach a Joule conversation or an SAP Analytics Cloud story, the right pattern is to publish the model's output as a Delta table and register it as a BDC external data product going Databricks→BDC, not to retrain an equivalent model inside SAP AI Core. SAP's own tabular foundation models (SAP-RPT-1.5, and TabPFN-3.5-Plus since its 2026-09-15 GA in AI Core) solve a different problem — in-context prediction on structured data without a training step — and are not a reason to duplicate a Mosaic AI model that already works.
One governance gap worth flagging to a client building on the SAP Business AI Platform (which since Sapphire 2026-05-12 unifies BTP, BDC and Business AI under one governed environment): a Databricks-originated data product surfaced through this connector does not automatically appear as a node the SAP Knowledge Graph can reason over. Someone has to model it as a proper Datasphere entity before a Joule agent can ground an answer on it with the same confidence as native SAP data — treat that modeling step as a deliverable, not an assumption, when scoping a Joule-agent-calls-Mosaic-agent architecture.
Why it matters
- The direct SQL query example is concrete proof the integration is genuinely zero-copy, not a marketing claim — reads go straight from BDC-managed storage to the Databricks compute cluster.
- OAuth 2.0 bearer-token registration per BDC tenant is the actual security mechanism — a specific detail worth verifying against a client's IAM model before scoping.
- Governance (owner, freshness SLA, classification) applies to every Data Product published either direction — this is a governed marketplace, not a raw data pipe.
Key points
- BDC Connect for Databricks is a zero-copy, bidirectional bridge built on the Delta Sharing open protocol — a Delta table becomes a queryable BDC Data Product and vice versa.
- No ETL pipeline, no staging copy: data stays in its authoritative store, so there is no second copy to secure, monitor or reconcile.
- It changes the architecture question from 'which platform wins' to 'how do both platforms share one governed truth'.
- Access control is enforced by the sharing side — model the data product's permissions before exposing it, not after.
- Reached general availability on 2025-10-06, the first production zero-copy bidirectional channel between BDC and a Databricks Lakehouse.
- Share registration on the Databricks side uses an OAuth 2.0 bearer token scoped per BDC tenant — rotate it on a schedule, don't treat it as a one-time setup step.
- It is a data bridge, not a model bridge: it does not route generative AI hub inference calls to Databricks, and vice versa — model outputs cross the boundary only as published Delta tables or BDC data products.
- It requires the consuming Databricks workspace to run on Unity Catalog; a workspace still on the classic Hive metastore cannot consume a Delta Share until migrated.
Common pitfalls
- Confusing zero-copy with real-time streaming — Signal: The client asks 'does BDC Connect give us live SAP data in Databricks?' and the consultant answers 'yes' without qualification. The client then designs a real-time fraud detection pipeline expecting sub-second SAP data updates from BDC Connect. Fix: BDC Connect via Delta Sharing is read-on-demand — a Databricks query reads from BDC storage at query time, with freshness governed by BDC's own refresh cadence (which for a Datasphere Analytical Dataset may be hourly or daily, not sub-second). For real-time streaming of SAP table changes, the correct architecture is SAP Change Data Capture (CDC) + Databricks Structured Streaming, not BDC Connect.
- Recommending BDC Connect to a client without an existing Databricks investment — Signal: The proposal slides show BDC Connect as an architecture component for a client whose analytics estate is SAP-only (BW/4HANA + SAC), because the architecture looks elegant in a diagram. Fix: BDC Connect adds Databricks platform licensing costs (DBUs) and Unity Catalog configuration overhead to a client who gains nothing from it. For SAP-only estates, Datasphere-native integration is the right answer. BDC Connect is justified only when Databricks is already a strategic platform commitment.
- Skipping the Unity Catalog setup assessment — Signal: The engagement starts with BDC Connect configuration, then discovers the client's Databricks workspace is on the classic Hive metastore, not Unity Catalog. Delta Sharing as a consumer requires Unity Catalog. The project stalls for 4-6 weeks while Unity Catalog is migrated. Fix: At project initiation, confirm the client's Databricks workspace is Unity Catalog-enabled. Unity Catalog migration is a non-trivial change (existing table references, access control policies, and notebook code must be updated). Scope this as a prerequisite workstream or disqualify BDC Connect as an option if the migration timeline is blocked.
- Ignoring the double-governance complexity in regulated industries — Signal: The architecture diagram shows BDC governance on one side and Databricks Unity Catalog governance on the other, described as a strength. In practice, the client's data governance team is not resourced to maintain two separate access control frameworks synchronized — the result is access decisions that contradict each other. Fix: Document which governance layer is the master for each access decision. The recommended pattern: BDC governs data classification and field masking (GDPR, sensitivity); Unity Catalog governs Databricks-user-level access (which team sees which shared table). Define the governance RACI before the first BDC Data Product is published.
- Equating Delta Sharing with Delta Lake feature parity — Signal: The Databricks team assumes that because BDC Connect uses Delta Sharing, they can use Delta Lake features (time travel, MERGE, OPTIMIZE) on the shared BDC Data Product in their Databricks environment. Fix: Delta Sharing is a read-only protocol over Delta Lake tables. A Databricks consumer of a BDC Delta Share can query the current version of the shared table, but cannot perform time travel (accessing historical versions), MERGE, or OPTIMIZE on the shared data — those operations are controlled by the BDC side. Databricks consumers can cache the shared table into their own Delta Lake if they need time-travel or write capabilities, at the cost of creating a copy.
Decision framework
| Decision | Option A | Choose A when | Option B | Choose B when |
|---|---|---|---|---|
| BDC Connect vs SLT replication for Databricks access to SAP data | BDC Connect (zero-copy Delta Sharing) | Use when Databricks needs business-ready, governed SAP analytics data (KPIs, financial aggregates, HR planning). Avoids ETL pipelines, eliminates the 'which version is authoritative' problem, and preserves BDC governance on every read. Requires BDC in the architecture (HANA Cloud or Datasphere semantic layer). | SLT replication + Databricks-native ingestion | Use when Databricks needs high-frequency, low-latency streams of raw SAP table changes for ML feature engineering (e.g., live order-line changes for demand-sensing models). SLT delivers change-data-capture; BDC Connect does not stream raw table changes. Also use when BDC is not in the client's roadmap. |
| When to recommend BDC Connect vs Datasphere-only integration | BDC Connect | Client has an existing Databricks investment (active workloads, licensed DBUs, data engineering team). The business case requires ML models trained in Databricks to consume SAP semantics, or Databricks analysts to query SAP KPIs without learning a new tool. Platform licensing cost is already carried. | Datasphere-native (Direct Access, BW Bridge, or HANA Smart Data Access) | Client has an SAP-only analytics estate or is evaluating Databricks but has not committed. Datasphere-native integration is architecturally simpler, has a longer track record, and does not require the Databricks Unity Catalog setup and Delta Sharing configuration overhead. |
| Governance model for a shared BDC Data Product | BDC Data Product with column masking (sensitive fields) | Use when the Data Product contains fields subject to GDPR, internal access control, or regulatory classification (cost data, personal identifiers, compensation data). BDC applies masking at the endpoint before Delta Sharing; Databricks consumers never see the masked values regardless of their Unity Catalog permissions. | BDC Data Product with full field exposure | Use for operational KPIs and aggregated analytics data with no field-level sensitivity (product revenue by region, inventory levels, order backlog by category). Simpler to configure and reduces query-time overhead; appropriate when the Unity Catalog subscriber list already limits access to cleared users. |
| OEM Databricks licensing vs separate Databricks subscription | SAP OEM path (Databricks bundled in BDC contract) | Use when the client wants a single vendor contract and SAP is the primary relationship. The $250M SAP–Databricks partnership makes this commercially viable; SAP resells Databricks DBUs as part of the BDC CU bundle for qualified deals. Simplifies procurement but limits Databricks edition choice. | Direct Databricks subscription + BDC Connect integration | Use when the client has an existing Databricks Enterprise or Premium contract with negotiated rates. Connecting BDC Connect to an existing Databricks workspace is technically straightforward; the client registers the BDC Delta Share in their existing Unity Catalog and the commercial relationship stays separate. |
| BDC Connect read direction for the primary use case | BDC Data Product consumed in Databricks (SAP → Databricks direction) | The most common production pattern in 2026: Databricks data scientists and analysts query SAP KPIs (revenue, inventory, workforce data) from BDC as a Delta Share, combining them with Databricks-native data sources (web analytics, IoT, third-party market data) for ML and advanced analytics. | Databricks Delta table consumed in BDC/SAC (Databricks → SAP direction) | Use when the client's ML or data engineering output (trained model scores, external market data enrichments, IoT sensor aggregates) needs to feed back into SAP Analytics Cloud dashboards or BDC Analytical Applications. Databricks registers the output table as a BDC external Data Product; SAC reads it via a live BDC connection. |
How SAP compares
| Capability | SAP | Snowflake | Databricks | Microsoft Fabric |
|---|---|---|---|---|
| Zero-copy mechanism to BDC | Every BDC connector shares the same source-side unit: a governed Data Product published from Datasphere or BW4HANA, with Datasphere staying the authoritative owner regardless of which lakehouse reads it. | BDC Connect for Snowflake (C168) — same governance-contract pattern, built on Snowflake's own secure data sharing rather than Delta Sharing; no Unity Catalog dependency. | BDC Connect for Databricks (this card) — Delta Sharing, Unity Catalog required on the consuming side, OAuth 2.0 bearer token scoped per BDC tenant. | BDC Connect for Fabric (C167) — OneLake Mirroring for SAP; a Datasphere table surfaces as a read-only OneLake shortcut, near-real-time rather than Delta Sharing's on-demand read. |
| Commercial relationship | — | Partner-connected: dual SAP + Snowflake commercial relationship (C168/C169); Snowflake credits billed separately from any SAP contract. | OEM-inside-BDC: SAP can resell Databricks DBUs as part of the BDC Capacity Unit bundle for qualified deals — the tightest commercial integration of the three. | Additive, not bundled: BDC Capacity Units (SAP) + Fabric F-SKU (Microsoft), billed separately with no cross-vendor bundle. |
| Best-fit AI/ML workload | — | High-concurrency, storage/compute-separated BI and reporting; Cortex Analyst/Agents (C175/C176) run on the Snowflake side. | ML training, feature engineering and Mosaic AI agent tool calls (C173) that need SAP data alongside non-SAP lakehouse data in the same job — the strongest ML-in-agent story of the three. | Power BI Copilot-centric workloads (C177) wanting governed SAP actuals without migrating the BI layer to SAC. |
| Read latency profile | — | Read-on-demand at query time, comparable to the Databricks pattern. | Read-on-demand at query time (seconds) — not a continuous mirror. | Near-real-time OneLake mirror (seconds to minutes) — closer to continuous refresh than the Databricks or Snowflake on-demand read. |
Facts worth quoting
- BDC Connect for Databricks reached general availability on 2025-10-06, the first production zero-copy bidirectional channel between SAP Business Data Cloud and a Databricks Lakehouse.
- The integration reuses the Delta Sharing open protocol unmodified — the same reader mechanism Databricks already uses for Databricks-to-Databricks sharing across organisational boundaries, so no separate reader implementation is required on either side.
- Share registration on the Databricks side uses an OAuth 2.0 bearer token scoped per BDC tenant, which means access to a given share can be rotated or revoked independently of the underlying SAP data product.
Sources
- Databricks — Announcing GA of SAP BDC Connect to Databricks
- SAP — BDC Connect for Databricks technical reference
- Constellation Research — SAP × Databricks launch coverage
- Sharing SAP S/4HANA Data with Databricks Using BDC Connect and Delta Sharing — SAP Community (Technology Blog Posts by Members)
- Why SAP Databricks Genie Is a Game‑Changer: Key Benefits and Expert Tips — SAP Community (Technology Blog Posts by SAP)
- How to use SAP Business Data Cloud Capacity Unit Estimator for SAP BDC Connect for Databricks? — SAP Community (Technology Blog Posts by SAP)
- Triggering SAP Databricks Jobs through SAP Datasphere Task Chains — SAP Community (Technology Blog Posts by SAP)
- Top SQL Queries Every SAP Databricks Admins Should Keep in your Back Pocket — SAP Community (Technology Blog Posts by SAP)
- Integrating SAP Databricks with SAP CPQ and SAP Datasphere for Analytics & Reporting — SAP Community (Technology Blog Posts by SAP)
- SAP Databricks - Data Classification : A Practical Guide for Modern Data Governance — SAP Community (Technology Blog Posts by SAP)
- Where ORD Is Used in SAP Datasphere, SAP Business Data Cloud & SAP Databricks — SAP Community (Technology Blog Posts by SAP)
- Integrating SAP Databricks SCIM API with SAP Identity Services — SAP Community (Technology Blog Posts by SAP)
- Breaking SAP Data Barriers with Datasphere & Databricks: A Medallion Journey (Bronze, Silver, Gold) — SAP Community (Technology Blog Posts by SAP)
- SAP Databricks – Best Practices SQL Query Performance Tuning Tips — SAP Community (Technology Blog Posts by SAP)
- Automating Invoice Predictions from SAP Databricks with SAP Datasphere Task Chain — SAP Community (Integration Blog Posts)
- Hands-on Tutorial (SAP) Databricks triggering ML in SAP Datasphere — SAP Community (Technology Blog Posts by SAP)
- SAP Databricks Alerts 🔔 : Send Email Notification for Real‑Time Monitoring Made Simple — SAP Community (Technology Blog Posts by SAP)
- How to Store and Manage SAP Secrets Using SAP Databricks CLI: A Developer’s Guide — SAP Community (Technology Blog Posts by SAP)
- Developing HANA ML models with SAP Databricks — SAP Community (Technology Blog Posts by SAP)
- SAP Databricks CLI: Unlocking Automation and Efficiency in Data Engineering — SAP Community (Technology Blog Posts by SAP)
- SAP Databricks : Query History [ Deep Dive Into Its Features, Benefits, and Practical Usefulness ] — SAP Community (Technology Blog Posts by SAP)
- SAP Databricks - Notebook supported Programming languages — SAP Community (Technology Blog Posts by SAP)
- Connecting SAP Analytics Cloud to Databricks model serving endpoint — SAP Community (Technology Blog Posts by SAP)
- Expose BW 7.5 objects as Data products in Databricks — SAP Community (Technology Blog Posts by SAP)
- How to provision SAP BDC Connect for Databricks? — SAP Community (Technology Blog Posts by SAP)
- Integrating Non-SAP Semi-Structured Invoices Data with SAP BDC Using SAP Databricks — SAP Community (Integration Blog Posts)
- New Customer Influence Sessions for SAP Business Data Cloud — SAP Community (Technology Blog Posts by SAP)
- SAP Business Data Cloud – Leveraging SAP Databricks’ assistant to easily build predictive scenarios — SAP Community (Technology Blog Posts by SAP)
Guides that answer with this page
These guides cite this page as one of the sources their answer rests on.