AI & Analytics Legends The knowledge platform for SAP Analytics
Academy module

Vector Databases in SAP Context: the HANA Cloud Vector Engine

Vector retrieval architecture in SAP context: SAP sources flow to embedding, to HANA Cloud vector store, to hybrid dense and sparse retrieval, to RAG and Joule consumption, with a scale threshold decision to an external ANN store — architecture diagram for Vector Databases in SAP Context: the HANA Cloud Vector Engine, Analytics Legends Academy module M137

As of 2026-10-04

In SAP landscapes vectors belong where they stay governed: the SAP HANA Cloud vector engine, now positioned as the AI database of SAP Business Data Cloud. It offers REAL_VECTOR (1–65,000 dimensions) and HALF_VECTOR columns, COSINE_SIMILARITY and L2DISTANCE, in-database embeddings with VECTOR_EMBEDDING (model SAP_NEB.20240715, 768 dimensions, five languages, NLP option required), HNSW approximate-nearest-neighbour indexes and CROSS_ENCODE reranking. Earlier claims of 'full scan only' are obsolete. Choose the managed document grounding service for documents, self-managed HANA Cloud tables when you need permission filters, hybrid search or reranking, and an external store only for a specific requirement. Measure Recall@k, MRR, latency and leakage on your own data.

What you will learn

  • Create a column table with a REAL_VECTOR or HALF_VECTOR column, populate it with VECTOR_EMBEDDING and query it with COSINE_SIMILARITY and L2DISTANCE, using correct function names.
  • Create an HNSW vector index with explicit BUILD and SEARCH CONFIGURATION and report the measured change in latency and Recall@5 on a 20-query test set.
  • Implement permission-filtered retrieval where the permission attributes are applied in the same SQL statement as the similarity search, with zero leakage in tests.
  • Build hybrid retrieval combining vector similarity with HANA full-text search and reciprocal rank fusion or CROSS_ENCODE reranking, and compare MRR against vector-only.
  • Decide between the managed document grounding service, self-managed HANA Cloud tables and an external vector store for a given scenario, with a written rationale.

Module overview

A vector database stores embeddings — numeric representations of text, images or records — and returns the items closest to a query vector. That capability underpins semantic search, retrieval-augmented generation (RAG), deduplication and recommendation. In an SAP landscape the question is rarely "which vector database is best in general" but "where should vectors live so that they stay governed, secure and close to the business data". Since 2024 SAP's answer is the SAP HANA Cloud vector engine, and in 2026 SAP positions SAP HANA Cloud as "the AI database" of SAP Business Data Cloud. This module teaches the engine's actual SQL, the design decisions that make retrieval work on SAP data, and when an SAP-managed or external alternative is the better choice.

Prerequisites

  • SQL on SAP HANA (column tables, full-text search basics)
  • Review core concepts first: C154, C319, C027
  • Recommended: module M136 (LLMs on Enterprise Data)

Outcomes

  • Write correct HANA Cloud vector SQL: types, similarity functions, VECTOR_EMBEDDING, HNSW index.
  • Design chunking and metadata for structured SAP objects and documents.
  • Enforce permission filters inside retrieval and prove it with tests.
  • Evaluate retrieval with Recall@k, MRR, latency and leakage on customer data.
  • Correct outdated claims (no ANN index, no in-database embeddings) in customer discussions.

Full module available to members. The full module adds: the decision framework · the end-to-end scenario walkthrough · the KPI scorecard · the anti-patterns · the code blocks · the knowledge check · the diagrams.

Open in the app →