Embeddings
As of 2026-09-27
What is Embeddings?
Sentence embeddings power RAG by matching meaning, not syntax — but the trade-off is stark: a large model (3,072d, $0.13/M) versus small (1,536d, $0.02/M) versus a local 384d model, and mixing models in one index silently breaks retrieval.
What it is
An embedding is a dense real-valued vector that encodes the semantic content of a token, sentence, or document in a high-dimensional space, such that semantically similar inputs cluster geometrically. Embeddings are the bridge between discrete text and the continuous mathematics of neural networks.
The problem they solve is the inability of exact-match search to find 'net revenue' when the user asks 'total income' — embeddings encode meaning, not syntax, enabling semantic retrieval at scale.
Three generations have dominated: static word embeddings (Word2Vec, Mikolov 2013; GloVe) assign one vector per word regardless of context; contextual embeddings (BERT, GPT) produce a different vector for each token depending on surrounding context; sentence embeddings average or pool contextual representations into a single document vector. Modern RAG systems use sentence embeddings. OpenAI's text-embedding-3-large produces 3,072-dimensional vectors at $0.13/M tokens; text-embedding-3-small produces 1,536d at $0.02/M. Matryoshka Representation Learning (MRL, Kusupati 2022) trains embeddings so that the first k dimensions of a 3,072d vector are a usable 256d embedding — enabling storage compression with only ~5% NDCG@10 degradation. SAP HANA Cloud's vector engine natively stores 1,536d cosine-similarity vectors and returns top-k results under 10ms at 1M documents.
Choose text-embedding-3-large when retrieval quality is critical (financial reports, regulatory text); choose text-embedding-3-small or a local model (all-MiniLM-L6, 384d) when cost or latency dominates. Never mix embeddings from different models in the same index — the vector spaces are incompatible.
Why it matters
- Matryoshka Representation Learning compresses a 3,072d vector down to a usable 256d with only ~5% NDCG@10 loss — a concrete storage-cost lever.
- SAP HANA Cloud's vector engine returns top-k results under 10ms at 1M documents with native 1,536d cosine similarity — a real performance baseline to quote.
- Model incompatibility across an index is an unforced error: embeddings from different models occupy incompatible vector spaces.
Key points
- Sentence embeddings (not word embeddings) are the correct unit for RAG over SAP documents — one vector per paragraph, not per token.
- Cosine similarity, not Euclidean distance, is the standard metric for embedding similarity; HANA Cloud uses cosine by default.
- Matryoshka embeddings allow truncation from 3,072d to 256d with only ~5% NDCG@10 loss — large storage savings.
- Never mix embeddings from different models in the same index; the vector spaces are not aligned.
- HANA Cloud vector engine: 1,536d cosine, top-k < 10ms at 1M docs — no external vector DB required for most SAP RAG workloads.
- Pin the embedding model version explicitly in the ingestion pipeline configuration — an unpinned dependency can silently switch models on a routine library upgrade, producing a corrupted, mixed-model index with no error thrown.
- HANA Cloud's Vector Engine is not the only retrieval mechanism in the same database — a native Knowledge Graph Engine (RDF/SPARQL) grounds relationship-shaped questions better than embedding similarity ever will.
- A cosine similarity score is a ranking signal, not a calibrated confidence measure — never present a raw similarity number to a business user as 'how sure the model is'.
Terms used on this page
- Embedding
- A dense real-valued vector representing a text unit in a learned semantic space.
- Cosine similarity
- cos(θ) = (A·B)/(|A||B|): measures angle between two vectors; 1.0 = identical direction, 0 = orthogonal, −1 = opposite.
- Matryoshka Representation Learning (MRL)
- Training technique that makes the first k dimensions of a large embedding usable as a standalone smaller embedding, enabling flexible dimensionality.
- RAG
- Retrieval-Augmented Generation: architecture that embeds a document corpus, retrieves the top-k relevant chunks at query time, and appends them to the LLM prompt.
- Vector index
- An approximate nearest-neighbour (ANN) index (e.g. HNSW, IVF) that enables sub-linear search over large embedding collections.
Sources
- Mikolov et al. 2013 — Efficient Estimation of Word Representations (Word2Vec)
- Su et al. 2021 — RoFormer: Enhanced Transformer with Rotary Position Embedding
- Press et al. 2022 — Train Short, Test Long: ALiBi
- Kusupati et al. 2022 — Matryoshka Representation Learning
- OpenAI text-embedding-3 model card
- SAP HANA Cloud vector engine documentation
- Peng et al. 2023 — YaRN: Efficient Context Window Extension of LLMs
- SAP Sapphire Orlando 2026 - AI database for AI agents and apps: Overview of SAP HANA Cloud — SAP Community (Technology Blog Posts by SAP)
- Unlocking custom AI use-cases with SAP HANA Cloud — SAP Community (Technology Blog Posts by SAP)
- SAP HANA Cloud: Expert-Guided Implementation Workshop Series — SAP Community (Technology Blog Posts by SAP)
- What’s New in SAP HANA Cloud – March 2026 — SAP Community (Technology Blog Posts by SAP)
- A Use Case for HANA Cloud Knowledge Graph: AI‑Driven Tender Analysis — SAP Community (Technology Blog Posts by Members)
- SAP Generative AI Hub: RAG on SAP Data with HANA Cloud Vector Store (Part 4 of 6) — SAP Community (Artificial Intelligence Blogs Posts)
- BM25 and Hybrid Search for RAG on SAP HANA Cloud (without PAL) — SAP Community (Artificial Intelligence Blogs Posts)
- Migrating to SAP HANA Cloud: What Actually Gets Better (Part 1 of 2) — SAP Community (Technology Blog Posts by SAP)
- LangGraph Checkpoint Saver for SAP HANA Cloud — SAP Community (Artificial Intelligence Blogs Posts)
- New Machine Learning, NLP and AI features in SAP HANA Cloud 2025 Q3 — SAP Community (Technology Blog Posts by SAP)
- Beyond Vectors: The Next Evolution of RAG on SAP BTP using SAP HANA Cloud — SAP Community (Artificial Intelligence Blogs Posts)
- New Machine Learning, NLP and AI features in SAP HANA Cloud 2025 Q4 — SAP Community (Technology Blog Posts by SAP)
- Help future-proof your database landscape with SAP HANA Cloud — SAP Community (Technology Blog Posts by SAP)
- JOIN US: Meet SAP HANA Cloud @ SAP TechEd 2025 — SAP Community (Technology Blog Posts by SAP)
- Step-by-Step: Implementing Generative AI with SAP HANA Cloud — SAP Community (Technology Blog Posts by Members)
- What’s New in SAP HANA Cloud – September 2025 — SAP Community (Technology Blog Posts by SAP)
- Building a RAG Bot on SAP BTP With Hana Cloud Vector Engine and AI Core — SAP Community (Technology Blog Posts by SAP)
- New Machine Learning and AI features in SAP HANA Cloud 2025 Q2 — SAP Community (Technology Blog Posts by SAP)
- Fast-Track Your AI Journey with SAP HANA Cloud’s Vector Engine Through Our New Expert Guided service — SAP Community (Technology Blog Posts by SAP)
- Semantic Querying with SAP HANA Cloud Knowledge Graph using RDF, SPARQL, and Generative AI in Python — SAP Community (Technology Blog Posts by SAP)
- 🚀 SAP AI Core Agent QuickLaunch Series 🚀 - Part 4 RAG Basics ①: HANA Cloud VectorEngine&Embedding — SAP Community (Technology Blog Posts by SAP)
- 🚀 秒速で学ぶ SAP AI Core Agent 開発 🚀 - Part4 RAG 基礎 ①: HANA Cloud VectorEngineと埋め込み処理 — SAP Community (Technology Blog Posts by SAP)
- Making RAG work better: Implementing Parent Document Retriever with SAP HANA Cloud Vector Engine — SAP Community (Technology Blog Posts by SAP)
- New Machine Learning and AI features in SAP HANA Cloud 2025 Q1 — SAP Community (Technology Blog Posts by SAP)
- HANA Cloud’s VECTOR EMBEDDING in CAP and Comparison with OpenAI Embedding — SAP Community (Technology Blog Posts by SAP)
Full card available to members. What the full card adds: the full decision framework · the SAP vs Snowflake / Databricks / Fabric comparison · the common pitfalls and their fix · the cheat sheet · the architecture schemas · the code blocks · the facts worth quoting.