Audio and Speech Foundation Models in SAP Workflows
As of 2026-09-27
What is Audio and Speech Foundation Models in SAP Workflows?
Whisper-based transcription cuts SAP Concur expense-note submission time from 4.2 to 1.1 minutes per line item by turning voice notes that were previously invisible to SAP analytics into structured fields.
What it is
Audio and speech foundation models convert spoken language into structured text—automatic speech recognition, or ASR—and increasingly perform semantic tasks directly on the audio signal itself, such as speaker diarisation, emotion detection, and language identification. The model that defines this category is Whisper, released by OpenAI in 2022: a transformer sequence-to-sequence model trained on 680,000 hours of multilingual audio, achieving word-error rates competitive with commercial ASR APIs at essentially zero marginal cost once self-hosted. The SAP pain point Whisper addresses is structural: audio generated by field technicians, call-centre agents, expense-report narration, and meeting recordings is invisible to SAP analytics for one simple reason—it lives in MP3 or WAV files, not in table fields, and nothing in a standard S/4HANA landscape can query an audio waveform.
Why it matters
- Whisper is trained on 680,000 hours of multilingual audio and matches commercial ASR APIs at zero marginal cost once self-hosted.
- Call-centre QA diarises and scores every call against a compliance checklist, feeding structured scores into a CRM case field and an SAC dashboard.
- Field-service voice notes get transcribed and entity-extracted to pre-fill S/4HANA PM notifications, removing manual data entry for technicians.
Key points
- Whisper large-v3: trained on 680k hours multilingual audio, open-source, runs 50× realtime on A10G GPU.
- No native diarisation — chain pyannote.audio for speaker separation in call-centre QA use cases.
- SAP Concur expense narration: 4.2 min → 1.1 min per line item in enterprise pilots.
- Streaming latency: 300-800 ms additional buffering vs batch — avoid for real-time IVR.
- A 'meeting intelligence' pattern (transcription + summarisation + structured write-back to SAP collaboration records) is architecturally straightforward to build on Whisper plus the generative AI hub's orchestration service, but a named, GA'd SAP product for it was not independently verified at a primary SAP source during this review — scope it as a custom build until a named product page is confirmed.
- Named-entity extraction downstream of transcription — not the transcription itself — is where the SAP business value sits in every integration pattern in this card; a raw transcript changes nothing in an S/4HANA or Concur process on its own.
- Code-switching and regional accent variation are documented Whisper weak points — validate word-error rate on representative audio from the actual global user base before committing to an accuracy SLA.
Terms used on this page
- Whisper
- OpenAI 2022 sequence-to-sequence ASR model. Trained on 680,000 hours of multilingual audio. Available as open-source weights (small/medium/large-v3). No native diarisation.
- WER
- Word Error Rate — the standard ASR accuracy metric. WER = (substitutions + deletions + insertions) / reference word count.
- Diarisation
- Separating audio by speaker identity — 'who spoke when'. Whisper does not include this natively; pyannote.audio is the standard companion library.
- SAP Field Service Management (FSM)
- SAP's mobile field-service application — the verified, named entry point for the voice-note-to-PM-notification integration pattern this card describes, unlike some 'meeting intelligence' wrapper products that were not independently confirmed at a primary SAP source during this review.
- Automatic Speech Recognition (ASR)
- The general category of models converting spoken audio into text; Whisper is the current open-weight reference model, alongside commercial APIs (Google Speech-to-Text, Azure AI Speech).
- Named-entity extraction (post-ASR)
- The downstream step that turns a raw transcript into structured fields (expense type, fault code, compliance flag) — this card's pitfalls section identifies this step, not transcription accuracy, as where SAP business value actually sits.
- Streaming vs batch ASR
- Streaming ASR trades some accuracy for sub-second latency (live captioning, real-time IVR); batch ASR processes complete audio files for higher accuracy at higher latency. Whisper's architecture favours the batch case.
- Code-switching
- Mid-conversation switching between languages, common among multilingual field teams and global call centres — a documented accuracy weak point for general-purpose ASR models including Whisper.
Sources
- Whisper paper — Radford et al. OpenAI 2022
- SAP Concur Expense — SAP Help Portal
- SAP Field Service Management — SAP Help Portal
- pyannote.audio speaker diarisation library
- SAP Generative AI — official product page
- Google Cloud — Speech-to-Text documentation
- Microsoft Learn — Azure AI Speech overview
- OpenAI — Whisper GitHub repository
- SAP Help Portal — Generative AI Hub overview, SAP AI Core (2026)
- Bredin et al. — pyannote.audio: neural building blocks for speaker diarization, arXiv:2104.04045 (2021)
- SAP Help Portal — Orchestration service, content filtering and data masking (2026)
- NVIDIA — A10 Tensor Core GPU specifications (inference hardware reference)
Full card available to members. What the full card adds: the full decision framework · the SAP vs Snowflake / Databricks / Fabric comparison · the common pitfalls and their fix · the cheat sheet · the architecture schemas · the code blocks · the facts worth quoting.