AI Red-teaming + Safety Evaluation — Adversarial Prompts, Jailbreaks, Harm Eval
As of 2026-10-06
What is AI Red-teaming + Safety Evaluation?
Red-teaming is now a compliance requirement, not a research nicety — EU AI Act Art. 9/15 documentation for a high-risk system runs 200-800 person-days, and the minimum scope for SAP deployments spans five specific checks.
What it is
AI red-teaming is the structured adversarial probing of a model and its deployed system to surface unsafe, biased, leak-prone, or jailbreak-susceptible behaviours BEFORE production exposure. In 2026 it is no longer a research curiosity — under EU AI Act Art. 9 (risk management) and Art. 15 (accuracy, robustness, cybersecurity) high-risk system providers must demonstrate red-team coverage; the Art. 9-49 documentation effort runs 200-800 person-days per high-risk system per the Munich applied-AI study cited in the Analytics Legends research ledger.
A red-team campaign has four work streams. Capability probing surfaces what the model can do — including capabilities the provider did not intend (writing exploit code, synthesising biothreat instructions, producing CSAM). Prompt-injection and jailbreak probing tests whether system prompts, retrieved-context boundaries, and tool-use guards can be bypassed by hostile user input. Bias and harm evaluation runs scenario suites (BBQ, RealToxicityPrompts, HarmBench, DecodingTrust) and surfaces disparate performance across demographic axes. Misuse evaluation probes specific high-risk uses (election manipulation, fraud, targeted harassment) and documents refusal-rate plus residual leak rate.
The canonical methodology is now codified. Anthropic's Responsible Scaling Policy commits to red-team gates at defined capability thresholds; OpenAI's preparedness framework and DeepMind's Frontier Safety Framework follow similar structure. NIST AI RMF Govern-Map-Measure-Manage functions map directly onto the red-team workflow. The MITRE ATLAS knowledge base catalogues adversarial tactics and techniques. The OWASP LLM Top 10 enumerates the production-deployment risks.
Why it matters
- Cost is quantified, not abstract: the Munich applied-AI study cited puts Art. 9-49 documentation effort at 200-800 person-days per high-risk system.
- The methodology is already codified across labs (Anthropic's RSP, OpenAI's preparedness framework, DeepMind's Frontier Safety Framework) and standards bodies (NIST AI RMF, MITRE ATLAS, OWASP LLM Top 10) — this isn't ad hoc.
- For SAP specifically, the minimum red-team scope names five concrete checks: prompt-injection over retrieved-content boundaries, Joule tool-use data leaks across tenants, refusal-rate on HR/payroll queries, bias in candidate-screening outputs, and OWASP LLM-01 jailbreak resistance.
Key points
- Four work streams — capability probing, prompt-injection / jailbreak, bias and harm evaluation, misuse evaluation; each documented in the Art. 9 risk register.
- EU AI Act obligations — Art. 9 (risk management) + Art. 15 (accuracy, robustness, cybersecurity) require demonstrated red-team coverage for high-risk systems by 2027-12-02 (Digital Omnibus, in force since 2026-07-27) (Annex III) / 2028-08-02 (product-embedded). Per the Analytics Legends research ledger.
- Frameworks — Anthropic Responsible Scaling Policy, OpenAI Preparedness, DeepMind Frontier Safety Framework, NIST AI RMF 1.0, MITRE ATLAS, OWASP LLM Top 10.
- Eval suites — BBQ (bias), RealToxicityPrompts, HarmBench, DecodingTrust, ToxiGen; supplement with domain-specific scenarios.
- SAP-specific scope — prompt-injection over retrieved content, Joule tool-use tenant-crossing, sensitive-data refusal rate, candidate-screening bias, OWASP LLM-01 jailbreak resistance.
- Tabular foundation models (SAP-RPT-1.6, TabPFN-3.5-Plus, both current in Sep 2026) need demographic-parity/disparate-impact statistical testing, not text-based bias suites, when used to score or rank people.
Terms used on this page
- Red-teaming
- Structured adversarial probing of an AI system to surface unsafe, biased, leak-prone or jailbreak-susceptible behaviours before production deployment; codified under EU AI Act Art. 9 risk-management for high-risk systems.
- Jailbreak
- Adversarial prompt pattern that bypasses a model's safety training or system-prompt guardrails, causing it to produce content it would normally refuse; catalogued by OWASP LLM-01 and tracked in HarmBench.
- Prompt injection
- Attack where hostile content embedded in a tool result, retrieved document, or user input rewrites the model's effective instructions; OWASP LLM Top 10 #1 risk in production deployments.
- Responsible Scaling Policy (RSP)
- Anthropic's voluntary commitment to define AI Safety Level (ASL) thresholds and to require specific red-team gates, deployment controls, and security measures before crossing each threshold; emulated by OpenAI's Preparedness and DeepMind's Frontier Safety frameworks.
- Demographic parity / disparate impact
- Statistical fairness tests comparing a model's positive-outcome rate across protected groups; the appropriate bias-evaluation method for a tabular scoring model, distinct from text-generation bias suites like BBQ.
- NIST AI RMF (Govern-Map-Measure-Manage)
- NIST's four-function structure for AI risk management; an inventory/classification tool such as SAP's AI Agent Hub covers Govern/Map, while Measure requires actual red-team or evaluation evidence.
- MITRE ATLAS
- A knowledge base of adversarial tactics and techniques against AI systems, structured on the same model as the MITRE ATT&CK framework for traditional cybersecurity.
Sources
- Anthropic Responsible Scaling Policy
- EU AI Act — Regulation (EU) 2024/1689, Art. 9 (risk management) + Art. 15 (accuracy, robustness, cybersecurity)
- NIST AI Risk Management Framework 1.0
- OWASP Top 10 for Large Language Model Applications
- SAP Community — SAP-RPT-1.6, Tabular Orchestration and RPT Playground API now available (Sep 2026)
- SAP News Center — TabPFN-3.5-Plus now available in SAP AI Core (15 Sep 2026)
- Prior Labs — TabPFN 3.5 changelog
- MLflow — LLM/model evaluation documentation (custom fairness metrics pattern)
- Microsoft — Responsible AI Standard
- Microsoft / Azure — PyRIT, Python Risk Identification Tool for generative AI (GitHub)
- SAP News Center — AI agents work at scale: AI Governance Assistant, EU AI Act + NIST classification (22 Sep 2026)
- NIST — AI 600-1 Generative AI Profile (risk categories and red-teaming actions for generative AI)
- EU AI Act — Article 55, obligations for providers of general-purpose AI models with systemic risk (documented adversarial testing)
- Mazeika et al. — HarmBench: a standardized evaluation framework for automated red teaming and robust refusal (arXiv 2402.04249)
- NVIDIA — garak, LLM vulnerability scanner (probe catalogue for jailbreaks, prompt injection, leakage)
Full card available to members. What the full card adds: the full decision framework · the SAP vs Snowflake / Databricks / Fabric comparison · the common pitfalls and their fix · the cheat sheet · the architecture schemas · the code blocks · the facts worth quoting.