Synthetic Data Generation for Training — Self-Instruct, Self-Play
As of 2026-09-27
What is Synthetic Data Generation for Training?
Synthetic data generation escapes the ceiling of human-labelled data by having a capable teacher model generate and a student model learn — Self-Instruct proved a 175B teacher could bootstrap a competitive 7B student from just 175 seed prompts.
What it is
Synthetic data generation is the practice of using a capable "teacher" model to manufacture training examples — question-answer pairs, code translations, tool-call traces, judged dialogue turns — that then train a "student" model, instead of paying humans to write every example by hand. It became the default enterprise fine-tuning strategy once frontier models got good enough to write consistently correct enterprise-domain data faster than any labeling team could hire and calibrate.
Why it matters
The bottleneck in almost every fine-tuning project is not compute or algorithm — it is the seed corpus. A team that wants a model fluent in translating ABAP into SQL, or answering questions against a specific SAP Business Data Cloud model, typically has a handful of hand-crafted examples and no realistic path to hiring annotators who understand both the domain and the target format well enough to produce thousands more. Synthetic generation turns that handful of seed examples into tens or hundreds of thousands of training pairs at a cost measured in API tokens rather than annotator headcount.
How it works
Three families dominate practice, and they solve different problems.
Self-Instruct prompts a teacher model with a small set of seed examples and asks it to write more examples "in the same style" — same format, same difficulty range, same domain. The output is filtered for near-duplicates and length outliers before it becomes training data. This is the cheapest, broadest approach: it maximizes coverage of instruction variety.
Why it matters in practice
- Three named failure modes recur across every approach: mode collapse, hallucination amplification (teacher errors become student training labels), and verbatim copying that creates downstream copyright exposure.
- The guardrails are specific, not generic: aggressive deduplication (MinHash + SemDeDup), factual filtering against ground truth, and provenance logging for EU AI Act Art. 10 + Art. 53 audit trails.
- For SAP analytics, the canonical pattern is concrete: generate 50K-200K ABAP-to-SQL and BW-to-Datasphere pairs from a hand-crafted seed set, filter hard, then LoRA-fine-tune (C258) on the result.
Key points
- Three families — Self-Instruct (teacher generates instructions from seeds), Self-Play (model plays both roles), Distillation (mimic a stronger teacher's outputs).
- Failure modes — mode collapse, hallucination amplification, verbatim copying from teacher training data; mitigated by aggressive deduplication and factual filtering.
- Guardrails — MinHash + SemDeDup deduplication, factual filtering against ground-truth KB, EU AI Act Art. 10 + Art. 53 provenance logging on every sample.
- SAP pattern — frontier API generates 50K-200K ABAP-SQL / BW-Datasphere / Q&A pairs from hand-crafted seeds; filter aggressively; LoRA-fine-tune (C258).
- Economics — 50-200x cheaper than human labelling at 2026 frontier-API rates; quality gap closes every model release.
- SAP's generative AI hub model library reaches Claude, OpenAI GPT models, Gemini and Mistral through one governed orchestration service — route synthetic-data-generation calls through it so data masking strips real customer content from seed prompts before they reach an external model.
- Log provenance inside SAP AI Core's own audit trail, not a side spreadsheet — it is the artefact an EU AI Act Art. 10/53 governance review will actually trust.
- From 2026-10-01, ABAP AI moves from SAP's free promotional tier to commercial, AI-Units-metered access — price a 50K-200K-sample synthetic-data pipeline against that date, not a free-tier proof-of-concept assumption.
Terms used on this page
- Self-Instruct
- Method (Wang et al. 2023) where a teacher LLM bootstraps an instruction dataset from a small seed set by generating new instruction-response pairs in the same style; the 7B Alpaca was the first widely-recognised demonstration.
- Self-Play (SPIN)
- Training procedure where the same model generates and then evaluates its own outputs, with a reward signal or judge selecting preferred outputs for the next training round; reduces dependence on external labels.
- Distillation
- Training a smaller student model to mimic the outputs of a larger teacher (Hinton et al., 2015); shifts inference cost from serving time to training time and produces compact models with much of the teacher's capability.
- Mode collapse
- Failure mode of synthetic data generation where the teacher emits a narrow band of patterns repeatedly; the student then over-specialises to that band and fails on out-of-distribution inputs.
- Data masking (orchestration service)
- A configurable pipeline module in SAP's generative AI hub orchestration service that strips or pseudonymises sensitive fields before a prompt reaches an LLM endpoint — the relevant control when seed examples for synthetic-data generation are drawn from real customer ABAP code, tickets or configuration data.
- Provenance logging
- Recording which model generated a sample, from what prompt, and when, for every synthetic training example; required under EU AI Act Art. 10 (data governance) and Art. 53 (GPAI transparency), and the record a governance reviewer will actually ask to see.
- Golden set
- A small, human-labelled, held-out evaluation set used to check a synthetic-data-trained model's real quality; never trust a synthetic-only evaluation, since it can share the same blind spots as the synthetic training data.
Sources
- arXiv: Self-Instruct — Aligning Language Models with Self-Generated Instructions (Wang et al. 2023)
- arXiv: Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models (Chen et al. 2024 — SPIN)
- Anthropic Responsible Scaling Policy — training-data governance
- NIST AI Risk Management Framework 1.0
- SAP Business AI — official product page
- arXiv — Bai et al. (Anthropic), "Constitutional AI: Harmlessness from AI Feedback" (2022, fetched 2026-09-27)
- arXiv — Hinton, Vinyals & Dean, "Distilling the Knowledge in a Neural Network" (2015, fetched 2026-09-27)
- Stanford CRFM — Alpaca: A Strong, Replicable Instruction-Following Model (2023-03-13, fetched 2026-09-27)
- SAP Help Portal — Orchestration service (data masking, content filtering) in the generative AI hub (fetched 2026-09-27)
- SAP Help Portal — generative AI hub model access overview, SAP AI Core (fetched 2026-09-27)
- SAP Community — SAP Joule for Developers: end of the promotional period and the path forward (2026, fetched 2026-09-27)
- EUR-Lex — Regulation (EU) 2024/1689 (EU AI Act), consolidated text incl. Art. 10 and Art. 53 (fetched 2026-09-27)
Full card available to members. What the full card adds: the full decision framework · the SAP vs Snowflake / Databricks / Fabric comparison · the common pitfalls and their fix · the cheat sheet · the architecture schemas · the code blocks · the facts worth quoting.