Instruction Tuning (SFT) — From Pre-trained to Helpful
As of 2026-10-09
What is Instruction Tuning (SFT)?
SFT turns a raw pretrained LLM's habit of continuing text into instruction-following by fine-tuning on curated (instruction, response) pairs — but it teaches behavior only, never new knowledge.
Instruction tuning — usually called Supervised Fine-Tuning, or SFT — is the training step that turns a raw, pre-trained language model into something that behaves like an assistant. A pre-trained model has only ever learned to predict the next word in enormous amounts of text; it has no built-in notion that a question should be followed by an answer rather than by more questions. Ask a purely pre-trained model "What is the capital of France?" and it may just as plausibly continue with "What is the capital of Germany?" because, statistically, that is what often follows a question in its training corpus. SFT is what teaches the model that a question deserves an answer, an instruction deserves compliance, and a request deserves a helpful, well-formatted response.
How it works
The mechanism is simple relative to what it produces. A curated dataset of instruction-response pairs — anywhere from ten thousand to a few million examples — is used to continue training the pre-trained model, using the same next-token-prediction loss as pre-training, just on a small, carefully written dataset instead of a scrape of the open internet. The model's weights shift toward producing the demonstrated style of response given a demonstrated style of instruction. Crucially, SFT does not add new world knowledge — that lives in the enormous pre-training corpus. SFT teaches behavior: format compliance, tone, refusal patterns, instruction-following.
Why it matters
- Unprompted, a pretrained model answers "What is the capital of France?" with more questions, not an answer — SFT is what fixes that specific failure.
- InstructGPT's 1.3B SFT model beat raw 175B GPT-3 in human ratings — proof that alignment work outweighs raw scale.
- SFT can only imitate demonstrations, not rank quality — it inherits every confidently-wrong example in its training set, which is exactly the gap RLHF, Constitutional AI, and DPO close.
Key points
- SFT teaches behaviour (instruction-following, format, tone, refusals), NOT knowledge — knowledge comes from pre-training.
- Dataset shape — 10K to 1M (instruction, response) pairs; quality dominates quantity past ~50K examples (LIMA 2023).
- Foundational works — FLAN (Google 2021) multi-task; InstructGPT (OpenAI 2022) free-form human-written; Self-Instruct / Alpaca teacher-bootstrapped.
- Key result — 1.3B SFT-tuned model preferred over 175B raw GPT-3 by humans (InstructGPT 2022) — alignment > scale on usefulness.
- Limits — no notion of quality ranking, can't learn from negatives, propagates demonstration errors; needs RLHF/DPO/CAI on top.
- SAP's generative-AI-hub guidance favors prompting plus grounding as the default before fine-tuning — most client requests (fresh facts, current numbers) are grounding problems, not the stable-format problems SFT is built to fix.
- SAP Domain Models (SAP-ABAP-2, S/4HANA/Ariba variants) are SAP's own continued-training / SFT-adjacent product on top of an existing foundation model — Early Adopter Care today, GA targeted Q3 2026 (past due: no GA announcement found as of 9 October 2026), not a committed delivery date yet.
Terms used on this page
- Supervised Fine-Tuning (SFT)
- The instruction-tuning phase that uses standard supervised next-token prediction on a curated dataset of (instruction, response) pairs to teach a pre-trained LLM to follow instructions and behave like an assistant; the foundation phase before any preference optimisation.
- FLAN
- Google's 2021 work (Finetuned Language Net) that demonstrated multi-task instruction tuning over 60+ NLP datasets formatted as instructions produces strong zero-shot generalisation to unseen tasks.
- InstructGPT
- OpenAI's 2022 paper (Ouyang et al.) that combined SFT on human-written free-form demonstrations with RLHF; the canonical recipe behind ChatGPT and the proof that a 1.3B aligned model beats a 175B raw model on usefulness.
- LIMA
- Meta 2023 result (Less Is More for Alignment) showing 1,000 carefully curated SFT examples produce a model competitive with 50K+ examples — quality dominates quantity in instruction data.
- Continued pre-training
- A further training pass on a domain-specific corpus layered onto an already pre-trained foundation model, distinct from SFT's (instruction, response) pair format; SAP Domain Models use this approach on SAP's own ABAP/S/4HANA/Ariba corpus rather than a from-scratch pre-training run.
- Prompting-plus-grounding default
- SAP's generative-AI-hub design preference for solving a capability gap via orchestration-service prompting and retrieval (document/database grounding) before reaching for fine-tuning; reflects that most client-reported gaps are fresh-facts problems, which grounding solves more cheaply than SFT.
Sources
- Ouyang et al. — Training language models to follow instructions with human feedback (InstructGPT, OpenAI 2022)
- Wei et al. — Finetuned Language Models Are Zero-Shot Learners (FLAN, Google 2021)
- Zhou et al. — LIMA: Less Is More for Alignment (Meta 2023)
- Stanford Alpaca — instruction-tuned LLaMA via Self-Instruct
- SAP Generative AI — official product page
- SAP Help Portal — SAP AI Core generative AI hub, model access and orchestration
- SAP Help Portal — generative AI hub orchestration service (prompting-plus-grounding default before fine-tuning)
- SAP Community — SAP Joule for Developers, end of the promotional period and AI-Units metering from 2026-10-01
- SAP Help Portal — metering and pricing for generative AI on SAP AI Core
- SAP News Center — Joule Studio, enterprise-scale agentic development (SAP Domain Models context)
- Self-Instruct: Aligning Language Models with Self-Generated Instructions — Wang et al., arXiv
- Scaling Instruction-Finetuned Language Models (Flan) — Chung et al., arXiv
- Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks — Wang et al., arXiv
- LoRA: Low-Rank Adaptation of Large Language Models — Hu et al., arXiv
- QLoRA: Efficient Finetuning of Quantized LLMs — Dettmers et al., arXiv
- SFT Trainer — Hugging Face TRL documentation
Full card available to members. What the full card adds: the full decision framework · the SAP vs Snowflake / Databricks / Fabric comparison · the common pitfalls and their fix · the cheat sheet · the architecture schemas · the code blocks · the facts worth quoting.