AI & Analytics Legends The knowledge platform for SAP Analytics
Concept card

Model Distillation — Train Small Models From Large Teachers

Model Distillation — Train Small Models From Large Teachers — Analytics Legends section illustration for the SAP Analytics knowledge base (concepts, studies, Academy)

As of 2026-09-27

What is Model Distillation?

Black-box distillation — running a closed-API teacher on your domain prompts, then fine-tuning an open-weight student on the results — is the only mode that works with frontier vendors, needing no access to their internals.

What it is

Model distillation is a training technique that transfers the learned behaviour of a large, expensive teacher model into a smaller, faster student model. The goal is not to reproduce the teacher's weights — that is impossible without access to them — but to reproduce its output behaviour on the domain of tasks that matter to you. The student trains on examples the teacher has already processed, learning to imitate the teacher's responses while remaining far cheaper to run at inference time.

Why it matters

  • Logit distillation gives richer training signal via KL divergence over the teacher's full probability distribution, but requires the teacher to expose logits — possible with Llama, not with GPT-4 or Claude.
  • Hidden-state distillation is the most data-efficient mode but needs architectural similarity between teacher and student, making it impractical when distilling from a proprietary architecture to an open-weight one.
  • The enterprise case for distilling over pure prompting is scale economics: every frontier-API call keeps paying the provider, while a distilled student runs far cheaper at inference.

Key points

  • Train a small student (7B-8B) to imitate a large teacher (70B+) on task-specific prompts — 80-95% quality at 5-10% inference cost.
  • Three modes — black-box (text outputs only, works with API teachers), logit (full distribution, open-weight teacher), hidden-state (deepest signal, architectural similarity).
  • Black-box distillation against GPT-4/Claude is the operative enterprise pattern — no GPU training infra beyond standard fine-tune.
  • Wins on narrow well-defined tasks (extraction, classification, structured generation); fails on open-ended reasoning + world-knowledge tasks.
  • Synthetic-data generation is distillation by another name — DeepSeek-R1-distill (2025), Phi-4 (Microsoft 2024) are public examples.
  • OpenAI's Model Distillation feature operationalises this workflow end-to-end inside one API: `store: true` on Chat Completions captures teacher input-output pairs as stored completions, then used directly as a fine-tuning dataset for a smaller model.
  • Check the teacher's licence before building the pipeline: Meta's Llama Community License and most closed-API terms of service restrict using a model's outputs to train a directly competing model — a legal review belongs before the first distillation run, not after deployment.
  • Databricks Mosaic AI Training and Snowflake Cortex Fine-tuning both expose managed fine-tuning jobs that can consume a teacher-generated dataset directly, so an enterprise already on either platform does not need separate distillation infrastructure.

Terms used on this page

Distillation
Training a small student model to imitate a large teacher model on a task-specific dataset, transferring the teacher's behaviour at a fraction of the inference cost.
Teacher
The large, expensive model whose outputs (or output distributions) form the training signal for the student; typically 70B+ parameters or a frontier API.
Student
The small model being trained; typically 7B-8B parameters; runs at 10-100x lower cost than the teacher after distillation.
Synthetic data generation
Using a large model to generate training examples (prompt-response pairs) that are then used to fine-tune a smaller model; modern enterprise face of distillation.
Chain-of-thought distillation
Training a student on the teacher's full reasoning trace rather than only its final answer — the technique behind DeepSeek-R1-distill and similar reasoning-model students; preserves more of the teacher's step-by-step problem-solving behaviour than output-only distillation.
Stored completions
OpenAI's mechanism for capturing input-output pairs from live API traffic (via a `store: true` flag on Chat Completions) to build a distillation or evaluation dataset without a separate logging pipeline.

Sources

  1. Hinton, Vinyals, Dean — Distilling the Knowledge in a Neural Network (the founding paper, 2015)
  2. DeepSeek-AI — DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (R1-distill family, 2025)
  3. Microsoft — Phi-4 Technical Report (2024)
  4. HuggingFace — TRL (Transformer Reinforcement Learning) distillation recipes
  5. SAP Help Portal — SAP Autonomous Suite documentation
  6. OpenAI — Model Distillation in the API (product announcement)
  7. OpenAI Cookbook — Leveraging model distillation to fine-tune a model
  8. Microsoft Learn — Stored completions and distillation in Azure OpenAI / Foundry
  9. Meta — Llama 4 Community License Agreement
  10. Databricks — Introducing Mosaic AI Model Training for fine-tuning GenAI models
  11. Snowflake Docs — Fine-tuning (Snowflake Cortex)
  12. Snowflake — Cortex Fine-tuning general availability release note (2025-02-07)
  13. Databricks Docs — Mosaic AI Training finetuning reference

Full card available to members. What the full card adds: the full decision framework · the SAP vs Snowflake / Databricks / Fabric comparison · the common pitfalls and their fix · the cheat sheet · the architecture schemas · the code blocks · the facts worth quoting.

Open in the app →