AI & Analytics Legends The knowledge platform for SAP Analytics
Concept card

Mixture-of-Experts (MoE) Architectures — Mixtral / DeepSeek / Grok

Mixture-of-Experts (MoE) Architectures — Mixtral / DeepSeek / Grok — Analytics Legends section illustration for the SAP Analytics knowledge base (concepts, studies, Academy)

As of 2026-10-06

What is Mixture-of-Experts (MoE) Architectures?

Mixtral 8x7B matches dense 70B-parameter models on most benchmarks while running roughly 3x faster because only 2 of its 8 experts (13B of 47B parameters) fire per token — the same architectural trick behind DeepSeek-V3's 671B-total / 37B-active model.

A Mixture-of-Experts (MoE) transformer replaces the dense feed-forward layer inside a transformer block with a bank of N parallel expert networks plus a small gating network, called a router, that selects the top k experts for each token. Only the selected experts actually compute for that token, so a model built with eight experts and top-2 routing fires roughly one quarter of its feed-forward parameters per token. The result is a model that carries the parameter count, and much of the learning capacity, of a very large dense network while paying the inference cost of a much smaller one. MoE is the architectural reason a 47-billion-parameter Mixtral 8x7B model can match dense 70-billion-parameter models on most public benchmarks while running several times faster per token.

Why It Matters

Enterprise AI budgets are dominated by inference cost, not training cost, because a model is trained once and queried millions of times. A dense model pays its full parameter cost on every single token it generates; an MoE model pays only for the experts the router actually activates. For an SAP analytics or Joule-style deployment routing thousands of natural-language queries against enterprise data every day, that difference compounds directly into GPU-hour and API-token spend. MoE is also why open-weight models have been able to close the gap with proprietary frontier models on quality benchmarks without matching them on raw dense parameter count. DeepSeek-V3, at 671 billion total parameters but only 37 billion active per token, was the first open-weight model publicly reported to reach GPT-4-class quality on a broad benchmark suite.

Why it matters in practice

  • Top-k routing means only a fraction of FFN parameters compute per token — 25% for Mixtral's top-2-of-8 — decoupling model quality from inference cost.
  • All 47B Mixtral parameters still have to be loaded into GPU memory even though only 13B activate per token, which is why MoE economics work on multi-GPU servers but break on single-GPU edge deployments.
  • Load balancing is a real production risk: if the router learns to favour one expert, the rest go to waste, requiring auxiliary losses to enforce balance.

Key points

  • MoE replaces dense FFN with N experts + router selecting k per token — Mixtral 8x7B: 47B total / 13B active (top-2 of 8).
  • Frontier examples — Mixtral 8x7B (Apache 2.0, Dec 2023), DeepSeek-V2 (160 experts top-6, May 2024), DeepSeek-V3 (671B/37B active, GPT-4 class, Dec 2024), Grok-1 (314B, Mar 2024).
  • Quality scales with total params (cheap to add experts); cost scales with active params — decouples the two.
  • Memory catch — all params must be GPU-resident; works on multi-GPU servers, breaks on single-GPU edge.
  • Production challenges — load balancing (auxiliary loss), expert parallelism (all-to-all GPU shuffle bottleneck), training instability (discrete routing decisions, z-loss + warmup help).
  • SAP's own tabular foundation models (SAP-RPT-1.6, TabPFN-3.5-Plus, both GA on the generative AI hub) are dense in-context-learning transformers, not MoE — do not extrapolate Mixtral/DeepSeek cost-versus-quality intuition onto SAP-RPT.
  • For a hosted generative AI hub call, AI Units billing is opaque to which experts fired on a given request — the MoE memory-versus-active-parameter trade-off only becomes the customer's problem when self-hosting an open-weight MoE checkpoint on SAP AI Core.

Terms used on this page

Mixture-of-Experts (MoE)
Transformer architecture where the dense FFN layer is replaced by N parallel experts plus a gating router that selects k of N experts per token; only the k chosen experts compute, decoupling total parameters (quality) from active parameters (cost).
Router / gating network
A small learned network (typically a linear layer + softmax) that scores each token against each expert and selects the top-k for actual computation; the discrete top-k selection is the source of MoE's training instability.
Top-k routing
The strategy of choosing the k highest-scoring experts per token; k=2 is canonical (Mixtral, Grok-1), k=6+ used in finer-grained MoEs like DeepSeek-V2.
Load-balancing loss
An auxiliary training objective that penalises the router for sending disproportionate token volume to any one expert; prevents expert collapse where most experts become unused.
Expert parallelism
The distributed-training and inference strategy of placing different experts on different GPUs; requires all-to-all token shuffling between GPUs per MoE layer, creating network bottlenecks.

Sources

  1. Mixtral of Experts (Mistral AI, 2023)
  2. DeepSeek-V2 technical report
  3. DeepSeek-V3 technical report (Dec 2024)
  4. Hugging Face MoE explainer blog
  5. xAI Grok-1 model card
  6. SAP Business AI — official product page
  7. arXiv — Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity (Fedus et al., 2021)
  8. arXiv — GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding (2020)
  9. arXiv — DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models (2024)
  10. SAP News Center — TabPFN-3.5-Plus now available in SAP AI Core (2026-09)
  11. Mistral AI — Mixtral of Experts announcement (2023): 47B total / 13B active parameters
  12. SAP Help Portal — Generative AI hub in SAP AI Core (hosted model endpoints)
  13. SAP Help Portal — Metering and pricing for generative AI on SAP AI Core (AI Units)

Full card available to members. What the full card adds: the full decision framework · the SAP vs Snowflake / Databricks / Fabric comparison · the common pitfalls and their fix · the cheat sheet · the architecture schemas · the code blocks · the facts worth quoting.

Open in the app →