Reasoning Models — Claude Extended Thinking, GPT-o4, Gemini Thinking
As of 2026-10-05
What is Reasoning Models?
Reasoning models spend 10-100× more inference compute per hard query via an internal chain-of-thought, and the right architecture routes only the 5-15% of genuinely hard agent invocations to them, not every call.
What it is
Reasoning models are the 2026 architectural inflection that the Stanford HAI AI Index 2026 documents as the technical underpinning of the agentic-AI category. They differ from baseline LLMs in one crucial axis: at inference time they spend additional compute on an internal chain-of-thought before producing the final answer — the model's 'thinking' is itself a computational step, not just a prompt-engineering trick. The three vendor families with production-grade reasoning modes in 2026 are Anthropic's Claude Extended Thinking, OpenAI's current reasoning-tier models (branded across a GPT-5.x/GPT-6 reasoning lineup as of September 2026 — OpenAI's own o-series naming has already been superseded once, so verify the current model name on developers.openai.com before quoting one to a client), and Google's Gemini models running with a configurable thinking_level parameter (Gemini 3.x and 2.5-series, per ai.google.dev, checked 2026-09-27).
The mechanism is variable inference-time compute. Where a baseline LLM produces a token stream in roughly constant time per token, a reasoning model can spend 10× to 100× more inference compute on a hard problem before emitting the final answer, by generating an internal reasoning trace (visible or hidden depending on the vendor's policy) that the model uses to refine its conclusion. The benchmark signature is unmistakable: on GPQA Diamond, reasoning models reach 70-85% vs 40-50% baseline; on competition mathematics (AIME, IMO subsets), reasoning models cross thresholds baseline models cannot reach at any prompt length.
Why it matters
- The benchmark signature is unmistakable: GPQA Diamond jumps from 40-50% baseline to 70-85% with reasoning enabled, and reasoning models cross competition-math thresholds baseline models never reach.
- The cost is real — 30-120 seconds per hard answer versus 2-5 seconds for a baseline LLM, and 5-20× higher token cost per query.
- Each vendor ships a distinct variant into a distinct product: Claude Extended Thinking (visible traces, SAP's autonomous supply-chain agents), GPT-o4 (hidden by default, Microsoft Copilot Studio), Gemini Thinking (visible, Databricks Mosaic AI).
Key points
- Reasoning models spend variable inference-time compute on internal chain-of-thought before answering — the 2026 architectural inflection per Stanford HAI AI Index 2026.
- Benchmark signature on GPQA Diamond — 70-85% reasoning vs 40-50% baseline (up from a 2023 GPT-4 baseline of 39% in the original GPQA paper); the gap is the canonical 'reasoning mode value' metric.
- Three vendor families — Claude Extended Thinking (Anthropic, visible traces, used by SAP for supply-chain agents) · OpenAI's current reasoning tier (hidden by default, opt-in summaries, naming evolves — verify on developers.openai.com before quoting) · Gemini with a configurable thinking_level parameter (Google, Vertex AI, Databricks Mosaic AI).
- Trade-off — 30-120s latency vs 2-5s baseline · 5-20× per-query cost, billed as generative-AI-hub token consumption, NOT covered by SAP's flat 0.02-AI-Units-per-action figure for SAP-delivered agents.
- Trace visibility is a governance axis, not just a UX detail — Claude and Gemini expose a visible/summarised trace, OpenAI's tier keeps it hidden by default; this matters for EU AI Act Art. 6/Annex III technical-documentation obligations (2027-12-02).
- Hybrid routing pattern — Joule Studio 2.0 (June 2026) routes automatically between reasoning and baseline; the architecture decision a consultant defends with the customer's CFO and, separately, with their compliance team.
- Only the 5-15% of genuinely hard agent invocations justify reasoning mode; a cheap complexity classifier in front of the routing decision, instrumented and reviewed against human-verdict outcomes, is the practical implementation pattern.
Terms used on this page
- Extended thinking
- Anthropic's term for the variable-inference-time-compute mode of Claude where the model generates a visible chain-of-thought trace before producing the final answer; the thinking-budget parameter is controllable by the calling agent.
- OpenAI reasoning tier
- OpenAI's current reasoning-model tier (per developers.openai.com, checked 2026-09-27); reasoning tokens are hidden by default and billed as output tokens, with an opt-in `summary` parameter and reasoning-effort levels up to 'xhigh'/'max'. Naming has already changed once since the o-series; verify the current model name before quoting one in a proposal.
- Gemini thinking_level
- Google's configurable reasoning-effort parameter (low/medium/high) available on Gemini 3.x and 2.5-series models (per ai.google.dev, checked 2026-09-27); exposes a partially visible reasoning trace via `thought` steps with an optional summary; pricing combines output and thinking tokens.
- Hybrid routing
- The architectural pattern where an agent platform routes each invocation between a reasoning model (5-15% of hard questions) and a baseline LLM (85% routine) to balance accuracy vs latency vs cost; Joule Studio 2.0 (June 2026) implements this natively.
- Reasoning tokens
- The additional tokens a reasoning model generates internally before producing its visible answer; billed as output tokens across all three vendor families, which is the mechanical reason reasoning mode costs materially more per query than a baseline call.
- Trace-visibility governance
- The distinction between a reasoning model's chain-of-thought being visible/summarised by default (Claude, Gemini) versus hidden by default with only an opt-in summary (OpenAI's current tier); relevant to whether a reasoning trace can serve as technical documentation evidence under EU AI Act Art. 6/Annex III obligations.
Sources
- Stanford HAI AI Index Report 2026 §1 Technical performance — reasoning-model inflection
- Anthropic — Claude Extended Thinking documentation
- Analytics Legends news — SAP autonomous supply chain agents (Claude embedded) Sapphire 2026
- SAP Business AI — official product page
- SAP Generative AI — official product page
- OpenAI Developer Docs — Reasoning models guide, current model lineup and token mechanics (checked 2026-09-27)
- Google AI for Developers — Gemini thinking mode documentation, thinking_level parameter (checked 2026-09-27)
- arXiv — GPQA: A Graduate-Level Google-Proof Q&A Benchmark, original 2023 baseline scores (2311.12022)
- SAP — AI Units pricing model, 0.02 AI Units per autonomous-agent action
- Gibson Dunn — EU AI Act Digital Omnibus, revised high-risk enforcement dates
- SAP Help Portal — generative AI hub metering and pricing for generative AI
- SAP News Center — SAP AI Agent Hub, cross-vendor agent/LLM governance (2026-09)
Full card available to members. What the full card adds: the full decision framework · the SAP vs Snowflake / Databricks / Fabric comparison · the common pitfalls and their fix · the cheat sheet · the architecture schemas · the code blocks · the facts worth quoting.