LLM Reasoning Playbook

Worked Stack Examples

Four end-to-end examples of composing frameworks across axes, plus a stack-selection cheatsheet to match your problem to a proven stack.

This page is a pattern library. Find the example closest to your problem and copy its stack, or use the cheatsheet just below to match your task shape to a proven composition. New to composing? Build one from scratch in the guided Build Your Recipe workflow.

Stack-selection cheatsheet

If your task is…Consider the stack
Strategic decision needing live facts + defensibilityStep-Back × ReAct × Self-Consistency
Multi-step quantitative / financial / algorithmicLeast-to-Most × PAL × Self-Refine
Synthesis over many messy documents, high factual stakesThread of Thought × Chain-of-Verification
Autonomous agent that must improve over retriesReAct × Reflexion
Creative strategy exploring many options then convergingTree/Graph of Thoughts × Self-Refine
Build an artifact from real data (charts, reports, code)Step-Back × Plan-and-Solve × Skeleton × [PAL in ReAct] × Self-Refine

The four worked examples. Each expands one archetype in full (the situation, why this mix, the actual prompt, a trimmed trace, and the payoff). Jump to the one closest to your problem:

The one rule. Use at most one framework per axis, and apply them from the outside in: the frameworks that break down or reframe the problem shape the prompt; grounding frameworks (tools, code) run inside the loop; and sampling or checking frameworks wrap the whole thing.


Example 1: Market-Entry Analyst

Stack: Step-Back (E) × ReAct (C) × Self-Consistency (B)

The situation. "Should a mid-size US grocery chain jump into the online meal-kit market in 2026?" This needs three things at once: real, current facts (market size, competitors), a solid way to think it through, and an answer steady enough to defend in a boardroom.

Why this mix.

Assembled prompt (per sample):

xml<task>Recommend GO or NO-GO: mid-size US grocery chain entering the online
meal-kit market in 2026, with justification.</task>
<tools>web_search(query), get_market_report(sector)</tools>

<phase name="step_back">
First state the general framework that governs any market-entry decision
(e.g., market attractiveness × competitive intensity × fit with existing
capability × unit economics). Do not decide yet.
</phase>

<phase name="react">
Loop until you can decide. Each cycle:
Thought: what does the framework tell you to investigate next?
Action: web_search(...) or get_market_report(...)
Observation: <environment-supplied>
Ground every claim in an Observation; never fabricate one.
</phase>

<output>Final Answer: GO or NO-GO + 3-bullet justification tied to the framework.</output>

Wrapper (Self-Consistency, B1):

Run the prompt above N=7 times, temperature=0.7.
Extract GO / NO-GO from each. Return the majority verdict + vote split.

Abbreviated trace (one sample):

step_back → Framework: attractiveness × intensity × capability-fit × unit economics.
Thought: size the market first.
Action: get_market_report("US meal-kit 2026")
Observation: ~$12B, growth slowing to ~5% YoY, top-3 hold ~60% share.
Thought: intensity is high; check capability fit / differentiation.
Action: web_search("grocery chain private-label meal kit cold-chain advantage")
Observation: existing refrigerated logistics = real cost edge vs pure-play.
Thought: unit economics, CAC vs pure-plays?
Action: web_search("meal kit customer acquisition cost 2025")
Observation: CAC high (~$90–130), churn heavy.
Final Answer: NO-GO (standalone). Justification: (1) attractiveness fading, 
slowing growth, entrenched top-3; (2) high CAC + churn breaks unit economics;
(3) capability-fit real but better captured as an add-on to existing e-commerce,
not a standalone entry.

The votes: 5 said NO-GO, 2 said GO → NO-GO (5 of 7). The two GO votes both leaned on the cold-chain edge; showing the split tells the analyst exactly where the decision is shaky.

The payoff. The answer rests on a principle (Step-Back), on real facts (ReAct), and is steadied against luck (Self-Consistency). The 5-to-2 vote split is a bonus deliverable, it points straight at the weak spot in the case.


Example 2: Quantitative Word Problem

Stack: Least-to-Most (D) × PAL (C) × Self-Refine (F)

The situation. "A company's revenue grew 12% in Year 1, fell 5% in Year 2, and grew 20% in Year 3. If Year 3 revenue was $4.2M, what was the starting revenue, and what was the 3-year growth rate (CAGR)?" It's several steps plus exact math, exactly where plain Chain of Thought tends to slip.

Why this mix.

Assembled prompt:

xml<task>Find starting revenue and 3-year CAGR. Y1 +12%, Y2 -5%, Y3 +20%,
Y3 revenue = $4.2M.</task>

<phase name="decompose">
Ordered subproblems, easiest first:
  1. Express Y3 revenue as start × product of growth factors.
  2. Solve for start.
  3. Compute CAGR from start and Y3 over 3 years.
</phase>

<phase name="solve_with_code">
Write Python that implements the subproblems in order. Reasoning in comments,
computation in code. Final line: answer = {"start": ..., "cagr": ...}.
</phase>

<phase name="self_refine">
Given the interpreter result, critique the SETUP (are growth factors and the
CAGR exponent correct?). If flawed, rewrite the code and re-run. Max 2 rounds.
</phase>

Model-written code (PAL):

python# Sub 1: Y3 = start * 1.12 * 0.95 * 1.20
factor = 1.12 * 0.95 * 1.20          # cumulative growth factor
# Sub 2: solve for start
start = 4_200_000 / factor
# Sub 3: CAGR over 3 years = (Y3/start)^(1/3) - 1
cagr = (4_200_000 / start) ** (1/3) - 1
answer = {"start": round(start, 2), "cagr": round(cagr, 4)}

Interpreter result: {"start": 3287671.23, "cagr": 0.0846}

Self-Refine pass:

Critique: factor = 1.2768; start = 4.2M / 1.2768 = $3.29M, consistent.
CAGR check: (4.2/3.288)^(1/3) - 1 = 8.46%. Exponent 1/3 correct for 3 years.
Setup valid → no revision needed.
Final Answer: Starting revenue ≈ $3.29M; 3-year CAGR ≈ 8.46%.

The payoff. Least-to-Most keeps the model from skipping steps; PAL makes the numbers exact; and Self-Refine is safe here because it checks against a real interpreter, not the shaky self-checking the playbook warns about.


Example 3: Multi-Document Synthesis with Fact-Check

Stack: Thread of Thought (A) × Chain-of-Verification (C)

The situation. "Read these 6 uploaded audit reports and support tickets, summarize our vendor's security posture, and flag any contradictions." The input is long, messy, and spread across many sources, easy to lose a fact in the middle , and a made-up claim about a vendor is expensive.

Why this mix.

Assembled prompt:

xml<task>Summarize the vendor's security posture and flag contradictions across the
6 sources.</task>
<context>{{6 REPORTS + TICKETS}}</context>

<phase name="thread">
Walk through each source in turn. For each: summarize its security claims and
note anything that conflicts with earlier sources. Carry a running summary.
</phase>

<phase name="draft_synthesis">Synthesize an overall posture + contradiction list,
grounded only in the per-source notes.</phase>

<phase name="verify">
List verification questions that would falsify each key claim in the synthesis
(e.g., "Does Source 3 actually state SOC2 Type II, or Type I?"). Answer each in
ISOLATION against the sources.
</phase>

<phase name="revise">Correct any claim the verification contradicts. Output final.</phase>

Abbreviated trace:

thread → S1: SOC2 Type II (2024). S2: pen-test, 2 criticals open. S3: claims
"fully remediated", conflicts with S2. S4: MFA enforced. S5: ticket, one
critical still open as of Q1'26. S6: SLA doc, no security detail.
Running thread flags S3-vs-S5 conflict.

draft_synthesis → "Strong (SOC2 II, MFA), but open criticals; S3's 'fully
remediated' contradicts S5's open ticket."

verify → Q: "Does S3 say fully remediated?" → A: S3 says "remediation in
progress," NOT "fully remediated." (draft overstated it)
       Q: "Is a critical open in S5?" → A: yes, confirmed Q1'26.

revise → Corrected: S3 claims remediation *in progress*; S5 confirms ≥1 critical
still open → posture: solid controls, but unresolved criticals and inconsistent
remediation reporting.

The payoff. Thread of Thought surfaces the buried conflict between Source 3 and Source 5 that a one-pass summary would miss; Chain-of-Verification catches the draft's own overstatement ("fully remediated") before it goes out. You get long-input reliability and fact-checking in one pipeline.


Example 4: Support-Team Dashboard (a build task)

Stack: Step-Back (E) × Plan-and-Solve (D) × Skeleton-of-Thought (A) × [PAL inside ReAct] (C) × Self-Refine (F) · skip B · DSP light (G)

The build-task archetype, the same problem you build interactively in Build Your Recipe.

The situation. "Build a comprehensive set of charts in our support team's Dashboard sheet that show how the metrics we track, ticket volume, response time, agent load, CSAT, influence each other." Unlike the first three, this is a build task: it produces an artifact, not a single answer. That changes the recipe in two ways the earlier examples don't show, it skips an axis, and it nests two frameworks on one axis.

Ground it first (Step 0). Before any framework, inspect the real sheet and write the rules a good answer must respect: the sample is small (so links are a maybe, not a fact); several metrics can move together without one causing the other (a product launch spikes volume and sinks CSAT at once); compare a day to the day before, not to itself (use a lag), yesterday's backlog drives today's response times; rank the days instead of trusting exact numbers, and show the dot plot; and the 1–5 CSAT ratings are rankings, not real amounts, so don't do math on them. This list is the part no framework gives you, and it steers every pick below.

Why this mix.

Assembled prompt (sketch):

xml<principles>{{STEP0_PRINCIPLES}}</principles>

<phase name="step_back">State what makes an "influence" valid here (lagged, rank-based,
confound-aware). Do not chart yet.</phase>

<phase name="plan">Design the full chart spec: list every panel (pair, lag, chart type)
before building. (Skeleton: headings first, then expand each panel.)</phase>

<phase name="build_loop">
For each panel, loop (ReAct):
  Thought: which pair + lag to test?
  Action: read_range(...) then run_python(...)   // PAL: compute rank corr + lag in code
  Observation: <environment-supplied numbers>
Never fabricate a correlation; compute it.</phase>

<phase name="self_refine">Critique the panel set: drop spurious/confounded pairs,
flag small-n. Revise. (Verifier = the recomputed stats.)</phase>

<stimulus>Favor leading indicators / early-warning relationships.</stimulus>

Abbreviated trace:

step_back → valid influence = lagged, rank-based, robust to shared events (launches).
plan/skeleton → panels: [backlog→response_time +1d], [response_time→csat +1d],
                [agent_load→response_time same-day], [volume→csat +1d] …
build_loop → Action: run_python(spearman(backlog[:-1], response_time[1:]))
             Observation: rho=0.48, n=63 → keep (lagged, moderate)
             Action: run_python(spearman(volume, csat))  Observation: rho=-0.61 but
             same-row + both spike at the launch → confound suspected
self_refine → drop the same-row volume↔csat panel (confounded); keep the +1d lag
              version; annotate small-n on all.
Final → 5 panels, each lagged rank-correlation + scatter, small-n flagged.

The payoff. Two moves the Q&A examples never make: it skips Sampling (nothing to vote on) and nests PAL inside ReAct on the grounding axis. But the real win is the principles payload from Step 0, it's what turns "charts that correlate" into "charts that don't lie." See the full derivation in Build Your Recipe.


Match your problem to a stack at the top of this page, then build it in Build Your Recipe.