Methodology walkthrough 9 min read

How to Read Scenario Intelligence: Component Deltas, Carried Assumptions, and Issue Transitions

How to interpret a DurationX comparison scenario beyond its total score: which components moved, what stayed the same, and why a zero-change result is not a blank result.

DurationXPublished July 23, 2026Worked with real, current sample figures

Published by DurationX (legal entity: Dijin Teknoloji Limited Şirketi, Ankara, Türkiye), the vendor-independent, human-reviewed pre-CAPEX assumption-readiness audit for long-duration energy storage. DurationX is run today as a single founder-operated pilot, so this article — like every DurationX report and public page — is attributed to the operating entity rather than to an individual. See Security & Data Handling for the full operator profile, human-review scope, and vendor-independence statement.

A comparison scenario that shows only “readiness 75, down 7 from baseline” asks you to trust a number. Scenario Intelligence is the deterministic bundle that lets you check it instead.

Every DurationX report compares the customer’s own baseline against up to four alternatives — a duration alternative, a technology-class alternative, or both. Each comparison scenario carries the same structured bundle, built by worker/src/pipeline.ts from the same deterministic scoring and benchmark-evaluation code the baseline itself uses. This guide walks through every part of that bundle using the real figures from the current worked sample.

Executive summary

The short version

  • Every comparison scenario states which assumptions changed and which were explicitly carried from the baseline — never an unlabeled second project.
  • A scenario’s component-level deltas always reconcile exactly to its total readiness delta, including the small rounding residual that bridges them — nothing is hidden in an unexplained gap.
  • Confidence is checked separately from readiness and usually does not move — when it does, the scenario states the exact evidence-coverage fact that changed.
  • Every finding transitions in one of four ways — resolved, introduced, unchanged, or no longer evaluable — and "resolved" does not always mean the score improved.
  • A scenario whose total score does not change still explains why, in one of four fixed ways — never an unexplained duplicate of the baseline.
Worked example used throughout this guide. The current worked decision brief is a fictional 70 MW / 40-hour compressed-air storage case (Meridian LDES 40h (sample), readiness 82, confidence High/83). It compares two alternatives: a 28-hour duration alternative, and a hydrogen technology-class alternative. No real customer, company, or project appears in it.

Changed vs. carried assumptions

Every scenario states, in units, exactly which field changed — for example target_duration_hours: 40 → 28 — and separately lists which decision-relevant assumptions were explicitly carried forward unchanged: technology class, CAPEX low/base/high, use case, reliability requirement, and grid/interconnection status (whichever of these was not itself the field varied by this scenario). A scenario is never an unlabeled second project — you always know exactly what is being held fixed.

The reconciliation identity

Every one of the eight scoring components has a before score, an after score, a delta, and that delta’s own weighted contribution to readiness (delta × weight ÷ 100). These weighted contributions are required to reconcile — exactly, not approximately — to the scenario’s total readiness change, through a fixed, disclosed chain:

Sum of weighted component deltas + rounding residual = base readiness delta

Base readiness delta + readiness-adjustments delta = total readiness delta

The rounding residual exists because each component score is itself rounded before it contributes to the total — rather than hide that residual inside an unexplained gap, the scenario states it as its own number, always within ±1 point.

Worked with the real hydrogen-alternative scenario from the sample above: technology-class fit moved from 100 to 73 (a −27-point component delta; weight 20, so −5.4 weighted points), and delivery maturity moved from 76 to 58 (a −18-point delta; weight 10, so −1.8 weighted points). Every other component was unchanged. That reconciles as:

Technology-class fit: −27 × 20 ÷ 100−5.4
Delivery maturity: −18 × 10 ÷ 100−1.8
Sum of weighted component deltas−7.2
Rounding residual+0.2
Base readiness delta−7
Readiness-adjustments delta0
Total readiness delta (82 → 75)−7

Nothing in that chain is estimated after the fact — every row is a real, disclosed number from the same report, and the identity is checked directly by an automated test on every change to this mechanism (worker/scripts/scenario-intelligence-smoke.mjs).

Confidence, checked separately

Confidence rarely moves between a baseline and a comparison scenario, because neither duration nor technology-class identity feeds the confidence model directly — only through the timeline and procurement evaluation status. In both worked scenarios above, confidence stayed at High/83, and the stated reason is a real, structural fact, never a placeholder: “Unchanged — this scenario does not vary any structured field or dataset-sourcing input the confidence model reads.” When confidence does change, the scenario names the specific evidence-coverage component that flipped — never a generic “confidence changed” statement with no reason attached.

Re-evaluated vs. reused benchmarks

For each of the four benchmark-backed dimensions, a scenario states plainly whether that comparison was independently re-run for this scenario’s own duty point, or legitimately reused unchanged from the baseline — never silently carried over without saying so:

  • CAPEX position is always independently re-evaluated for every comparison scenario, since it depends on both technology class and duration. In the 28-hour duration scenario, it re-evaluates to “within observed central band” — a real change from the baseline’s “aggressive” position at 40 hours.
  • Interconnection-to-energization and interconnection-to-COD timeline evidence are legitimately reused unchanged in both scenarios — a real methodological fact, since neither metric is keyed by duration or technology class.
  • Procurement-timeline evidence is reused unchanged when technology class does not vary (the duration scenario), and independently re-evaluated when it does (the hydrogen-class scenario).

How findings transition

Every canonical finding on the baseline is diffed against the same scenario’s own, independently re-evaluated findings — never the baseline’s findings copied forward. Each one lands in exactly one of four states:

StateMeaning
ResolvedThe finding existed on the baseline and a real re-evaluation under this scenario removed it.
IntroducedThe finding did not exist on the baseline and a real evaluation under this scenario produced it.
UnchangedThe finding exists on both the baseline and this scenario — its size may still differ.
No longer evaluableThe finding existed on the baseline, but this scenario’s own re-evaluation could not be completed — distinct from "resolved," which means the risk was actually checked and cleared.

The duration-alternative scenario shows exactly this distinction in practice: the baseline’s “CAPEX benchmark position (aggressive)” finding transitions to resolved at 28 hours — a real re-evaluation, not a guess — while every other finding stays unchanged. Because that CAPEX finding was already scored neutral-by-policy (no readiness points either way), resolving it does not move the total score — the reconciliation identity above stays at zero even though a real finding changed status. Reading the transition alongside the component deltas, rather than instead of them, is what avoids the wrong conclusion here.

The hydrogen-class scenario shows the opposite direction: a new Technology Class Fit finding is introduced (it did not exist on the CAES baseline), while the CAPEX exception finding stays unchanged — it re-evaluates to the same “aggressive” position at hydrogen’s own duty point, so it is neither resolved nor newly introduced.

Why a stable scenario still explains itself

A scenario whose total readiness does not change is required to state exactly why, in one of four fixed ways: the varied assumption does not affect any scored rule; the alternative remains in the same decision band; an offsetting pair of component changes netted to zero; or the required evidence for comparison is not evaluable. The 28-hour duration scenario above states two of these at once, verbatim: “The varied assumption does not affect a scored rule under the current rubric. The alternative remains in the same decision band.” A stable scenario is never an unexplained duplicate of the baseline.

Two comparability caveats apply to every scenario, disclosed rather than assumed: a benchmark exception declared against the baseline is not automatically re-applied to a hypothetical scenario, and cross-assumption reconciliation findings are evaluated once, against the declared baseline only — never separately re-run per scenario.

Scope and limitations

Scenario Intelligence recalculates the same deterministic scoring and benchmark logic the baseline uses — it does not introduce a second, independently-reasoned model, and the sum of every component delta is checked to reconcile exactly to the total, on every scenario, every time. It does not recommend which scenario to pursue, does not imply that any alternative is more likely, feasible, or investable, and is capped at four alternatives alongside the baseline (five total), per the fixed scope of this audit.

See it in the real report

Every figure above, in context.

Read the full worked decision brief these scenarios are taken from, or the scoring methodology behind the components themselves.

Frequently asked questions

If a benchmark flag "resolves" in a scenario, does the readiness score go up?

Not necessarily. A benchmark-flag finding tied to a metric that is scored "neutral by policy" carries no readiness points either way, so resolving it can leave the total score exactly where it was. "Resolved" describes the finding’s own status, not an assured score change — always read it alongside the component deltas, not instead of them.

Does DurationX recommend which scenario to pick?

No. Scenario Intelligence states what changed, what stayed fixed, and what the deterministic consequences were. It does not rank scenarios or recommend one over another — that judgment, like every product decision here, belongs to the customer and its advisers.

Is a baseline benchmark position ever silently reused for a scenario that should re-trigger it?

No. CAPEX position is independently re-evaluated for every comparison scenario’s own duty point, and procurement timing is re-evaluated whenever technology class changes — both are stated explicitly as "re-evaluated" rather than "reused" whenever that recomputation happens, precisely so a stale baseline value is never mistaken for a fresh one.

Source note

Sources and methodology note

This article describes DurationX’s own scenario-generation mechanism. Its primary sources are therefore the real, currently-shipping code, specification, and worked example, not third-party research:

  1. 1worker/src/pipeline.ts — the `scenarioIntelligence` bundle: component deltas, the reconciliation identity, confidence notes, benchmark-lookup notes, issue transitions, carried assumptions, and the stable-scenario explanation.
  2. 2worker/src/scoring.ts — `generateScenariosR2UsefulnessGated`, the candidate search and usefulness gate that selects which alternatives are shown at all.
  3. 3Docs/specs/02-scoring-and-methodology.md §7.2 — the owning specification for every field named above.
  4. 4durationx.com/methodology, Section 6 — the public scenario- and memo-derivation summary.
  5. 5durationx.com/sample — the real, current worked example every figure in this article is taken from.