Methodology

What the audit tests, how the two scores are computed and interpreted, what evidence backs the benchmark, and what has — and has not yet — been independently validated.

Who operates DurationX, the scope of human review, vendor independence, and report-security mechanics are stated once, on the home page’s Trust and accountability section and in full on Security & Data Handling — not repeated here. This page is scoped to one question: how the audit itself is built, tested, and interpreted.

Section 1

What the assumption-readiness audit tests

DurationX does not return a bare number. Every paid audit runs a fixed, versioned set of controls against the project you submit: whether required assumptions are present and internally consistent, whether the declared technology class and duration are coherent, whether the CAPEX range is complete and — where a public benchmark applies — positioned against it, whether a declared benchmark exception is traceable to evidence, whether comparison scenarios are recalculated rather than guessed, and whether the final score, findings, and memo reconcile with each other before a person signs off.

Section 2 names all thirteen control categories and what each one confirms. The readiness score in Section 3 summarizes the result of these controls — it is not a substitute for reading them, and no control is skipped to make the findings look larger or smaller than they are.

Section 2

Audit control protocol and result statuses

The audit is organized into thirteen disclosed control categories. Each one points to the specific deterministic sub-check, benchmark comparison, cross-field reconciliation rule, or human review step that actually performs it — no control duplicates another’s logic, and every category below traces to real code, not a marketing label.

  1. Required-assumption completeness. Confirms every required assumption field is declared, present, or explicitly marked pending.
  2. Unit and range validity. Confirms declared power/duration/CAPEX figures use a consistent unit basis and produce a coherent normalized value.
  3. Cross-field consistency. Confirms the interconnection capacity and value-stream declarations are internally consistent with the declared target power and streams.
  4. Technology, use-case, and duration coherence. Confirms the declared target duration is coherent with the selected technology class's benchmarked duration range.
  5. CAPEX scope, range, and benchmark applicability. Confirms the declared CAPEX range is complete, internally ordered, and — where evaluable — positioned against a matched public benchmark duty point.
  6. Reliability and cycling coherence. Confirms the declared reliability requirement and cycling pattern are coherent with the selected use case and the technology class's benchmarked life data.
  7. Grid and timeline benchmark applicability. Confirms the declared grid/interconnection timeline is internally consistent and, where evaluable, positioned against benchmarked interconnection lead times.
  8. Delivery and procurement evaluability. Confirms the declared procurement/delivery timeline is internally consistent and, where evaluable, positioned against benchmarked procurement lead times.
  9. Evidence origin and substantiation status. Confirms every material finding's evidence trail is visible, typed, and dated where applicable, with no evidence overstated as independently verified.
  10. Project-exception traceability. Confirms a declared benchmark exception carries a rationale, a declared evidence origin, and an exact evidence request to close it.
  11. Scenario robustness and recalculation. Confirms comparison scenarios are only shown when they produce a decision-relevant change versus the baseline, recomputed through the same benchmark path.
  12. Score, issue, memo, source, and PDF reconciliation. Confirms every customer-visible priority (top failures, required next inputs, roadmap) resolves to a real, traceable finding, and that the delivered PDF's hash and signature match the persisted record.
  13. Prohibited-claim and document-integrity review. Confirms the delivered narrative contains no prohibited claim or phrasing, and that the document's hash and signature are independently re-checkable.

Each control resolves to exactly one of four results:

ResultMeaning
SupportedThe declared assumption and, where a public benchmark applies, the matched comparison both check out.
Conditional — evidence pendingA finding is real but rests on weaker evidence — for example a declared exception with no external corroboration yet. Not wrong; not yet substantiated.
UnsupportedA check genuinely fails: a declared rationale does not match any real deviation, a required ordering is contradictory, or a prohibited statement was caught.
Not evaluatedNo applicable benchmark, source, or input exists to test this control for your specific project. Never guessed, never treated as a penalty.

This catalog and its underlying logic are implemented and verified against real fixtures today (worker/scripts/audit-control-register-smoke.mjs, 119 checks, 0 failed — re-run for this page). Every category’s own check already runs on every delivered report through the scoring, cross-assumption reconciliation, and human review steps described in Sections 3, 5, and 8 below. The separately itemized, per-category register — the same thirteen results shown individually, each with its own evidence references — renders as its own appendix in the authenticated report view once a report is produced under the report format built to carry it; where a report predates that format, the same checks’ results still reach you through the readiness score, the top assumption failures, Path to Ready, and the review checklist, just not as one separately itemized list.

Section 3

How readiness and confidence are interpreted

Every report shows two numbers: a readiness score (0–100) and a confidence band (High / Medium / Low). Readiness measures how strong the current decision setup is under the inputs provided. Confidence measures how complete and defensible the evidence base is. The two are computed independently and are never merged into one figure.

Why these weights

The eight component weights below are fixed by the Product Constitution and cannot be changed by this audit without a locked-document revision. The three heaviest components — technology-class fit (20), duration justification (15), and CAPEX completeness (15) — carry sixty of the available hundred points because they are the assumptions a technology-class or vendor conversation is built on: get the class-duration relationship wrong and every other assumption inherits the error; leave CAPEX incomplete and the project cannot be compared to any benchmark at all. The remaining five components carry ten points each — real, scored dimensions, but each affects a narrower slice of the decision than the first three. This is a design choice about which incomplete assumption does the most damage to a diligence conversation, not a claim that the lighter components matter less in an absolute sense.

ComponentWeight
Technology-class fit20
Duration justification15
CAPEX completeness15
Value-stack clarity10
Reliability case definition10
Grid and interconnection realism10
Operational use-case consistency10
Delivery and procurement maturity10
Total100

Each component is the sub-weighted mean of a small number of sub-checks. Each sub-check is an ordered decision table: the customer’s inputs meet exactly one rule, which emits one of five fixed band values (100 / 70 / 40 / 25 / 10). No qualitative judgment at runtime — same intake plus same pinned rubric equals the same score, always. No single missing input can drive a component below 10, which is the structural reason the Constitution’s own rule holds in practice: missing CAPEX limits readiness, it never zeroes the project out.

Why these thresholds

The four outcome bands are also fixed by the Constitution. They are spaced so a project must clear real evidence and consistency thresholds — not merely avoid outright contradiction — to reach “Ready,” while a project with one serious, isolated gap can still land in a middle band rather than being zeroed out, because that gap is bounded by its own component’s weight and floor, never by a special case in the outcome mapping itself.

ScoreOutcome label
80–100Ready for next diligence step
60–79Conditionally ready
40–59Insufficient evidence
0–39Not ready

"Assumption set ready for the next diligence step" means the submitted assumptions meet the stated completeness and consistency thresholds under the cited rubric. It does not mean the project is feasible, safe, financeable, permitted, investment-ready, bankable, or likely to succeed.

Why confidence rules work the way they do

Confidence is field-by-field completeness (present-valid / present-weak / explicit-absent / missing) plus a sourcing factor for benchmark provenance — computed independently of readiness. The band cutoffs (High ≥ 75, Medium 45–74, Low < 45) are set so a project with several unresolved or missing fields cannot reach High confidence regardless of how well the fields that are present score on readiness, because confidence is answering a different question — is there enough evidence here to trust the comparison — than readiness, which asks how strong the project’s setup is.

Why missing evidence lowers confidence first, readiness second

When an assumption is missing, two things happen, in a fixed order. First, confidence drops immediately and in full for that field, because an absent input weakens the evidence base right away. Second, readiness drops only through the one scoring component that consumes that field, and only down to that component’s own floor — never to zero, and never affecting any other component. This ordering exists so an honestly incomplete project reads as “evidence is thin here” first, and only secondarily as a lower component score: a project that has simply not gathered every input yet is never told its assumptions are bad — it is told the evidence is not there yet.

Why conservative assumptions do not automatically earn readiness

A customer who declares wide, cautious ranges, defers optional fields, or avoids committing to a specific number is not rewarded with a higher score for being careful. Readiness measures the completeness and internal consistency of the assumption set actually submitted, not how prudent its author was; a vague or hedged answer scores the same modest way as an equally incomplete confident one. Being cautious is not a substitute for being complete, and being decisive is not a substitute for being evidenced.

Path to Ready — how points and thresholds combine

Path to Ready is arithmetic, not a second model. The rubric already computes, for every assumption gap, how many readiness points that gap costs. Path to Ready ranks those already-computed costs from largest to smallest, shows the distance between the current readiness score and the next band threshold, and states which evidence a re-audit would assess first. When several scored checks fail on the same component, they are shown as one combined action with their combined points.

What it is not: a prediction that a project will succeed, a promise that a re-audit will reach a particular score, or advice to invest or procure. Points shown are a deterministic re-evaluation ceiling, subject to accepted evidence — whether that evidence actually closes a gap is determined only by a re-audit against the same version-pinned rubric.

Section 4

Benchmark applicability, active sources, and coverage

The benchmark set is derived from public techno-economic sources, in this priority order:

  1. Official public techno-economic datasets: PNNL Grid Energy Storage Cost and Performance, NREL Annual Technology Baseline (ATB), IEA storage summaries where recent.
  2. Official government or national-lab reports (DOE).
  3. Official market rules or tariff documents (interconnection and procurement).
  4. Customer-provided project inputs (never a benchmark source — always the audit subject).

Every source we cite carries five fields: source name, source date, geography, currency year where relevant, and version or retrieval date. Vendor materials, supplier price lists, and privately-sourced pricing are excluded from the primary benchmark.

DurationX does not sell storage systems, rank vendors, accept supplier referral fees, or receive commissions from technology suppliers. The benchmark may include public, licensed, and clearly identified supplier-originated evidence; supplier-originated evidence is labeled and does not become "independent" merely because DurationX uses it.

Why the applicability rule exists

Two benchmark comparisons — the CAPEX position and the grid/interconnection timeline position — are only shown when your project’s specific duty point (technology class, power, duration, geography) has a real, applicable comparison in the pinned dataset. The comparison never interpolates or extrapolates a value between two real data points, and never fits a curve across neighboring cells to invent one: it is always either the nearest real matched study point (CAPEX) or the country’s own published typical/stretched band (grid timeline). If no real, applicable comparison exists for your project’s specific duty point, the metric is marked not evaluable — never estimated, never guessed, and never scored as a penalty. A benchmark comparison that silently interpolated across dissimilar projects would be a worse kind of wrong than an honest “not evaluable.”

Active sources — current benchmark dataset

Generated directly from the source registry backing dataset v1-review-2026-07-06, not maintained by hand. A source below feeds exactly one disclosed role; none is implied to support every output.

SourceRole in this datasetEdition / yearGeography
PNNLCost curve — primaryESGC v2024 · 2024US
NRELCost curve — secondary cross-checkATB v2024 · 2024US
LBNLGrid-interconnection / commissioning queue-timelineQueued Up 2026 · 2026US
BLSCurrency-year (CPI) normalization input — not a directly displayed benchmark figureannual-average-through-2024 · 2024US
ECBCurrency (FX-to-EUR) normalization input — not a directly displayed benchmark figureannual-average-through-2024 · 2024EU

Current data coverage

Of the ten canonical technology classes, 7 — Compressed-air energy storage, Gravity storage, Hydrogen (electrolyzer + storage + fuel cell), Lithium-ion LFP, Lithium-ion NMC, Thermal storage (molten-salt / phase-change), Vanadium redox flow — have real cost-curve coverage drawn from the active sources above. 3 — Iron-air battery, Liquid-air energy storage, Zinc-bromide flow — do not yet have a qualifying independent cost source; public figures for these are limited to single-vendor claims or wide, inconsistent secondary estimates, not benchmark-grade data. An audit against one of these classes is scored on every other component as normal, but the CAPEX-benchmark comparison for that class is marked not evaluable rather than estimated.Evaluated class set for paid checkout today: v1-seven.

Interconnection queue-timeline coverage (how long grid interconnection typically takes) is currently limited to: US. A project outside these geographies is scored on every other component as normal; the grid-timeline benchmark comparison for it is marked not evaluable, never estimated from a different geography.

Procurement lead-time data (how long equipment delivery typically takes) does not yet have a qualifying independent source for any technology class. This sub-check is not scored against any figure until a reliable source is identified — it is never estimated or approximated.

Section 5

Exceptions and evidence treatment

A benchmark comparison that lands outside the typical range is not automatically an error. It becomes a project exception: a declared rationale, a declared evidence origin (customer, vendor, or an independent calculation), and an explicit statement of what additional evidence would substantiate or close it.

Public benchmark data describes a population of comparable projects, not any one specific project — a project can have a genuine, well-supported reason to sit outside the typical range (a distinctive site condition, a different commercial structure, a real vendor quote below the population median), and treating every deviation as a defect would penalize defensible engineering and commercial judgment. So the rule is to disclose the deviation, name the evidence that would substantiate it, and never invent an automatic penalty for it — this is why benchmark exceptions remain possible by design, not merely tolerated as an edge case.

Evidence origin is tracked but never silently upgraded: a customer-declared rationale stays labeled customer-declared even when it names a vendor or an independent calculation as its basis — it is confidence-relevant traceability information, never treated as independent verification of the underlying fact.

Every report states its confidence basis in exactly these terms (lib/confidence-basis.ts) — a fixed sentence naming what the band does and does not mean, plus a plain count of how many findings rest on a customer declaration versus an externally-benchmarked comparison versus a customer-uploaded file reviewed. Uploaded evidence never improves confidence merely by existing: it can affect a finding only through an explicit, reviewable link, never by file count alone — see the axis-isolation evidence in Section 7 below.

Section 6

How scenarios and the committee memo are derived

Up to five scenarios — the customer’s own baseline plus up to four alternatives — and the committee memo are not a second, independently-authored analysis. Both are deterministic projections over the same findings described above: a scenario re-runs the same benchmark and scoring logic against the changed assumption only, states which components moved and by how much, and states plainly when a benchmark lookup was re-run versus legitimately carried over unchanged from the baseline. A scenario whose total score does not change still states why — the changed variable does not affect any scored rule, an offsetting pair of components netted to zero, or the required evidence for comparison is not evaluable.

The committee memo assembles the same findings into one structured briefing note — current readiness position, conditions before deeper diligence, principal uncertainties, project exceptions, an action/evidence sequence, and a scope boundary — and is checked, before delivery, against a compliance scan that blocks any recommendation to proceed, hold, or invest. The memo may only frame a question about the assumption file’s readiness for the next diligence step, never a go/no-go call. See the worked demonstration of both on the sample report.

Section 7

Calibration and replay evidence

These tests prove deterministic reproducibility and internal consistency — the same pinned intake, dataset, rubric, and report-format version always produce the same readiness score, confidence band, scenario deltas, and reconciliation result, replayed byte-for-byte, every time. They do not, and cannot, prove that the methodology’s conclusions are empirically or externally valid — whether these specific weights, thresholds, and applicability rules are the objectively correct way to judge long-duration storage assumption readiness is a question only an independent energy-storage subject-matter review can answer, and that review has not yet happened (rung 5 below).

  1. 1 · Worked example

    A fictional project is run through the same deterministic scoring, benchmark resolution, canonical issue/evidence, scenario, and committee-memo generation code a paid audit uses, then rendered with the same report template. No real customer, company, or project appears in it. See the worked sample.

  2. 2 · Internal test evidence

    Before any scoring or report-rendering change ships, it is replayed against fixed internal test cases with known, pinned expected results — the same code path a paid audit uses. This is internal engineering verification by the team that built DurationX, not an independent audit. All counts below were produced by re-running the named test in this session, not copied from an earlier record:

    What it provesTestResult
    Readiness/confidence across all four outcome bandsr2-golden-replay.mjs — DB-free unit-level replay of the scoring engine42/42 passed. r1: 91 (High) · 68 (Medium) · 47 (Low) · 26 (Low). r2: 91 (High) · 68 (Medium) · 52 (Low) · 31 (Low).
    Same four cases, through the real end-to-end pathreal-http-golden-replay.mjs — real form submission → intake acceptance → job claim → dataset hydration → pipeline62/62 passed. Case A: r1 90 / r2 94 (both High). Case B: r1 75 (High) / r2 81 (Medium).
    Band-boundary crossings at the 39/40, 59/60, 79/80 cutoffsreadiness-band-boundary-smoke.mjs — DB-free unit test against the real outcome-label function16/16 passed. Every integer readiness value 0–100 maps to exactly one of the four locked labels, monotonically, with no skipped or duplicated band.
    Missing-evidence and evidence-origin axis isolationconfidence-basis-smoke.mjs + pv06-04-confidence-evidence-origin-shadow-calibration.mjs36/36 and 20/20 passed. Confidence is byte-identical whether a declared exception is customer-, vendor-, or independently-sourced, and whether 0 or 3 files are uploaded — a separate, non-scoring traceability field (not confidence) is where that distinction shows up instead.
    Scenario monotonicity and reconciliationscenario-intelligence-smoke.mjs61/61 passed. The sum of a scenario’s component-level deltas reconciles exactly to its total readiness delta, and every zero-change scenario carries one of the four permitted explanations, never an unexplained duplicate of the baseline.
    Timeline/CAPEX applicability and not-evaluable handlingtimeline-evaluator-smoke.mjs + audit-control-register-smoke.mjs68/68 and 119/119 passed. Real aggressive / within-band / conservative-or-long classifications and real not-evaluable outcomes; no fabricated benchmark position is ever produced.
    Historical replay stability (r1 isolation)report-architecture-reorg-smoke.mjs107/107 passed. A report produced before a given field or section existed renders that absence byte-for-byte identically on replay — never a fabricated or backfilled value.

    The first two rows are deliberately both cited because they test different things: r2-golden-replay.mjs calls the scoring engine’s functions directly against a hand-built benchmark set — a fast, DB-free unit-level check, not a full customer-path replay, per its own header comment. real-http-golden-replay.mjs is the real end-to-end path: an actual form submission through intake acceptance, job claiming, dataset hydration, and the pipeline, exactly as a paid audit runs. Neither substitutes for the other.

  3. 3 · Design-partner case note

    Not yet available. No consented, anonymized design-partner case note has been published.

  4. 4 · Consented customer outcome

    Not yet available. No customer has yet consented to describe a real, resulting next-step decision.

  5. 5 · Independent methodology review

    Not yet completed. The current status, scope, and date are stated once, in Section 8 below — not repeated here.

Section 8

Human quality control, AI role, versioning, signing, and providers

Versioning, signing, and verification

For a pinned intake, dataset, rubric, rules, and normalization inputs, the numeric scoring calculation is deterministic. Explanatory narrative may be AI-assisted and can vary; it is constrained by the stored calculation record and reviewed before release.

Every delivered report pins the exact dataset, scoring-rubric, template, and prompt versions used to produce it, plus the underlying model and a verification reference. Report metadata is signed with Ed25519 over a canonical JSON payload that includes the SHA-256 hash of the PDF file. Because rubric, dataset, rules, and template versions are revised over time, multiple versions can be active across different reports at once — this page does not carry a single current version number; the exact versions and evaluation date that apply to a given audit are shown only inside that report.

The public verifier at /verify re-reads the PDF bytes from private storage, recomputes the SHA-256 hash, and re-checks the signature. A tampered file or tampered metadata fails. The verifier never returns customer content — only valid or invalid.

AI role and human quality control

Numeric scoring is deterministic and version-pinned; no AI is involved in that calculation. Narrative explanations are produced under a version-pinned AI process — the model, prompt, and template versions are recorded on every report — and are reviewed by a person before release.

The score itself is not a project-success probability or a professional certification — it measures assumption completeness and internal consistency under the stated rubric, nothing broader. “Human-reviewed” and “manually quality-checked” mean exactly this, and nothing broader:

Before release, a person checks the report for input-to-output consistency, source/provenance display, prohibited claims, unresolved compliance flags, calculation reconciliation, and document integrity. This quality-control review is not an engineering review, investment recommendation, legal opinion, safety certification, or independent verification of every customer-provided fact.

Independent methodology review

A separate independent subject-matter reviewer is intended to review the identified methodology, benchmark compilation rules, rubric, and golden cases under a published scope and conflict-of-interest declaration. The reviewer does not certify individual customer projects.

Independent methodology review pending. The identified methodology, rubric, compiler rules, and benchmark dataset version have not yet been reviewed by an independent energy-storage subject-matter reviewer. This page will be updated with the reviewer identity, review date, covered versions, scope, and any conditions or exceptions once that review is complete and recorded.

Reviewer roles

Three roles are kept separate and are never collapsed into one another: Report quality reviewer: checks a specific report’s quality-control checklist; no specialist licence is implied. Methodology reviewer: independently reviews the versioned benchmark/scoring framework under a written scope and conflict-of-interest declaration. Customer advisers: licensed engineers, fire/safety professionals, interconnection specialists, lawyers, tax advisers, financial advisers, insurers, lenders, and procurement advisers retained by the customer for project-specific diligence.

Providers and data flow

ProviderRole
Supabase (EU)Database, authentication, private PDF storage
VercelWeb app and API
Hetzner (EU)Analysis worker
PaddleBilling (merchant of record)
ResendTransactional email
AnthropicReasoning and compliance LLM passes

Section 9

Scope, limitations, and legal

Customer project data is confidential by default. Raw customer data is not reused for model improvement without explicit opt-in; the opt-in default is off. Deletion requests are executed within thirty days of receipt with an audit-trail record kept.

DurationX provides a vendor-independent, human-quality-controlled assessment of the readiness of assumptions submitted for a long-duration energy-storage decision. The assessment is tied to the project, location, inputs, benchmark dataset, methodology, rubric, and report versions identified in the report and to information available on the stated evaluation date. DurationX does not independently verify every customer-provided fact and does not provide engineering design or sign-off, fire or electrical safety certification, interconnection approval, permitting advice, vendor selection or procurement advice, legal or tax advice, investment advice, a valuation, a bankability opinion, or project due diligence. Scores and labels are not probabilities of success and do not guarantee cost, schedule, performance, availability, revenue, returns, financing, regulatory approval, or any other outcome. Public and licensed benchmark information may be incomplete, delayed, model-based, geography-specific, or subject to source restrictions and revision. Missing or non-comparable data are reported as not evaluable and must not be interpreted as favourable or unfavourable evidence. Decisions remain the responsibility of the customer and its qualified advisers.

This is the same wording printed in every report’s disclaimer section — see the Terms of Service for the full service definition, and Section 8 above for AI role, human review scope, and provider data-flow detail.

Request the Assumption Readiness AuditSee sample brief