Methodology walkthrough 6 min read

How to Read DurationX's Calibration Evidence: Reproducibility, Not Validity

What DurationX’s internal test evidence actually proves about its scoring methodology — and the specific, different question it cannot answer on its own.

DurationXPublished July 23, 2026Internal test evidence · independent review status stated below

Published by DurationX (legal entity: Dijin Teknoloji Limited Şirketi, Ankara, Türkiye), the vendor-independent, human-reviewed pre-CAPEX assumption-readiness audit for long-duration energy storage. DurationX is run today as a single founder-operated pilot, so this article — like every DurationX report and public page — is attributed to the operating entity rather than to an individual. See Security & Data Handling for the full operator profile, human-review scope, and vendor-independence statement.

Every DurationX report carries the same short phrase: scoring is deterministic. That phrase is easy to test and easy to over-read. This note explains exactly what the tests behind it prove, and the one thing they cannot.

The distinction: reproducibility versus validity

Methodology, Section 7 states this distinction as the single most important framing for reading any of DurationX’s internal test evidence, and it is worth restating here in full rather than paraphrased, since it is the exact claim — and the exact limit on that claim — this whole note is about:

These tests prove deterministic reproducibility and internal consistency — the same pinned intake, dataset, rubric, and report-format version always produce the same readiness score, confidence band, scenario deltas, and reconciliation result, replayed byte-for-byte, every time. They do not, and cannot, prove that the methodology’s conclusions are empirically or externally valid — whether these specific weights, thresholds, and applicability rules are the objectively correct way to judge long-duration storage assumption readiness is a question only an independent energy-storage subject-matter review can answer.

Reproducibility is a claim about the software: given the same pinned versions and the same inputs, the same numbers come out, every time, with no hidden randomness or silent drift. Validity is a claim about the world: whether those numbers are actually the right way to judge whether a project’s assumptions are ready for deeper diligence. A methodology can be perfectly reproducible and still be wrong about what matters — the two questions are independent, and conflating them is precisely the failure mode this framing exists to prevent.

What the calibration evidence actually tests

The full evidence table — exact test names and the exact pass counts from the most recent run — is published in one place, kept current as the suite grows: Methodology, Section 7. It is not repeated here, since a second copy would drift the moment the suite changes. What follows is a plain-language summary of the six categories it spans:

  1. Golden cases across all four outcome bands. Hand-worked projects with known expected results are replayed through the scoring engine, both as a direct unit-level call and through the real end-to-end HTTP path (intake, acceptance, job claim, dataset hydration, pipeline) — kept explicitly separate, never conflated.
  2. Band-boundary crossings. Every integer readiness score from 0 to 100 is checked against the outcome-label cutoffs, confirming the mapping is monotonic with no skipped or duplicated band at the 39/40, 59/60, and 79/80 boundaries.
  3. Missing-evidence and evidence-origin axis isolation. Confidence is checked to be byte-identical regardless of whether a declared exception names a customer, vendor, or independent evidence origin, and regardless of how many files are uploaded — confirming that origin and upload count affect traceability, never the confidence score itself.
  4. Scenario monotonicity and reconciliation. Every scenario's component-level score deltas are checked to reconcile exactly to its total readiness delta, and every zero-change scenario is checked to carry one of a fixed set of permitted explanations — never an unexplained duplicate of the baseline.
  5. Timeline/CAPEX applicability and not-evaluable handling. Real aggressive, within-band, and conservative-or-long classifications, and real not-evaluable outcomes, are checked directly — confirming no benchmark position is ever fabricated when a comparison genuinely does not apply.
  6. Historical replay stability. A report produced before a given field or section existed is checked to render that absence byte-for-byte identically on replay — confirming an older report is never silently backfilled with a value it never had.

Two of these categories are worth reading side by side: the golden-case replay is explicitly run two ways — once as a fast, DB-free unit-level call directly against the scoring engine, and once through the real end-to-end HTTP path an actual paid audit uses (intake submission, job claim, dataset hydration, and the full pipeline). Neither substitutes for the other; both are published because they prove different things — the first proves the scoring function itself is deterministic, the second proves the whole delivery chain around it does not introduce drift.

Why this framing matters for a buyer

A deterministic, well-tested engine is a real and checkable property — it means a report can be replayed, a correction can be verified, and a stated score is never the product of an unrepeatable judgment call. It is not, by itself, evidence that the rubric’s specific weights and thresholds are the objectively correct way to weigh a long-duration storage project’s assumptions. Treating internal test evidence as if it answered that second question would be exactly the kind of overclaim DurationX’s own publication rules exist to prevent.

Independent methodology review is the step that can speak to validity. Its current, real status:

Independent methodology review pending. The identified methodology, rubric, compiler rules, and benchmark dataset version have not yet been reviewed by an independent energy-storage subject-matter reviewer. This page will be updated with the reviewer identity, review date, covered versions, scope, and any conditions or exceptions once that review is complete and recorded.

This status is queried live from the same record Methodology, Section 8 shows — it is not a separate claim maintained by hand, and it will update automatically the moment a real review is recorded.

Scope and limitations

This note explains DurationX’s own internal engineering verification. It is not an independent audit of that verification, and it does not itself substantiate the methodology’s weights or thresholds — see Methodology, Section 3 for the published rationale behind those choices, which is a different kind of evidence again (design reasoning, not test evidence). Nothing in this note changes, or is evidence for changing, any scoring weight, threshold, or rubric value.

See the evidence in full

The complete, current test table.

Read the full seven-row calibration evidence table, the worked sample it is tested against, or the rationale behind the rubric itself.

Frequently asked questions

If the tests all pass, does that mean the methodology is correct?

No. Passing tests confirm the engine is deterministic and internally consistent — the same inputs always produce the same output, and that output does not contradict itself across scenarios, bands, or replays. It does not confirm that the underlying weights, thresholds, and applicability rules are the objectively correct way to judge assumption readiness. That is a separate, harder question for an independent energy-storage subject-matter review.

What would count as validity evidence, as opposed to calibration evidence?

An independent subject-matter reviewer examining the rubric, benchmark compilation rules, and golden cases against their own domain expertise and reaching an independent, published judgment — not a re-run of DurationX’s own test suite. See the current review status below.

Source note

Sources

This note describes DurationX’s own internal verification. Its primary source is therefore the published evidence itself, not third-party research:

  1. 1durationx.com/methodology, Section 7 — the current, complete calibration-evidence table and the reproducibility-vs-validity framing quoted above.
  2. 2Docs/specs/09-test-and-release.md — the owning test specification these scripts implement.
  3. 3durationx.com/methodology, Section 8 — the live independent-review status shown above.