Skip to content
safemode.space

Validation study

Six independent raters were given 307 mapping units from this corpus and the instruments' own definitions, and asked what relationship each pair carries. This page reports what they returned, including where they disagreed with the corpus and where they disagreed with each other.

Competing interest

  • The sole author of this study owns the product the study evaluates.
  • The study has not been independently analysed.
  • No rater took part in computing or framing the results.

This statement appears before the findings rather than after them, so that what follows is read with it in hand.

Commercial activity and curation more generally are covered separately, on the about page.

What was measured

Every figure on this page describes the corpus as it stood at one moment: fingerprint 634a026855441b1d0fb92280e37f78c64449657fb367129931b88acdb20cdc3d, drawn on 2026-08-06 under master seed safemode-validation-2026-08-06. The corpus has been curated since. These numbers are not restated as it changes, because a completed measurement that moves is not a measurement.

raters6
rated rows returned557
distinct pairs rated307, matching the answer key exactly
shared core units, rated by all six50
extension rows257
Krippendorff alpha, core, nominal0.762
bootstrap interval, 5,000 resamples0.672 to 0.840, median 0.759
alpha with the two skipped ratings excluded0.775
pairwise agreement, 15 pairs74% to 88%
units where all six agreed29 of 50, with 13 at five of six and 4 each at four and three

One figure carries its denominator or does not appear. Alpha computed over high-confidence units alone is 1.000, and that is 17 units of the 50-unit core. 9 more were excluded because only one rater marked them high, and 24 carried no high-confidence rating at all. A perfect coefficient over a third of the core, reported without saying so, would be the most misleading number this study can produce.

The convergence, and what tempers it

10 shared units drew a unanimous mitigates from all six raters against a corpus that records addresses. Read alone, that is 10 mappings this platform graded more weakly than six independent readers would.

It does not read alone. Those units are also among the least confident judgements in the study, on the raters' own confidence field, and the corpus was often equally unsure:

  • 5 units: every rater said moderate, and the corpus records high. This is the divergence.
  • 2 units: every rater said moderate, and the corpus records moderate. Nothing diverges here at all.
  • 1 unit: every rater said high, and the corpus records high. This is agreement.
  • 2 units: the raters did not agree on confidence.

Say plainly which way that cuts, because it does not cut this platform's way. The counterweight is 5 units and not seven, so the convergence stands with less of an offset than a first reading of it suggested, not more. A page that revised that figure downwards while keeping the tempering language would be doing exactly what this study was run to detect.

Finding one: the spread is the finding, not the average

Agreement with the corpus is 51.0% on the shared core (153 of 300) and 37.7% on the extensions (97 of 257). Both averages are uninformative, and the breakdown is why. The averages are published beside it rather than instead of it, because each half is misleading without the other.

stratumnagree%
extension-B, NIS2 Article 21463780.4%
extension-D, supply chain463371.7%
extension-C, NIS2 implementing regulation Annex461737.0%
extension-F, jamming and RF27622.2%
extension-E, EU Space Act4648.7%
extension-A, CRA Annex I4600.0%

By the relationship the corpus recorded: triggers obligation 42 of 48 (87.5%), addresses 197 of 461 (42.7%), relates to 11 of 48 (22.9%). By instrument on the shared core: NIS2 Directive 84.6%, NIS2 implementing regulation 60.6%, EU Space Act 33.3%, Cyber Resilience Act 26.9%.

The corpus and six independent readers agree almost completely on reporting obligations and on NIS2 Article 21, and agree on nothing whatever on CRA Annex I. A single headline agreement rate would average those two facts into a number describing neither.

Finding two: doctrine, not error

102 rater comments were read and classified, one verdict per comment, across two symptoms: 82 where a rater answered mitigates against a recorded addresses, and 20 where a rater answered relates to against the same recorded value.

All 102 argue from the text of the provision they were shown. None imports a boundary the corpus draws elsewhere and none misreads a provision. The shape is the same across all six raters and all four instruments: where a provision mandates an outcome, a property or a named technical control, the raters treat it as an obstacle; where it mandates a policy, a procedure or a category of measures, they do not. That is the instrument's own rule applied literally, which is what raters told to use only the instrument's definitions are supposed to do.

So the disagreement is a boundary question rather than a quality problem, and this corpus's own doctrine already reaches part of it. The two findings are one finding read twice: agreement collapses exactly where the boundary question lives. The methodology page sets out where that line is drawn and why.

What was asked for, and what was done about it

One rater set four conditions on the study's authorship. One is adopted in narrowed form and three are recorded and not adopted:

  • A prominently declared competing interest, not in a footnote. adopted, narrowed to the authorship of this study rather than to commercial interest in general
  • Agreement statistics computed by someone who is neither a rater nor interested in the outcome. Recorded, not done.
  • Blind adjudication of disagreements. Recorded, not done.
  • Rater contributions declared under CRediT. Recorded, not done.

Those conditions were set on a methods paper rather than on this page, and the same reply attaches others: that the rater writes the methods and limitations sections unsoftened, and that the corpus, the drawing script, the seeds and the complete rating data are deposited with a DOI before submission rather than on acceptance. Publishing this page discharges none of them. Whether a methods paper follows is not decided here.

What does not publish

The raw instruments do not publish. Twelve files of individual judgements are a different consent from an aggregate figure, and the raters gave the second. They are archived and available on request. No rater is named anywhere on this page, at their own request, asked a second time after the ratings were returned.

The arithmetic is checkable without them, which is the point. The answer key is a committed file, so the agreement figures can be recomputed by anyone holding the repository. The alpha implementation is checked against a published worked example with heavily missing data, against an identity with a second coefficient that does not share its machinery, and against both bounds. Every returned file's digest is recorded, so a later edit to one is visible rather than arguable.

Corpus 2026.08.24-1, built 2026-08-24 from 226 techniques, 308 regulation articles, 125 ENISA controls, 2,610 framework controls, and 90 countermeasures.