Our last case file argued that the AliExpress fine was about proof, not policy. True, but "proof" is a word everyone nods at and nobody defines. Here is the version an auditor actually accepts.
A score you can't re-run is an opinion with a number attached
Say you scored a claim in March. By June, all that survives is the score itself. You can't say which sources you had, what weighed against them, or whether any of it still stands. The number isn't evidence anymore — it's a memory of evidence.
Reproducibility is what closes that gap. The same input produces the same explainable result, whether the person asking is a colleague, an editor, or a regulator.
What actually gets written down
Every TrustMark™ run persists three things, not one:
- The score, with its per-dimension breakdown and weights.
- The evidence — every source, tiered by integrity, each citation traceable back to what was fetched.
- The reasoning — the chain from raw signal to sub-score to verdict, in plain language.
All of it versioned and timestamped. Re-run it next month and you get the same number back, or you get a different one with a diff showing exactly what moved and why.
When the score changes, that's the feature
Evidence isn't static. A source issues a correction. New reporting lands. A domain's track record shifts. A system that returns the same verdict forever isn't stable — it's stale.
So claims and sources can be re-checked on demand, and the trail records the delta instead of overwriting the past. What you defend is never really the number. It's the record of how you arrived at it, including the part where you changed your mind.
The auditor's version of this
Three obligations converge on the same artefact:
- DSA Article 37 — independent audits. Reproducible re-scoring is what makes a mitigation measure checkable rather than asserted.
- EU AI Act — explainability and measurable confidence in AI-generated output, with hallucinations caught by multi-source corroboration rather than vibes.
- ISO/IEC 42001 — evidence for AI management-system controls, produced continuously instead of assembled the week before certification.
None of them ask you to be right about everything. They ask you to show your working.
Export the trail, attach it to the decision, hand it to whoever asks. That's the deliverable — not the score.
The question is never what you scored. It's whether you can still show your work six months later.
