Skip to content
GenomarkerCORPUS AUDIT · METHODOLOGY

Methodology & corpus selection

An adversarial test we run on ourselves.

The audit exists to answer one question honestly: does the assurance model find real failures in real published science — and does it stay silent when there is nothing to find?

  1. Corpus selection

    Published computational biology analyses, 2019–2025, with public data and computational methods sections. No Genomarker involvement in the original work; no cherry-picking by expected outcome — selection criteria are fixed before re-execution.

  2. Spec reconstruction

    Each paper’s methods are transcribed into a canonical AnalysisSpec exactly as published — including its ambiguities. Where the publication under-specifies, the ambiguity itself becomes a finding, not a silent guess.

  3. Re-execution protocol

    The reconstructed spec runs through the same assurance kernel, the same versioned rules, and the same replay machinery used for customer analyses. Nothing bespoke, nothing softened.

  4. Classification

    Findings are classified by dimension and materiality. “Material” means capable of changing the paper’s interpretation — a deliberately conservative bar, applied by rule, not by editorial judgment.

  5. Disclosure ethics

    Aggregate results are public. Per-paper findings are shared with the original authors first, anonymized in public reporting, and never framed as accusations — the audit tests the assurance model, not the scientists.

  6. Illustrative figures — audit in progressVersioning

    This is report v3.6 of a living audit. Figures remain marked illustrative until the corpus reaches its pre-registered n = 40; every revision is archived and diffable.

Back to the audit report

PRE-REGISTERED PROTOCOL · SHA-PINNED REVISIONS