Abstract
What this paper establishes
Monitorscape Research studies institutional behaviour using records that can be incomplete, revised, duplicated, or removed at the source. This standard defines the controls required to turn those records into public evidence: explicit questions, versioned cohorts, field-level provenance, pre-specified transformations, human validation, reproducible outputs, and visible corrections.
At a glance
Key points
- 01
Every numerical claim must identify its observation unit, cohort, denominator, time period, source coverage, and transformation.
- 02
A source statement, a Monitorscape observation, an extracted value, and an analytical inference are separate evidence classes.
- 03
Missingness and unresolved links are reported as results because they determine what a dataset can support.
- 04
Material revisions produce a dated new version and a correction note that explains whether the conclusions changed.
Monitorscape Research examines how public rules move through institutions and become operational obligations. The raw material for that work is administrative evidence: web pages, registers, procedure histories, documents, publication identifiers, and the successive observations Monitorscape records from them. Those records are valuable precisely because they are longitudinal. They are also vulnerable to source redesigns, silent corrections, missing dates, duplicate publication, and changes in terminology.
This standard defines the minimum evidence needed for Monitorscape to make a public research claim. It applies to quantitative reports, data notes, working papers, research protocols, and methodological evaluations.
Start with a falsifiable question
Every study begins with a written protocol that defines the decision problem before outcome distributions are examined. The protocol records:
- the primary research question and any secondary questions;
- the observation unit;
- the covered institutions, source surfaces, and time period;
- inclusion, exclusion, and censoring rules;
- primary and secondary measures;
- planned comparisons and sensitivity analyses;
- known source changes or coverage gaps; and
- the conditions required for publication.
Exploratory analysis is allowed and often necessary. When it changes the planned method, the release labels the change as exploratory and explains when and why it was introduced.
Separate four evidence classes
Monitorscape publications distinguish the origin of a statement because different origins support different levels of confidence.
| Evidence class | Meaning | Example |
|---|---|---|
| Source fact | Information stated directly by an authoritative public source. | A ministry page states that comments close on 18 October. |
| Platform observation | Information about when or how Monitorscape captured a source. | The page was first observed by Monitorscape on 9 October. |
| Extracted field | A structured value parsed or classified from source material. | The date phrase is normalised to 2026-10-18. |
| Derived measure | A value calculated from one or more supported fields. | The observed comment window is nine calendar days. |
An observation timestamp does not become a source publication date. An extracted category does not become an institution’s own classification. A derived lifecycle status does not replace the procedural language in the official record.
Public tables and figures identify derived measures. The technical appendix defines each transformation and the fields on which it depends.
Preserve provenance at field level
The minimum provenance for a research field is the source identity, source URL or document reference, observation time, extraction method, and transformation version. Where the source provides its own stable identifier, Monitorscape retains it.
Source pages may be overwritten, removed, or moved. When permitted, the research evidence package retains a content fingerprint or snapshot reference sufficient to establish which version was analysed. The public release respects source licensing, personal-data obligations, and restrictions on redistributing documents.
If two official surfaces disagree, the study does not choose a value silently. The conflict is resolved using a study-specific precedence rule or remains marked as unresolved. The quality appendix reports the frequency and direction of those conflicts.
Define the cohort before calculating the result
A research result is inseparable from the population that produced it. Each release publishes a cohort flow that begins with all discovered records and shows every material inclusion, exclusion, deduplication, linkage, and validation step.
The release states:
- the number of source objects discovered;
- the number meeting the study’s subject-matter scope;
- the number removed as duplicates or non-normative material;
- the number eligible for each measure;
- the number missing each required field; and
- the final denominator behind every reported percentage.
Records do not disappear because their lifecycle is incomplete. Active, withdrawn, rejected, or unresolved records are retained when they provide information and are treated using an explicit status or censoring rule.
Treat linkage as a measured operation
Regulatory research frequently joins a proposal to a consultation page, parliamentary dossier, final instrument, or later amendment. A wrong link can create a precise but fictional timeline.
Studies use deterministic identifiers first. Candidate matches based on titles, text similarity, institution, subject, or timing require validation. Every linkage method publishes its yield, reviewed precision, unresolved rate, and known failure modes. A model confidence score is not a substitute for evidence that two records represent the same regulatory object.
When an analysis depends on several links, the release reports the complete-chain yield rather than only the success rate of each individual step.
Validate extraction and classification
Rules, parsers, and language models can accelerate extraction. Public claims still require a validation design suited to the consequence of error.
The default process is a stratified review across institutions, years, document types, and relevant outcome ranges. Reviewers use a written coding guide. For material fields, two independent reviews or an equivalent adjudication control are preferred. The release reports agreement, error categories, and corrections applied to the full cohort.
Validation metrics match the task. Classification reports class-level precision and recall where feasible. Date extraction reports exact-match accuracy and error magnitude. Linkage reports false-positive and unresolved-match rates. A single aggregate “accuracy” score is insufficient when rare errors can materially alter a result.
Report distributions and missingness
Regulatory timelines are usually skewed. Monitorscape therefore reports medians, quantiles, and distributions before means. Threshold claims include both numerator and denominator. Small cells are suppressed or aggregated when disclosure or stability requires it.
Missing data is not described only as a technical limitation. Missingness can reveal a public-information problem and can vary systematically by institution, instrument type, or year. Every release reports field availability across the analytical strata used for interpretation.
Imputation is used only when the assumption is defensible and pre-specified. Observed and imputed results are shown separately. A legally derived date records the rule and instrument classification used to derive it.
Distinguish description from causation
Most Monitorscape studies describe institutional behaviour. A difference between institutions, years, or instrument types does not by itself establish why the difference occurred.
Causal language requires a design capable of ruling out credible alternatives. Otherwise, the publication uses descriptive terms such as “was associated with,” “coincided with,” or “was observed alongside,” and names the main sources of confounding.
Institution rankings receive additional review. They are withheld when case mix, source availability, censoring, or institutional mandate makes a simple comparison misleading.
Make computational work reproducible
A quantitative publication must be reproducible from a versioned research extract or an equivalent protected evidence package. The analytical release records:
- data snapshot or cohort version;
- source-coverage matrix;
- schema and data dictionary;
- deterministic cleaning and transformation code;
- environment and dependency versions;
- random seeds where applicable;
- figure and table generation code; and
- a manifest linking public outputs to the calculation that produced them.
The public package includes code and aggregate data whenever licensing, privacy, and security constraints allow. If row-level data cannot be released, Monitorscape explains the restriction and publishes sufficient aggregate checks for readers to assess the result.
Disclose the role of AI
Monitorscape may use AI to classify documents, extract candidate dates, summarise text, propose record links, or assist with code and prose. A publication discloses AI use when it materially affects the evidence or analysis.
The research team remains responsible for every public claim. AI-generated fields used in an analysis receive task-appropriate validation. AI-generated prose is checked against the cited source and analytical output. Models are not listed as authors and model confidence is not presented as statistical certainty.
Review before release
Each publication receives three checks:
- Domain review verifies that the institutional and legal interpretation is defensible.
- Methods review verifies cohort construction, transformations, statistics, and limitations.
- Reproduction review regenerates the stated figures and tables from the release package.
A programme note can be released before these checks because it makes no findings. A working paper identifies the review steps that remain open. A publication labelled Published has passed all checks required by its declared scope.
Correct visibly and version materially
Editorial corrections that do not affect meaning may be updated in place. A material correction receives a dated notice describing the error, the affected outputs, the correction, and whether the conclusions changed.
A data refresh, revised cohort, changed transformation, or new analytical method creates a new version. Earlier versions remain identifiable. Monitorscape does not silently replace a result while keeping the same version label.
Retractions remain visible with an explanation and a link to any replacement work.
Publication types
Monitorscape Research uses five labels:
- Research protocol: a question and method specified before findings are released.
- Data note: evidence about coverage, provenance, quality, or a defined dataset.
- Working paper: substantive analysis open to further review or revision.
- Published research: analysis that has passed the declared review and reproduction checks.
- Correction or revision note: a documented change to an earlier release.
The label appears with the title, date, version, and author. Readers should be able to understand the maturity of a claim without reading the full paper.
Scope and responsibility
Monitorscape Research describes public institutions, regulatory processes, and observable operating conditions. It does not provide legal advice and does not determine whether an organisation is compliant.
Research priorities, commercial relationships, funding, and material conflicts are disclosed when they could reasonably affect interpretation. Editorial conclusions remain the responsibility of the named authors and reviewers.
Suggested citation
Monitorscape Research. Research methods and publication standard, version 1.0. 23 September 2026.