Population Validity
Population validity is the claim that a causal effect observed in one group of people will replicate in another. It is the most commonly discussed dimension of external validity, yet it is routinely misunderstood as a demographic matching problem — as if the question were whether the target population looks like the study sample. The deeper question is whether the underlying causal mechanisms that produced the effect in the source population still operate in the target population, even when demographics differ. A drug tested on young adults may fail in the elderly not because they are older, but because their homeostatic set-points have shifted, altering the causal pathway.
The standard approach to population validity — inclusion and exclusion criteria — is structurally conservative. It attempts to guarantee validity by restricting generalization to populations that resemble the trial. But this strategy defers the hard question: what features of a population actually matter for the causal mechanism? Until we know the mechanism, demographic matching is educated guesswork dressed as rigor.
The relationship between population validity and selection bias is particularly fraught: the very criteria that make a trial internally valid — homogeneous participants, controlled conditions — systematically undermine its population validity by excluding the heterogeneity that defines real-world populations.
Population Validity as a Cross-Scale Problem
The failure to generalize from study populations to target populations is often treated as a sampling problem — a mismatch between the sample and the population. But from a systems perspective, it is a cross-scale interaction problem. The study population and the target population operate as distinct scales, each with its own emergent properties, and the causal mechanisms that operate at one scale may not transfer to another.
Consider a clinical trial conducted in a controlled hospital setting (the study scale) and the same intervention deployed in community health clinics (the target scale). The hospital setting imposes constraints — specialized staff, intensive monitoring, compliant patients — that constitute a slow-scale structure. The intervention's effect is measured within this structure. But when the intervention moves to community clinics, the slow-scale structure changes: staff turnover is higher, monitoring is less intensive, patient compliance is lower. The fast-scale dynamics (the intervention's immediate effects) may be the same, but the slow-scale constraints that modulate those effects are different. The result is a classic cross-scale failure: the effect at the study scale does not transfer because the slow-scale architecture that supported it has been replaced.
This framing explains why effect size heterogeneity is the norm rather than the exception in meta-analysis. Each study operates in a slightly different institutional, cultural, and biological context. These contexts are not noise around a true effect. They are the slow-scale structures that constitute the conditions under which the effect exists. An intervention does not have one true effect. It has a distribution of effects, each conditioned on the slow-scale architecture of the population in which it was measured.
The implication is radical for research design. Instead of asking whether a study population represents a target population — the demographic matching frame — we should ask whether the study population's slow-scale structure is a plausible model for the target population's slow-scale structure. This requires understanding not who the participants are but what institutional, environmental, and biological constraints operate on them. A trial in Swedish hospitals may generalize better to German hospitals than to Swedish community clinics, not because Germans are more like Swedes than Swedes are, but because German hospitals share the slow-scale structure of Swedish hospitals.
Population validity is not a statistical property of samples. It is a dynamical property of systems. The question is not whether we have drawn from the right population. It is whether the causal mechanism we have measured operates in a context sufficiently similar to the one we care about. Demographics are a poor proxy for context, and until we stop treating them as a good one, we will continue to be surprised when effects fail to replicate.