Causal Identification

Marketing questions look like prediction problems but almost always demand counterfactual reasoning. "What would sales have been without this campaign?" and "what is the incremental return on another pound of search spend?" are comparisons between the observed world and a world that did not occur. A model can predict the observed target well and still answer these questions wrongly, because media is deployed precisely when demand is expected to be unusual. This page sets out the reasoning that keeps an Epsilon analysis honest, and the principled variable-selection judgement it requires.

The Reasoning Chain

A defensible analysis follows one chain, in order:

\[\text{estimand} \;\to\; \text{identification assumptions} \;\to\; \text{estimator} \;\to\; \text{inference} \;\to\; \text{diagnostics}.\]

The temptation is to start at the estimator — write down a Turing model, fit it, and interpret the coefficients. Naming the estimand first prevents an association produced by a convenient regression from being reinterpreted after the fact as the effect you wanted.

Let $Y_t(\bar{a})$ be the potential outcome at time $t$ under a media history $\bar{a} = (a_1, \ldots, a_t)$. A typical estimand is the expected incremental outcome of replacing a reference plan $\bar{a}$ with an alternative $\bar{a}'$:

\[\tau_t(\bar{a}', \bar{a}) = \mathbb{E}\left[Y_t(\bar{a}') - Y_t(\bar{a})\right].\]

The estimand must specify the intervention precisely — a one-week increase, a sustained quarter-long increase, the removal of a campaign, or a reallocation — together with the population, geography, outcome scale, and time horizon. "Holding other things fixed" is not sufficient until it is clear which variables are held fixed and whether doing so is a coherent intervention. Carryover makes this sharper: a change in week $t$ can move outcomes for several weeks, so the estimand is often a cumulative effect over a window,

\[\tau_{t:t+H} = \mathbb{E}\left[ \sum_{h=0}^{H} Y_{t+h}(\bar{a}') - \sum_{h=0}^{H} Y_{t+h}(\bar{a}) \right].\]

Only after the estimand is named should the identification assumptions be stated: the conditions under which the observational data recover $\tau$. The estimator then implements those assumptions, inference quantifies uncertainty conditional on them, and diagnostics check the computation. Convergence statistics such as $\hat{R}$ and effective sample size, and posterior predictive checks, verify that the sampler approximated the posterior of the chosen estimator. They say nothing about whether the identifying assumptions hold. Posterior intervals quantify uncertainty conditional on the model and its assumptions; they do not automatically include uncertainty about whether those assumptions are true. If identification fails, a well-converged model returns a precise estimate of the wrong quantity.

Media Assignment Is The Central Threat

Campaign timing and intensity are endogenous to brand strategy. Budgets respond to forecasts, launches, promotions, competitor activity, and managerial judgement, and those same factors often move the outcome directly. Writing the mean function schematically as

\[Y_t = \beta a_t + g(x_t) + u_t,\]

where $a_t$ is media, $x_t$ are observed common causes, and $u_t$ collects unobserved demand determinants, naïve regression fails to identify $\beta$ whenever media responds to those unobserved determinants:

\[\operatorname{Cov}(a_t, u_t) \neq 0.\]

Advertising scheduled ahead of an expected peak can be credited with demand that would have occurred anyway, biasing the effect upward. Where managers advertise defensively into weak periods that their private forecasts predict, the observed association can understate the effect or even carry the wrong sign. No model component decides which mechanism applies; it must be reasoned about from how spend was actually set.

Controls Chosen Causally, Not By Correlation

A covariate's correlation with media spend does not by itself make it a bad control. Confounders are often correlated with treatment precisely because they drive assignment. Dropping every variable associated with spend to avoid multicollinearity would remove many of the variables needed to block non-causal back-door paths — a common and damaging fallacy.

Good controls are ordinarily predetermined variables: quantities fixed or known before the media decision and not caused by the intervention. They earn their place by predicting the untreated baseline outcome, the assignment process, or both. Weather, holiday calendars, long-run demographic composition, and stable store characteristics such as store size are typical examples. Adjusting for the confounders $c$ that satisfy the back-door criterion recovers the interventional mean,

\[\mathbb{E}[Y \mid \operatorname{do}(a)] = \sum_{c} \mathbb{E}[Y \mid a, c]\, P(c),\]

so a strong predictor of baseline sales improves precision even when it is only weakly related to spend, and a predictor of assignment reduces confounding even when its direct association with the outcome is modest.

"Pre-treatment" is judged relative to the decision time, not the calendar. Contemporaneous weather may be predetermined for one-week-ahead planning, whereas a managerial forecast available to planners but absent from the data can leave residual confounding. Store size is predetermined for a short campaign but may itself respond to marketing over a multi-year horizon.

Bad controls are variables caused by treatment: post-treatment quantities, mediators, treatment-path managerial choices, and colliders. If media raises branded search which then raises sales, conditioning on search volume strips the media channel of the downstream demand it created. This is over-control or mediator bias, and it estimates a direct rather than a total effect. Colliders are subtler. Where a variable $c$ is a common effect, $a \to c \leftarrow u$, conditioning on $c$ opens a spurious path between $a$ and $u$. Algorithmically generated "campaign priority" or "optimisation" scores that react to both campaign volume and unobserved organic demand are realistic MMM colliders. M-bias is the related case in which adjusting for a pre-treatment collider of two latent causes opens an otherwise blocked path — so pre-treatment status is necessary but not sufficient for a safe control. A mediator is not automatically useful either: identification through it (the front-door strategy) requires the mediator to intercept the whole effect with no unblocked direct path and adequately controlled mediator confounding, conditions rarely met by adding intermediate metrics to an MMM.

Price And Competitor Advertising: An Explicit Trade-off

The hardest covariates — price, promotional depth, competitor advertising — are neither clean pre-treatment controls nor obvious mediators. They typically respond to the same latent demand shocks that drive your own media decisions, and may be jointly determined with spend in market equilibrium. Price is the canonical case. A structural pair,

\[Q_t = \alpha + \beta P_t + \gamma a_t + \varepsilon_t, \qquad P_t = \delta + \phi Q_t + \psi z_t + \eta_t,\]

shows the problem: because $P_t$ depends on $Q_t$, it is correlated with $\varepsilon_t$, so ordinary regression on $\beta$ — and, through it, on the media coefficient $\gamma$ — is inconsistent. This is simultaneity (endogeneity) bias.

Omitting price instead leaves an omitted-variable bias. Its sign follows the familiar linear intuition,

\[\operatorname{Bias}(\hat{\beta}_{\text{media}}) \approx \gamma_P \,\frac{\operatorname{Cov}(\text{media}, \text{price})} {\operatorname{Var}(\text{media})},\]

so if discounts (lower price, with a negative demand coefficient) coincide with higher spend (a negative covariance), omitting price tends to bias the media effect upward, crediting media for discount-driven sales. If high prices coincide with high spend and suppress demand, omission can bias the effect downward. These directions are conditional on the signs and timing of the relationships, not universal.

The analyst must therefore own an explicit trade-off, and state it:

  • Include the endogenous covariate and accept simultaneity, over-control, or collider bias in the media coefficient; or
  • Omit it and accept omitted-variable bias when it moves demand and covaries with media.

Neither choice is free. Documentation should record the assumed causal ordering, what planners knew when decisions were made, and how conclusions move under plausible alternative treatments of the endogenous covariate. Competitor advertising deserves the same care: activity triggered by the focal campaign is a mediator of strategic response, whereas activity driven by a shared market shock is a confounder, and the correct handling differs.

Classify The Problem First

Before choosing an approach, classify the problem along two independent dimensions.

The identification threat asks why a naïve regression would fail: confounding by observed or latent demand shocks, reverse causation, simultaneous determination of spend and outcome, endogenous campaign selection, collider or sample-selection bias, or omitted variables. Different threats need different remedies; adding seasonal flexibility addresses an omitted seasonal pattern but does nothing about spend that responds to a private sales forecast, and hierarchical pooling stabilises estimates without removing confounding shared across units.

The temporal structure asks how the effect unfolds: carryover (adstock), delayed response, saturation, anticipation, pre-trends, and overlap between campaigns. Temporal structure is not itself an identification strategy, but it interacts with one. Anticipatory purchasing moves outcomes before the recorded campaign, undermining assumptions that pre-campaign observations represent an unaffected comparison path. Pre-trends can reveal trajectories inconsistent with parallel-paths reasoning, though their absence does not prove it. Carryover, for example the recursion

\[\tilde{x}_t = x_t + \lambda\, \tilde{x}_{t-1},\]

means "holding spend fixed" must refer to an exposure stock over a window, not a single period; and when demand shocks are autocorrelated, the carryover parameter $\lambda$ can act as a sponge for persistent omitted confounding, making genuine media decay hard to separate from an omitted trend. Closely spaced campaigns can leave separate channel effects only weakly identified, because flexible adstock fits the overlap while leaving many decompositions observationally equivalent.

Two Assumptions That Are Easy To Forget

Two further conditions sit underneath any causal reading and are worth stating explicitly.

Overlap (positivity). A counterfactual spend point is only supported if the data contain comparable spend levels. Response curves and optimised allocations that stray far outside the observed range rest on extrapolation, not evidence; the relative bounds in Budget Optimisation exist to keep the solver inside the supported region.

Measurement and aggregation. Spend is a proxy for exposure, and national or weekly aggregation can mask heterogeneity and manufacture associations that do not hold at the level of the decision. Aggregation bias and proxy error are identification threats in their own right, distinct from confounding, and they are not fixed by a better sampler.

What Observational MMM Can And Cannot Deliver

Under credible assignment, measurement, temporal, and functional-form assumptions, observational MMM can estimate useful incremental effects, propagate parameter uncertainty, pool across units, and — most valuably — make its assumptions explicit. Its causal interpretation is nonetheless conditional on those assumptions.

Regularising priors, nonnegative-coefficient constraints, and saturation curves make an otherwise unstable model estimable. They are inductive bias, not identification. A sign constraint can prevent an implausible negative estimate but cannot distinguish a true positive effect from positive confounding; it merely forces a possibly biased parameter into an acceptable range. A narrow prior reduces posterior width while increasing dependence on the prior.

The strongest remedy is experimental calibration. Geo-lift and conversion-lift experiments supply design-based variation that breaks the link between media and unobserved demand, and their estimates can enter the model as calibration priors. The experiment must match the MMM estimand closely enough in population, treatment contrast, outcome, and horizon; otherwise transport assumptions are required. Where experiments are unavailable, rely on sensitivity analysis — refitting under alternative control sets, lag structures, and prior scales, and asking how large an unmeasured confounder would have to be to overturn a conclusion — and on triangulation against quasi-experiments and platform lift studies. Agreement across methods with different failure modes is more persuasive than precision from a single specification.

A well-fitted, well-converged Epsilon model is a disciplined summary of what the observed data and the stated assumptions imply. It is not, by itself, proof of a causal effect.