Blog Campaign measurement Guide

Measure incremental impact when attribution is incomplete

Choose a creator campaign holdout, geographic test or observational comparison, with an experiment worksheet and a worked example of an inconclusive result.

Two separated trays of purchase tokens beneath a split delivery funnel, with a balance resting near level between them.

To test influencer marketing incrementality, compare sales outcomes for groups assigned to receive different campaign conditions. Prefer a randomized holdout when you can control delivery. Use a geographic experiment when you can separate markets. If neither is possible, estimate a comparison with observational data and state the assumptions. Missing click attribution does not prevent every experiment, but you still need reliable sales outcomes for each group.

Decide what you can withhold

Google Analytics defines attribution as assigning credit for actions across touchpoints. A campaign experiment asks how purchases change when you change the campaign condition. A tracked order alone cannot answer whether that purchase would have happened anyway.

Start by naming the treatment. Is it the creator's public post, paid distribution of that post, a discount, or the whole package? Holding out paid distribution while everyone can see the public post estimates the additional effect of paid distribution. It does not estimate the full creator relationship.

Use this selection table before commissioning an analysis.

What you can controlDesign to considerMain limitation
Assignment of eligible people to campaign delivery or suppressionRandomized audience holdoutOrganic exposure, sharing and incomplete outcome matching can weaken the comparison
Delivery across separate marketsRandomized geographic testCross-market exposure and too few independent markets limit inference
No controlled delivery, but usable untreated comparison dataObservational counterfactual modelConclusions depend on assumptions about the comparison
Only sales before and after a public postDescriptive before/after reportTiming alone cannot separate the post from other causes

These are design recommendations. They do not imply that a particular platform or account gives you access to a holdout tool.

Randomize before delivery

For an audience holdout, define the eligible population first. Randomly assign people to campaign delivery or the control condition before launch. Keep their original assignments when analyzing results, including people assigned to treatment who never receive an impression.

Comparing people who clicked with people who did not click answers a different question. Clicking is a choice that can reflect existing purchase intent. Comparing only people reached also discards the protection of the original assignment.

Measure purchases for both assigned groups using the same rules. A holdout still fails if you can observe outcomes only for campaign clickers. Check whether lost identifiers, unmatched orders or missing records differ by group. Do not silently count every unmatched person as a nonbuyer.

For creator campaigns, write down likely contamination routes: a public post, a shared discount code, another creator's overlapping audience, or other paid distribution. If controls also receive the tested treatment, the comparison no longer represents a clean campaign-versus-no-campaign test.

Geographic tests need separate markets

Google Research's geo-experiment publication describes random assignment of non-overlapping regions, with geographic ad targeting implementing the conditions. This design can support aggregate sales measurement without tracing every purchase to a creator link.

For planning, examine markets' historical sales patterns and operational differences before assigning treatment. Record stock availability, regional promotions and changes in other media. Decide how orders enter a market, such as a consistent shipping-location rule.

A creator's home city does not establish where their audience sees the content. If a nationwide public post reaches treatment and control markets, local creator selection alone has not created a geographic holdout.

Ask an analyst to plan the number of markets and the analysis before launch. Regions are the randomized units. Thousands of orders inside two regions do not become thousands of independently randomized observations. The person-level calculation below must not be reused for that design.

Complete the experiment worksheet

Fill this in with the person who owns order data and the person who controls delivery. Resolve missing answers before committing the test budget.

Worksheet fieldDecision to record
Budget decisionWhat spending change will this result inform?
TreatmentWhich posts, ads, offers and dates change?
ControlWhat remains available without the tested treatment?
Assignment unitPerson, household or market; define eligibility before launch
Primary outcomeOne purchase definition, data source and denominator
Outcome coverageHow will outcomes be observed in both groups?
TimingStart, stop, purchase window and refund-maturity date
Detectable effectSmallest change worth funding; sample-size and power plan
AnalysisEstimator, uncertainty method and treatment of missing data
ExceptionsRules for stockouts, contamination and delivery failures
Decision ruleWhat evidence supports scaling, stopping or another test?

Choose the reporting window before seeing results. Late purchases and refunds can change the outcome, so a publication date alone is an incomplete window definition.

Keep test costs consistent with the decision. For the financial follow-through, build a complete campaign cost ledger before turning estimated additional purchases into a spending recommendation.

A positive estimate that remains inconclusive

The following data are hypothetical. Suppose a brand can randomize people into two groups and observe every person's purchase outcome. It tests paid creator-content distribution for 30 days. The primary outcome is whether an assigned person has at least one paid order remaining after cancellations and full refunds, assessed 14 days after the purchase window closes. Each person counts once.

Assume independent person-level outcomes, no spillover, and a fixed analysis date. Both groups continue receiving the same other marketing.

Hypothetical resultTreatmentControl
Assigned people5,0005,000
People with a qualifying purchase220200
People without a qualifying purchase4,7804,800
Purchase rate4.4%4.0%

The observed difference is 0.4 percentage points. Relative to the control rate, that is a 10% increase. Applied to the treatment group, the point estimate is 20 additional purchasers. None of those statements establishes that the true effect is positive.

NIST's large-sample two-proportion method gives a reproducible test of equal purchase rates. Here, both samples have ample purchase and nonpurchase counts for the normal approximation.

Hypothetical calculationExpressionResult
Pooled purchase rate(220 + 200) / (5,000 + 5,000)0.042
Standard error under equal ratessqrt(0.042 × 0.958 × (1/5,000 + 1/5,000))0.004012
z statistic(0.044 − 0.040) / 0.0040120.997
Two-sided normal p-value2 × (1 − normal CDF(0.997))0.319

At a predeclared two-sided 5% significance level, this result does not reject equal rates. It neither establishes lift nor proves that the campaign had no effect. The p-value is not the probability that the campaign failed.

Report the positive point estimate alongside the inconclusive test. Request an effect interval in the final analysis to assess the range against the minimum worthwhile effect. Use confidence intervals to describe uncertainty rather than treating the estimate as a guaranteed return. Do not keep extending the experiment until its p-value crosses a threshold; plan any follow-up separately.

When a controlled test is unavailable

An observational model can predict what sales might have been without the campaign. Its credibility depends on the comparison data.

The CausalImpact authors specify two important assumptions: control series remain unaffected by treatment, and their relationship with the treated outcome remains stable after launch. They also recommend testing imaginary intervention dates in pre-campaign data. A model that invents lift before the campaign gives you a reason to investigate.

For a creator campaign, an ostensibly untreated product can be a poor control if the promotion also increases demand for it. A concurrent regional sale can break the historical relationship between markets. Document these threats; a fitted model cannot remove them by declaration.

Modash's marketing mix modeling article discusses incomplete promo-code tracking and recommends checking models against randomized experiments. Use aggregate modeling to inform broader budget questions when the data support it. Do not assume a model that predicts total sales well has isolated one creator's contribution.

Before the next launch, complete the worksheet's treatment, control and outcome-coverage rows. If you cannot describe a credible comparison, label the report observational and decide what delivery control the next campaign needs.

Sources

  1. When Promo Codes Aren't Enough: Measuring Influencer Campaigns with Marketing Mix Modeling Modashaccessed Sep 30, 2026
  2. Get started with attribution Google Analytics Helpaccessed Sep 30, 2026
  3. Measuring Ad Effectiveness Using Geo Experiments Google Researchaccessed Sep 30, 2026
  4. How can we determine whether two processes produce the same proportion of defectives? NIST/SEMATECH e-Handbook of Statistical Methodsaccessed Sep 30, 2026
  5. CausalImpact: An R package for causal inference using Bayesian structural time-series models Google CausalImpact authorsaccessed Sep 30, 2026