Returning creators may perform better because you renewed the strongest first-round performers. A repeat influencer performance comparison must preserve the original roster, compare returners with their own earlier results, and separate changes in timing, spend, and briefs. These comparisons describe the selected roster. They cannot, by themselves, show that repeating a partnership caused better performance.
Start with the selection decision
Suppose your team renews creators who produced the most orders. Their second-round average is then compared with the average for everyone in the first round. The second group entered through a performance filter; the first did not.
This is survivor selection. The World Bank's explanation of selection bias describes the underlying problem: nonrandomly selected groups can differ systematically before the intervention being evaluated.
That selection can be useful for buying decisions. It still changes what a comparison means. "Our renewed roster produced more orders per creator" and "renewal increases a creator's orders" require different evidence.
Modash's long-term partnership guide recommends trials before longer commitments and describes using performance to decide who continues. It also discusses changing creative angles and adding paid distribution. Those operating choices are reasons to document both who returned and what changed.
Before calculating results, record each original creator's renewal status and decision reason. Use separate reasons for performance, price, availability, product fit, and missing evidence. A creator who declined a new offer should not silently become a performance rejection.
Freeze the outcome before comparing rounds
Choose one primary outcome and keep its definition unchanged. For a sales-oriented comparison, you might use attributed orders per creator during a fixed reporting period. Record the attribution model, channel scope, post date, reporting cutoff, and treatment of canceled or unpaid orders.
These details depend on your data source. Shopify's marketing-report documentation says its reports cover Online Store channel orders. Its last-click model includes direct visits; last non-direct click follows a different allocation rule. The reports also include canceled, pending, and unpaid orders while excluding test and deleted orders.
Therefore, an attributed order count from those reports is not a count of settled payments. Use the same model and order-status rules in both rounds. Keep unknown attribution visible rather than allocating it to whichever creator seems most likely.
A reporting period is also distinct from an attribution rule. "Orders recorded during the 14 days after publication" specifies when you count outcomes. It does not establish which earlier interaction receives credit. Write down both rules. The article on defining an affiliate attribution window helps resolve that choice before you compare creators.
The example below uses attributed orders only. It makes no comparison between platform view, reach, or engagement definitions.
A renewal cohort with the selection effect visible
The following data are hypothetical. Six creators each publish one first-round placement. The brand renews A and B because they had the highest attributed order counts. It also recruits G and H for round two.
Assume equal placement counts and budgets, the same order-counting rules, and a complete 14-day reporting period for every placement. Round two takes place in a later calendar period. These assumptions make the arithmetic comparable; they do not remove selection or calendar effects.
| Creator | Round one orders | Round two orders | Status |
|---|---|---|---|
| A | 80 | 72 | Renewed after strongest trial result |
| B | 60 | 56 | Renewed after second-strongest result |
| C | 40 | Not activated | Not renewed |
| D | 20 | Not activated | Not renewed |
| E | 10 | Not activated | Not renewed |
| F | 0 | Not activated | Not renewed |
| G | Not enrolled | 32 | New in round two |
| H | Not enrolled | 28 | New in round two |
Do not replace "Not activated" with zero. You did not observe C through F under a second paid placement. Their first-round results stay in the original roster, but they have no second-round placement outcome.
The same table supports several different statements:
| Comparison | Reproducible calculation | What it describes |
|---|---|---|
| Original trial average | 210 orders / 6 creators = 35 | All six trial creators |
| Returners before renewal | 140 / 2 = 70 | A and B already outperformed the trial average |
| Returners after renewal | 128 / 2 = 64 | A and B's second-round average |
| Returners versus original roster | (64 / 35 - 1) × 100 = 82.9% higher | A selected group versus the full starting group |
| Same returners across rounds | (64 / 70 - 1) × 100 = 8.6% lower | Change within A and B |
| Returners versus current newcomers | (64 / 30 - 1) × 100 = 113.3% higher | Two selected returners versus two new creators |
The returners beat the original roster while declining against their own baseline. Both statements are true. Neither calculation isolates the effect of another partnership.
The newcomer comparison shares the calendar period, but the brand has already tested the returners. Newcomers have not passed that same filter. Matching them on audience size would not erase the difference in how they were chosen.
Mark changes that the averages cannot explain
Build one row per creator per activation. Preserve cohort membership using the first activation date, rather than reclassifying past rows whenever a creator returns. For the mechanics of that choice, use a cohort definition that stays fixed.
Add these comparison flags alongside the outcome:
| Flag | What to record | How to handle it |
|---|---|---|
| Selection | Renewal reason and first-round result | Display the returners' baseline beside the full roster |
| Brief | Product, offer, format, deliverables, creative direction | Separate materially different assignments |
| Distribution | Paid support and placement count | Compare like-for-like activity before combining totals |
| Calendar | Publication week, promotion, stock availability | Compare within the same campaign period where possible |
| Measurement | Attribution model, cutoff, missing records | Separate incompatible definitions; disclose missing coverage |
These are recommended reporting controls, not platform requirements.
Calendar matching matters because timing can affect the measured outcome. NIST's randomized block design guidance explains how experiments hold important background conditions constant within blocks. Campaign weeks and offers are possible blocking factors when designing a test. Sorting historical rows into weeks does not turn them into a randomized experiment.
If every returner receives a stronger discount in a holiday sale and every newcomer posts afterward, mark the comparison as confounded. There is no overlapping offer-and-period group in that dataset. Do not invent a seasonal adjustment percentage to fill the gap.
Decide what evidence the renewal decision needs
For roster planning, descriptive evidence may be enough. Report the selected cohort's current output, its prior output, cost, and delivery record. You can decide that a returning creator is worth hiring without claiming that relationship length caused their performance.
For a causal question, define a test before choosing winners. One possible design assigns eligible creators to an additional early activation or no additional early activation, then gives both groups the same later activation. Randomize within planned campaign blocks and compare outcomes at the common later activation. This tests the effect of that extra prior activation among eligible creators, under those conditions.
Agree eligibility, timing, attribution, and the outcome in advance. Plan sample size and analysis with a statistician; keep dropouts in the assignment record. Audience overlap, cross-group exposure, and departures from the planned briefs still need attention. The distinction between attributed revenue and incremental sales also applies to the outcome you choose.
For your next renewal meeting, bring the full starting roster and calculate the returners' first-round average before showing their latest result. Put the renewal reasons and changed conditions beside those two numbers.



