MishaBook a demo

Aug 14, 2026

Common Cohorts Mistakes on Shopify

A cohort is a group of customers acquired or grouped by a shared characteristic (signup date, first purchase value, channel) tracked across time periods to measure retention, repeat purchase rate, or lifetime value. Cohort analysis isolates the effect of acquisition timing or segment on behavior.

What Cohorts Actually Measure

Cohort analysis answers: Do customers acquired in January behave differently than those acquired in February? Do high-AOV first-time buyers retain better than low-AOV buyers? The output is a table where rows are cohorts and columns are time periods (weeks, months post-acquisition), with a metric (repeat purchase rate, revenue per customer, churn rate) in each cell.

The core assumption is that cohort membership explains variance in behavior. If it doesn't - if January and February cohorts have identical retention curves - the cohort split adds no signal. Many Shopify operators build cohorts anyway, then misinterpret noise as insight.

Mistake 1: Cohorts Too Small to Read

A cohort with 15 customers acquired in a single week will show wild swings in repeat purchase rate (6.7% per customer = 1 repeat buyer). By week 4, one extra repeat purchase shifts the rate by 6.7 percentage points. This is statistical noise, not signal.

Threshold: Minimum 100 customers per cohort for repeat purchase rate or churn rate. Minimum 200 for LTV or AOV comparisons. If a Shopify store does $50k/month at $100 AOV, that's 500 customers/month - enough for weekly cohorts. At $20k/month, use monthly cohorts instead.

  • Count customers in each cohort before building the table
  • If any cohort has <100 customers, widen the grouping (week to month, channel-only to channel + device)
  • Document the sample size in the cohort table header

Mistake 2: Mixing Acquisition and Behavioral Definitions

A cohort must be defined by ONE dimension: acquisition date, acquisition channel, first purchase value, or geography. Mixing them - 'January email signups with AOV >$150' - creates a cohort that conflates two variables. If retention is high, is it because they were acquired in January, via email, or because they spent more? Impossible to know.

The fix: Build separate cohort tables. One by acquisition date (to spot seasonal trends). One by channel (to compare email vs. paid vs. organic). One by first-purchase AOV (to test if high-spenders retain better). Each table answers one question.

  • Define the cohort dimension before querying
  • One dimension per table
  • If you need to control for a second variable, filter it (e.g., 'email cohorts, excluding wholesale')

Mistake 3: Ignoring Seasonality and External Events

A January cohort acquired during New Year's resolutions will have different repeat rates than a July cohort acquired during summer. A March 2020 cohort acquired during COVID lockdowns will behave differently than March 2019. If the analysis doesn't account for this, the operator attributes seasonal or event-driven variance to the cohort itself.

Procedure: Compare cohorts within the same season or year. If comparing January 2024 to January 2023, note that external conditions (economy, competitor activity, product changes) may differ. If comparing January to July 2024, expect lower repeat rates in July due to seasonal purchase patterns, not cohort quality.

  • Group cohorts by season (Q1, Q2, Q3, Q4) if comparing across months
  • Note major events (product launch, pricing change, marketing shift) in the cohort period
  • Compare year-over-year cohorts (Jan 2024 vs. Jan 2023) to isolate seasonal effects

Mistake 4: Treating Correlation as Causation

If a paid-traffic cohort has 25% repeat rate and an organic cohort has 15%, the operator concludes 'paid traffic is higher quality.' But paid traffic may correlate with higher AOV, which independently drives repeat rate. Or paid traffic may skew toward existing customers (warm audience), not new acquisition. The cohort table shows correlation, not the causal mechanism.

Validation: Build a second cohort table controlling for the confounding variable. Split paid traffic by AOV (high vs. low), then compare repeat rates. If high-AOV paid and high-AOV organic have similar repeat rates, AOV is the driver, not channel.

  • Cohort tables show association, not causation
  • Before acting on a cohort insight, identify the mechanism (AOV, repeat frequency, product mix, customer segment)
  • Test the mechanism with a second cohort split

Mistake 5: Cohort Windows Too Short or Too Long

A 4-week cohort window (measuring repeat purchases in the first 4 weeks post-acquisition) is too short for most DTC brands - most repeat purchases happen in weeks 6 - 12. A 52-week window is too long - external events (new product, price change) will confound the signal. The window must match the repeat purchase cycle of the category.

For apparel and home goods, use 12 - 16 weeks. For consumables and supplements, use 8 - 12 weeks. For luxury goods, use 24 - 52 weeks. If the repeat purchase cycle is unknown, analyze the distribution of time between first and second purchase - use the 75th percentile as the cohort window.

  • Calculate median and 75th percentile time-to-repeat for the store
  • Set cohort window to 75th percentile (or 1.5x median)
  • If repeat cycle is <4 weeks, use weekly cohorts; if >12 weeks, use monthly cohorts

Mistake 6: Not Accounting for Survivorship Bias

If a cohort acquired in January includes 500 customers, but 50 have been refunded or churned by week 8, the repeat purchase rate is calculated on 450 survivors, not 500. This inflates the repeat rate and hides the fact that 10% of the cohort failed to survive. Survivorship bias is especially dangerous when comparing cohorts - if one cohort has higher churn before week 8, its repeat rate will appear artificially high.

Fix: Report repeat purchase rate as a percentage of the original cohort size, not survivors. Or build a separate churn/refund cohort table to track attrition independently.

  • Define the denominator: original cohort size or survivors at week N
  • If using survivors, document the attrition rate
  • Compare repeat rate across cohorts using the same denominator definition

Questions

FAQ

How do I know if my cohort analysis is valid?

Check three things: (1) Each cohort has at least 100 customers. (2) The cohort is defined by one dimension only. (3) The repeat purchase window matches the actual repeat cycle of the category. If all three are true, the table is valid. If repeat rates vary by >5 percentage points across cohorts, the difference is likely real (assuming no external events).

Should I build cohorts by acquisition date or acquisition channel?

Both, in separate tables. Acquisition date cohorts reveal seasonality and long-term trends. Acquisition channel cohorts reveal which channels drive higher-quality customers. Don't mix them in one table - it confounds the analysis. Start with acquisition date to establish baseline retention, then split by channel to compare.

What if my store is too small for cohort analysis?

If monthly cohorts have <100 customers, cohort analysis will be noisy. Instead, track repeat purchase rate and LTV as a single metric (not split by cohort). Once the store reaches 500+ customers/month, build monthly cohorts. Until then, focus on total repeat rate and segment by channel only if you have 200+ customers per channel per month.

Can I use cohort analysis to test the impact of a product launch or pricing change?

No - cohort analysis is observational, not causal. A cohort acquired after a price increase will have different repeat rates, but you can't isolate the price effect from other changes (seasonality, marketing spend, product mix). Use an A/B test or holdout group instead. Cohorts are useful for understanding long-term trends, not short-term interventions.

Want this on your account?

Thirty minutes. Bring the number that keeps you up.

More from the blog