Email Holdout Testing: How to Prove Your Retention Channel Is Actually Incremental

An email holdout test is the only way to measure how much revenue your retention program actually creates versus how much it takes credit for. Your ESP's attribution counts purchases within a set window after an open or click — including purchases that would have happened anyway. A holdout test reveals the real difference by withholding emails from a random group and comparing their purchase behavior to the group that received them.
Most DTC brands report that email drives a substantial chunk of revenue. That number comes from their ESP dashboard, and it is almost certainly inflated. Not because email isn't working — in our experience, it almost always is. But because attribution and incrementality are fundamentally different measurements, and most brands have never run the test that separates the two.
This guide covers how to build holdout tests in Klaviyo (which has no native feature for this), what each flow type requires, how to avoid the mistakes that invalidate most tests, and how to present the results so your CFO stops questioning whether retention is worth the investment.
Last updated: July 2026
What Is an Email Holdout Test and Why Does It Matter?
An email holdout test measures your retention program's true incremental impact by randomly splitting your audience into two groups: one that receives your emails and one that does not. The difference in purchase behavior between these groups — not your ESP's attributed revenue — reveals how much revenue your email program actually generates.
Holdout test is an experimental methodology borrowed from clinical trials: you withhold a treatment (email) from a randomly selected control group and compare their behavior to the group that received the treatment. Incrementality is the metric this test produces — the measurable lift in revenue, conversions, or engagement caused by your emails that would not have occurred without them.
Here is why this matters: Klaviyo defaults to last-touch attribution with a 5-day click and 5-day open window, according to Blossom's benchmark data. That means if a subscriber opens your email on Monday and purchases on Thursday, email gets full credit — even if that customer visits your site directly every Thursday to reorder. The attribution is technically correct (they did open the email), but the revenue is not truly incremental.
The real question is not "does email drive revenue?" — it is "how much revenue would we lose if we stopped sending?" A holdout test is the only way to answer that question with data instead of assumptions.
Revenue attribution is the system your ESP uses to assign credit for purchases to specific emails or flows. It tells you which emails preceded a purchase. Holdout testing tells you which purchases would not have happened without the email. Both measurements are useful. Only one measures your program's true value.
Holdout testing is also distinct from the broader attribution strategy a mature retention program needs. Attribution helps you compare channels against each other. Holdout testing proves whether a specific channel is worth running at all.
How Is a Holdout Test Different From an A/B Test?
An A/B test optimizes elements within your email program — subject lines, send times, offers — by comparing two versions of the same email. A holdout test answers a fundamentally different question: does the email itself drive behavior that would not have happened otherwise? A/B tests make your emails better. Holdout tests prove they work.
Think of it this way: an A/B test compares Email Version A against Email Version B. Both groups receive an email. A holdout test compares "received the email" against "received nothing." One group is held out entirely.
- A/B test question: "Does a subject line about free shipping outperform one about the product benefit?"
- Holdout test question: "Does sending this welcome flow at all cause more first purchases than not sending it?"
The practical difference matters for how you design each test. A/B testing is an optimization methodology that compares two variants of the same treatment to find the better performer. A/B tests can run on small splits because you are detecting the difference between two similar things. Holdout tests need enough people in the control group to detect the difference between something and nothing — a wider expected gap, but one that still requires statistical rigor to confirm it is real and not noise.
You need both in your measurement toolkit. A/B tests optimize what you send. Holdout tests prove that sending it matters.
How Do You Set Up a Holdout Test in Klaviyo?
Klaviyo has no native holdout test feature, but you can build one using a combination of custom profile properties, conditional splits within flows, and segment-based exclusions for campaigns. The key is tagging holdout group membership on each profile so it persists across all touchpoints and measurement periods.
Here is the step-by-step process for setting up a flow-level holdout test:
- Create a holdout group property. Add a custom profile property — something like "holdout_welcome_flow" — with a value of "true" or "false." This property determines which group each subscriber belongs to and ensures group membership stays consistent throughout the test.
- Assign profiles to the holdout group. Use Klaviyo's random sample segment to identify the profiles that will be held out. Create the segment using Klaviyo's sampling feature, then bulk-update those profiles with your holdout property set to "true."
- Add a conditional split to the flow. At the very beginning of the flow you are testing — before any emails fire — add a conditional split. A conditional split is a branching element in Klaviyo flows that routes profiles down different paths based on conditions you define, such as profile properties, event data, or random sampling (see Klaviyo's flow documentation for setup details). Profiles with the holdout property set to "true" take the yes branch, which leads directly to a flow exit with no emails. Everyone else continues through the flow normally.
- Build your measurement segments. Create two segments: one for profiles who entered the flow during the test period with the holdout property set to "true" (your control group), and one for profiles who entered during the same period with the property set to "false" (your test group).
- Define your measurement window. Decide in advance how long after flow entry you will measure purchase behavior. For most flows, you want to measure revenue generated within a defined window after the trigger event — not indefinitely.
- Pull results from Shopify, not Klaviyo. Since the holdout group never receives an email, Klaviyo cannot attribute revenue to them. Compare total purchase revenue between your two segments using your ecommerce platform's order data.
For Campaign Holdouts
Campaign-level holdout tests work differently because campaigns are one-time sends, not triggered flows:
- Create a holdout segment using a persistent profile property (same approach as above).
- Exclude the holdout segment from every campaign send during the test period.
- At the end of the test period, compare revenue per profile between the group that received campaigns and the group that did not.
Campaign holdouts are harder to run cleanly because you must exclude the holdout group from every single campaign. One missed exclusion contaminates the test. Consider running campaign holdouts separately from flow holdouts to keep the variables clean.
How Big Should Your Holdout Group Be?
Your holdout group needs to be large enough to produce statistically meaningful results but small enough that you are not suppressing too much revenue during the test. The right size depends on your total flow volume and how large a revenue difference you expect the test to detect.
Sample size is the number of profiles in each test group. Statistical significance is the confidence level that your results reflect a real difference rather than random variation — the threshold that lets you make decisions on the data without second-guessing whether the gap was just noise.
Blossom's testing methodology requires a minimum of 1,000 recipients per variant for flow tests, according to Blossom's benchmark data. That means your holdout group needs at least 1,000 profiles entering the flow during the test period to produce readable results.
Here is the tradeoff you are managing:
- Holdout too small: You will not have enough data to detect a meaningful difference between groups. The test runs for weeks and produces inconclusive results — wasted time and wasted opportunity.
- Holdout too large: You are suppressing emails from a significant portion of your audience, which means you are deliberately leaving revenue on the table during the test period. For high-performing flows, this cost adds up quickly.
Blossom's approach is to suppress roughly one in ten flow entrants as a holdout, according to Blossom's benchmark data. This keeps the holdout small enough to limit revenue impact while producing enough volume for reliable measurement — provided your flow sees sufficient traffic. If your flow does not generate 1,000 holdout entries within a reasonable test window, you likely need to extend the test duration or reconsider whether that flow has enough volume to test at all. Use your email marketing KPIs to estimate whether a given flow has the traffic to support a holdout.
How Long Should You Run a Holdout Test by Flow Type?
Different flows require different test durations because the underlying customer behavior cycles differ. A welcome flow holdout can produce reliable data faster than a winback holdout because new subscriber volume is continuous and the conversion event is well-defined. Winback holdouts take longer because re-engagement cycles are inherently slow.
Blossom's testing methodology requires a minimum 30-day runtime for any flow test, according to Blossom's benchmark data. But some flows need more. Here is how we approach each flow type in our retention programs:
Welcome Flow Holdout
- Why it is testable first: New subscriber volume is continuous (driven by your popup and acquisition), so the holdout group fills steadily. The conversion event — first purchase — is clearly defined and typically happens within a few weeks of signup.
- Duration: The standard 30-day minimum from Blossom's benchmark data usually produces reliable results here, provided your list growth supports the minimum sample size.
- Primary metric: Revenue per subscriber in each group, measured within a consistent window after signup.
- Watch for: Seasonal subscriber quality shifts. Subscribers acquired during a sale behave differently than those acquired during normal periods — run the test during a representative window.
Cart Abandonment Flow Holdout
- Why it is testable: Cart abandonment events happen frequently for most ecommerce brands, generating steady volume for both groups.
- Duration: The 30-day minimum is usually sufficient according to Blossom's benchmark data. Cart recovery typically happens within days of abandonment, so you do not need an extended measurement window.
- Primary metric: Cart recovery rate and revenue per abandoner in each group.
- Watch for: Subscribers who abandon multiple carts during the test period. Decide in advance whether you are counting unique abandoners or total abandonment events.
Winback Flow Holdout
- Why it is harder: Winback flows target lapsed customers, and re-engagement happens slowly. The flow itself may run for several weeks, and the purchase decision cycle for a lapsed customer is longer than for an active one.
- Duration: In our experience, winback holdouts need significantly longer than the standard minimum to produce meaningful results — the re-engagement cycle simply takes more time. Plan accordingly and resist the temptation to read results early.
- Primary metric: Reactivation rate (did they purchase again?) and incremental revenue per lapsed customer.
- Watch for: Lapsed customers who return organically through other channels. This is exactly what the holdout test is designed to detect — but it means you need patience.
Browse Abandonment Flow Holdout
- Why it requires volume: Browse abandonment has the lowest intent of any triggered flow. The incremental lift tends to be smaller, which means you need more data to detect it reliably.
- Duration: Plan for at least the 30-day minimum according to Blossom's benchmark data, and potentially longer for brands with lower site traffic.
- Primary metric: Revenue per browser in each group.
- Watch for: Overlap with cart abandonment. If a subscriber browses, then adds to cart, they may enter both flows. Your holdout design needs to account for this interaction.
Campaign Holdout
- Why it is different: Campaign holdouts measure the incremental value of your entire campaign program, not a single triggered sequence. The holdout group receives no campaigns but still receives triggered flows.
- Duration: Campaign holdouts need to run long enough to capture multiple campaign types and cadences. A single week will not tell you anything meaningful because it only captures a few sends. In our experience, a representative test window needs to span enough sends to reflect your normal campaign mix.
- Primary metric: Total revenue per subscriber in each group over the test period.
- Watch for: Missed exclusions. Campaign holdouts are the most operationally demanding because you must remember to exclude the holdout group from every send. One missed exclusion contaminates your data.
What Mistakes Invalidate Most Email Holdout Tests?
Most holdout tests fail not because the methodology is wrong but because of implementation mistakes that contaminate the results. These five failure modes are specific to ecommerce retention programs, and each one is preventable if you design for it upfront.
Cross-channel contamination. Your holdout group is not receiving emails, but they are still getting SMS messages, seeing retargeting ads, and receiving push notifications. If these other channels compensate for the missing emails, your holdout test will understate email's true incremental value. The fix: Decide upfront whether you are testing "email alone" or "email within the full channel mix." If you want a clean email-only measurement, consider suppressing the holdout group from SMS as well — but know this increases the revenue cost of running the test.
Insufficient test duration for slow-cycle flows. Running a winback holdout for a couple of weeks and drawing conclusions is like judging a diet after three days. Re-engagement takes time, and short tests produce unreliable data. The fix: Match your test duration to the behavior cycle of the flow you are testing. When in doubt, run longer rather than shorter.
Measuring proxy metrics instead of revenue. Open rates and click rates tell you whether people interact with your emails. They tell you nothing about incrementality. A holdout group that never receives emails cannot generate opens or clicks — the only comparable metric between groups is purchase behavior. The fix: Define revenue as your primary metric before the test starts. Track revenue per recipient (RPR) — the average revenue generated per profile in each group — using your ecommerce platform's order data, not your ESP's attribution reports.
Holdout group too small to detect the expected lift. If your holdout group has a few hundred people and the true incremental lift is modest, the difference between groups may be real but statistically undetectable. You end up with noise instead of signal. The fix: Calculate your minimum required sample size before starting the test. Blossom's testing methodology requires a minimum of 1,000 recipients per variant, according to Blossom's benchmark data.
Running tests across major promotional periods. A holdout test that spans Black Friday will be distorted by promotional behavior that does not reflect normal operations. The test group benefits from BFCM emails while the holdout group misses them — but both groups are influenced by the site-wide sale, heavier advertising, and seasonal purchase intent. The fix: Either avoid testing during major promotional windows entirely, or run the test long enough that the promotional period is a small fraction of the total measurement window.
How Do You Present Holdout Results to Your CFO?
The entire value of proving incrementality is making a business case for retention investment — and that means translating holdout data into three numbers your CFO already thinks in: incremental revenue, incremental cost per acquisition, and true channel return on investment.
Here is the framework we use when communicating holdout results to finance teams. The methodology aligns with how leading measurement frameworks approach incrementality across all digital channels — email is no exception.
The Three Numbers That Matter
- Incremental revenue: The revenue difference between the test group and the holdout group, normalized per recipient. If subscribers who received your welcome flow generated more revenue per person than those who did not, that per-person difference multiplied by your total flow volume is your incremental revenue. This is the number that tells your CFO "this is what we would lose if we turned this off."
- Incremental CPA: Your total email channel cost (ESP fees, agency fees, production costs) divided by the number of conversions that were truly incremental — not all conversions attributed to email, just the ones that would not have happened without it. This number is almost always stronger than your paid acquisition CPA, which makes email look compelling in a portfolio context.
- True channel ROI: Incremental revenue divided by total channel cost. This is the honest return on your retention investment, stripped of attribution inflation. It will be lower than what your ESP dashboard shows — and that is the point. A defensible number earns more budget than an inflated one nobody trusts.
Here is why this framing matters: a flow might show strong attributed revenue in Klaviyo, but when you run the holdout test, the incremental number is lower. That is normal and expected. Many of those attributed purchases would have happened anyway — the customer was already going to buy, and the email happened to land within the attribution window.
The defensible incremental number is strategically stronger than the inflated attributed number. When your CFO trusts the measurement, the conversation shifts from "prove email is worth it" to "how do we invest more in the channel with the strongest incremental ROI?"
Report holdout results alongside your retention marketing dashboard metrics. Attributed revenue tells you which emails perform well relative to each other. Incremental revenue tells you whether the program as a whole justifies the investment. You need both — and the holdout test is what connects them.
When presenting to finance, lead with customer lifetime value impact. Holdout tests on post-purchase and replenishment flows often reveal that email's biggest incremental contribution is not the first conversion — it is the second and third purchases that would not have happened without ongoing retention messaging.
The Measurement That Changes the Conversation
Holdout testing almost always confirms what retention marketers already believe: email is incremental. It drives revenue that would not exist without it. But the real number is typically lower than what your ESP attributes, and knowing the exact figure is far more powerful than defending a number everyone suspects is inflated.
Start with your highest-volume, highest-revenue flow — usually welcome or cart abandonment. Set up the holdout using the Klaviyo methodology above. Run it for the minimum duration your flow type requires. And when you have results, present them using the incremental revenue framework, not open rates and attributed revenue.
The brands that measure incrementality do not just prove email works. They know exactly how well it works, which flows contribute the most, and where to invest next. That precision is what separates a retention program from a sending schedule.
Frequently Asked Questions
These are the most common questions brands ask when planning their first email holdout test — covering setup mechanics in Klaviyo, holdout group sizing, the difference between holdout and A/B testing, and how to measure and report incremental revenue results to finance stakeholders and leadership teams.
What is a holdout test in email marketing?
A holdout test randomly withholds emails from a subset of your audience (the control group) and compares their purchase behavior to the group that received emails. The difference reveals how much revenue your email program truly generates versus how much it simply takes credit for through attribution. Unlike A/B tests that compare two versions of an email, holdout tests compare receiving email against receiving nothing — making them the standard for proving channel-level incrementality.
How do you set up a holdout test in Klaviyo?
You set up a holdout test in Klaviyo using custom profile properties and conditional splits, since Klaviyo has no native holdout feature. Tag a random sample of profiles with a holdout property, add a conditional split at the top of your flow that routes holdout profiles to a flow exit, then measure purchase behavior for both groups using your ecommerce platform's order data rather than Klaviyo's attribution reports.
How big should a holdout group be for email?
Your holdout group needs a minimum of 1,000 profiles entering the flow during the test period to produce statistically reliable results, according to Blossom's benchmark data. The standard approach is to hold out roughly one in ten flow entrants — large enough for meaningful data but small enough to limit the revenue you suppress during the test.
What is the difference between a holdout test and an A/B test?
An A/B test compares two versions of the same email to optimize elements like subject lines or offers — both groups receive an email. A holdout test compares "received email" against "received nothing" to measure whether the email itself drives incremental behavior. A/B tests improve your emails. Holdout tests prove they generate revenue that would not exist without them.
How do you measure email incrementality?
You measure email incrementality by comparing purchase behavior between your test group (received emails) and your holdout group (received nothing) over a defined measurement window. The key metric is revenue per recipient in each group, pulled from your ecommerce platform rather than your ESP. The difference is your incremental revenue. Report this as total incremental dollars, incremental CPA, and true channel ROI to communicate value in a language finance teams trust.
Get measurement frameworks and retention tactics like this in your inbox every week. Subscribe to our newsletter →
Need help implementing this?
Let us take the hassle of managing your email marketing channel off your hands. Book a strategy call with our team today and see how we can scale your revenue, customer retention, and lifetime value with tailored strategies. Click here to get started.
Curious about how your Klaviyo is performing?
We’ll audit your account for free. Discover hidden opportunities to boost your revenue, and find out what you’re doing right and what could be done better. Click here to claim your free Klaviyo audit.
Want to see how we’ve helped brands just like yours scale?
Check out our case studies and see the impact for yourself. Click here to explore.
Read Our Other Blogs

Klaviyo + Shopify Integration: The Complete Data Sync and Event Tracking Setup Guide



How to Repair Sender Reputation: The Recovery Playbook When Your Emails Land in Spam



Abandoned Cart Flow Strategy: Why Your Recovery Sequence Stops Working After Email One




Not Sure Where to Start?
Let's find the biggest retention opportunities in your business. Get a free Klaviyo audit or retention consultation.

























































































