Start with the causal question

Campaign reports show rising clicks and purchase conversions. Yet “would these customers have bought without this message?” remains difficult to answer. A frequently clicked message and a message that creates additional revenue are not the same thing.

To estimate the effect created by a campaign, randomly divide customers with the same eligibility at the same time and compare operating that campaign with not operating it. Keep other operating conditions as similar as possible and measure purchases and revenue for everyone originally assigned. Comparing clickers with nonclickers, or last month with this month, does not replace that comparison.[1]

The quantity of interest is the incremental effect. Rules assigning revenue to campaigns and experiments estimating the difference from a no-campaign outcome answer different questions.

Identify the question answered by each dashboard metric

Separating metric roles exposes gaps in CRM reporting. This table organizes measurement questions; it is not a shared field specification for every tool.

MetricQuestionWhat it cannot answer alone
Send volumeHow many messages were processed for sending?Did customers actually receive them?
DeliveryHow far along this channel is delivery confirmed?Did a person read it?
Opens/clicksHow many recorded interactions occurred?Did those interactions cause purchases?
Conversions/attributed revenueWhich purchases and amounts match the event and attribution rules?How much would have occurred without messaging?
Incremental purchases/revenueHow did business outcomes differ from a comparable control?Will effects persist for other customers, seasons or longer periods?

Even “delivery” varies by channel. OneSignal push Delivered means handoff to the push provider, distinct from device-level Confirmed Receipt. Its reported CTR is Clicks divided by Delivered. Email opens depend on tracking-image loading and may not correspond to actual reading.[2][3]

“Unique” does not necessarily mean distinct customers. OneSignal email unique clicks/opens use the Subscription unit. If your business counts buyers by customer ID, first reconcile the units in messaging and order reports.[4]

Attribution windows and experimental observation use different clocks

OneSignal's current conversion reporting grants direct attribution to the latest eligible interaction within the channel's attribution window. Several messages can receive influenced credit for one conversion. That rule allocates revenue to messages; it does not calculate revenue without messaging. Adding influenced performance across messages does not yield net incremental performance.[5]

Attribution windows begin from the interaction defined by their rule, such as a click or open. Experimental observation instead begins from a time available to both groups, for example 14 days after assignment. The control group has no campaign click, so click date cannot be the shared start.

Include purchases during observation regardless of acquisition path, including direct visits without a message click. Conversely, absence of a traceable messaging touchpoint does not prove a purchase was completely uninfluenced. Maintain attribution and experiment reports separately under different names.

How the service works

A consistent rhythm for planning and verified sending — Plan the campaign calendar, verify audiences before sending and use experiment results in the next operating plan.

A different comparison answers a different question

Copy A versus copy B identifies a better expression within a campaign. If both groups receive messaging, it does not establish that the campaign is better than sending nothing. That question needs a holdout excluded from the campaign.

A control excluding only one repurchase campaign estimates the effect of adding that campaign to existing operations. A long-term holdout excluding all marketing asks a different question about the entire CRM program. Do not turn either into an experiment that withholds necessary order confirmations or security notices.

If the campaign adds a new discount coupon, the tested intervention combines messaging and a benefit. It is not equivalent to holding the discount constant and changing only the message. Define the tested combination first; do not change the explanation after seeing results.

OneSignal documentation separately describes messaging A/B tests and randomly tagging users to form a control excluded from sending. The existence of A/B testing does not automatically complete incremental-revenue analysis. Define where control purchases are collected and the unit of analysis separately.[6]

A design for testing one repurchase reminder

This is a hypothetical design example. Thirty days after purchase, 1:1 assignment and 14-/28-day windows are explanatory assumptions, not industry averages or recommended defaults. Determine actual sample size and duration from purchase cycles, metric variability and the minimum effect that matters.

Design itemExample decision
Measurement questionDoes adding one repurchase reminder to existing operations increase net revenue per assigned customer?
Eligible customersThirty days since last purchase; no repurchase at assignment; verified customer ID, channel consent and reachability
Assignment unitCustomer ID; keep one customer's app/browser/email subscriptions in the same group and assign once
RandomizationRandom 1:1 split of the same eligible population; retain experiment ID, assignment time and group; do not rerandomize midway
Treatment/controlTreatment receives one reminder without an additional discount; control omits only this campaign. Other campaign, price and necessary-notification rules stay the same
ObservationInclude orders within 14 days after each customer's assignment; reflect their refunds through day 28. Finalize after the last customer's observation and the agreed collection-latency allowance end
Primary metricNet revenue per assigned customer: actual discounted merchandise payments less refunds, excluding taxes/shipping consistently
Secondary/guardrail metricsBuyer proportion, sending cost per assigned customer, new opt-outs and complaints; separately inspect refund size/timing
AnalysisRetain all originally assigned customers in each denominator; report mean difference and uncertainty. Fix primary metric, decision time and minimum scale-up effect before observing results
Reasons to defer conclusionsControl exposed to the same campaign; customers in both groups; order events missing on one side; unresolved allocation-ratio anomaly; incomplete observation/refund windows

“Net revenue” here is an experimental metric definition. Do not subtract a discount twice when using already-discounted payments. Measure messaging or separate gift costs outside this metric. Whether a modest revenue gain justifies additional cost and customer inconvenience is a separate business decision.

A purchase increase within 14 days may merely pull next month's purchases forward. If long-term repurchase or renewal matters, plan follow-up observation around that cycle from the outset. This short window is not evidence of long-term revenue impact.

If employees purchase through one organizational account or share benefits, randomizing organizations may be more appropriate than individuals. Analysis must then account for correlated outcomes within organizations. Treating all clicks/orders within randomized organizations as independent samples misstates uncertainty.[7]

Restricting analysis to clickers changes the comparison

Comparing only the 100 clickers among 1,000 treatment customers with the control no longer preserves randomization. It reselects customers by a post-assignment behavior. Excluding failed sends or nonreaders only from treatment creates the same problem.

This design defaults to comparing everyone in the originally assigned groups, an assignment-based analysis. Delivery and clicks become supporting metrics explaining how the intervention operated. Even correct randomization can be undermined by asymmetric outcome collection or analytical inclusion.[1]

Check these three areas jointly across sending and analysis settings.

One customer across devices and channels. OneSignal Journeys describes connecting subscriptions to one user through External ID. Without that connection, subscriptions may represent separate users, and one user can enter several Journeys. Apply the holdout to push and alternative sending paths for the same campaign, not just one email list.[8]

Receiving states and exceptions. Honor an opt-out by stopping subsequent sends. Do not silently remove dissatisfied customers from analysis to improve results. If withdrawal of required tracking permission prevents further observation, classify the outcome as missing. Check dashboard semantics too: OneSignal push Unsubscribed includes subscriptions unreachable at sending time, so it is not automatically “new customers who opted out because of this campaign.”[2]

Order data and other campaigns. Use the same order source and customer mapping for both groups, deduplicating by order ID. A fully observed customer with no order counts as zero; identity failures and collection outages do not become zero revenue. Apply the same eligibility/sending rules for other campaigns or design to separate interference. Record operating conditions if effects include interactions with another campaign.

Group sizes statistically inconsistent with the planned assignment or analytical ratio can signal sample ratio mismatch, or SRM. A difference of a few customers is not automatically failure, but investigate the cause before interpreting performance.[9]

How the service works

From customer data to campaigns ready to build — Review data and measurement, then define key campaigns and an implementation sequence around business goals.

What can a difference between 12% and 10% establish?

These are synthetic calculation data, not a real customer experiment. They illustrate conversion-rate interpretation, not verification of the preceding design's primary net-revenue metric.

ItemTreatmentControl
Originally assigned unique customers1,0001,000
Customers purchasing at least once in the specified period120100
Buyer proportion12.0%10.0%

The observed difference is 12% − 10% = 2 percentage points. Relative to the control rate, 2 ÷ 10 = 20%. Two percentage points and a relative 20% describe the same data in different units, not two separate achievements.

Assume independent customers, accurately observed purchases and one prespecified comparison at a fixed endpoint. The normal-approximation standard error and 95% confidence interval for the difference are as follows.[7]

Standard error = √(0.12 × 0.88 ÷ 1,000 + 0.10 × 0.90 ÷ 1,000)
               ≈ 0.0139857

95% confidence interval = 0.02 ± 1.96 × 0.0139857
                        ≈ −0.007412 to +0.047412
                        = −0.74 to +4.74 percentage points

The point estimate favors improvement, but the interval includes zero. This calculation neither establishes conversion improvement nor proves no effect. In particular, it cannot support “revenue increased 20%”: customer-level amounts incorporating order value and refunds are absent.

Do not apply this approximation unchanged to rare conversions, organization-randomized experiments or samples counting a customer repeatedly. A few high-value orders can greatly change revenue uncertainty even with identical purchase counts. Analyze net revenue separately using customer-level amount distributions.[7]

There is no rule that 1,000 per group is sufficient. First specify the smallest decision-relevant difference, metric variability and required power to detect it. Long purchase cycles require outcome-maturation time as well as recruitment time.[10]

Checking results and declaring success are different activities

Monitor incorrect sends, missing events and complaints during the experiment. But checking significance daily under ordinary fixed-horizon analysis and stopping on the first favorable day no longer preserves the original statistical decision properties. Prespecify suitable sequential analysis and stopping rules if interim results will determine termination. Stop operations separately to prevent harm; do not present early-stopped results as definitive revenue success.[11]

Avoid inspecting opens, clicks, purchases, revenue and many customer groups, then selecting only the best-looking result as the headline. Prespecify the primary metric and label exploratory findings separately. Microsoft's recent experimentation material also discusses accounting for metric correlation and relevance.[12]

Reporting conclusions can differ as below. These are examples of decision statements, not observed business results.

ObservationSupported report
Copy A/B test improved clicks, but no no-send control or purchase observation exists“Message interaction improved. Incremental revenue remains unverified.”
Appropriate control and complete observation; primary effect and uncertainty meet prespecified criteria; guardrails acceptable“Evidence supports an incremental effect for this audience, period and operating conditions.”
Estimated range includes both loss and meaningful benefit, or the necessary purchase cycle is incomplete“Current data cannot support a scale-up decision. Further observation or experimentation is needed.”

If customer IDs, orders for both groups and consistent exclusion rules are already managed, organize the comparison before buying a new analytics tool. If channels identify customers differently or organizational effects are computed from individual logs, repair the measurement structure first.

Under click-through rate in the next report, include the compared groups, the primary effect and its uncertainty, and what remains unresolved. Record the design in the CRM campaign measurement and experiment workbook. For external support defining metrics and necessary events, see IXC CRM consulting scope.