Measure impact with Proof
Run a fixed-duration holdout experiment with verified member identities and independent action tracking.
Proof compares new members who receive the Streakfox experience with a randomly assigned control group. It measures completed actions, not clicks on the streak widget.
Before starting
- Publish your program and verify a server or webhook completion. Browser-only check-ins cannot supply an independent baseline.
- Supply a stable member hash and a short-lived, server-signed identity token. The WordPress plugin handles this. Anonymous or unverified members do not see the widget while Proof is active.
- Collect the counted action for every member, even when Streakfox is hidden.
- Finish configuring rewards and follow-up first. Changing the program, rewards, outbound webhook or project timezone/domain invalidates an active result.
Start an experiment
Open Proof for the selected program. Choose a control group of 10%, 20% or 50% and a total duration of 14, 28 or 42 days. Confirm the tracking prerequisites and start.
Members are assigned once when the widget first requests state with a verified identity. The assignment is stored and stays the same across devices. Members with earlier program activity and test members are excluded from enrolment and keep their existing experience. Members who never make another action remain in their assigned group’s denominator.
The control group receives no widget, earned reward reveal or Streakfox milestone/retention webhook during the experiment. Your own unrelated communications continue. Avoid treatment-only messages outside Streakfox that bypass this assignment.
Enrolment closes seven days before the scheduled end. New members arriving after enrolment closes receive the normal experience and are excluded from the experiment. Existing assignments stay in force through the end.
The outcome
A member returns if the server records this program’s action more than 24 hours and no more than seven days after assignment. A member is included in the observed denominator only after their complete seven-day window. Test events, widget interactions and impressions are excluded.
This is a seven-day return window, not the exact-day D7 metric in Insights. It is not a revenue or churn measurement.
Events must also reach Streakfox before the experiment ends. Late deliveries cannot rewrite the final result. Control members who complete actions still count toward normal active-member usage; assignment alone does not create billable activity.
During collection, Proof shows group counts and observed rates without an impact verdict. After the scheduled end, it shows the difference in percentage points and a 95% Newcombe–Wilson interval. An interval crossing zero means no clear difference. Negative results are displayed as negative.
At least 100 fully observed members in each group are required for a verdict. This is a reporting safeguard, not a power calculation: the effect size, baseline and traffic still determine how informative a result is. A 10% control group may be too small for a low-traffic program. Plan the duration and split before starting; do not restart repeatedly to chase a positive result.
An unexpectedly imbalanced group split, changed treatment or deleted member data invalidates the result. Ending early restores the normal experience but does not produce a final impact claim. Configure a new experiment to try again.
What the result can establish
The comparison estimates the effect of the defined Streakfox experience on enrolled new members in this program. It does not establish the effect on existing members, all visitors, other customers or revenue. Customer-side identity mistakes, missing action tracking or messages sent outside the experiment can still invalidate the interpretation.
Method: Newcombe’s comparison of intervals for independent proportions. Data quality: Microsoft’s sample-ratio mismatch guidance.