Author
Abstract
A/B testing is central to digital marketing optimization, yet the speed of decision-making often outpaces statistical discipline. This paper studies two empirically measurable threats to experimental validity on public marketing data---peeking bias from continuous monitoring and false discovery inflation under multiple comparisons---together with a third threat, novelty-driven early lifts, which is evaluated by controlled injection of temporal decay rather than by direct observation. A systematic experimental framework is proposed that layers mixture sequential probability ratio testing, CUPED-style variance reduction using pre-treatment covariates (a generalization of the original CUPED design, which requires a same-metric pre-period that is not available for every outcome here), temporal stratification with injected-novelty stress tests, and stratified subgroup screening with Benjamini-Hochberg control. The protocol is illustrated on two public marketing datasets used as proxies for broader digital experimentation rather than as native paid-search logs: a widely used public copy of the Criteo Uplift dataset (13,979,592 rows) and the Hillstrom MineThatData email campaign dataset (64,000 customers, three arms). In 1,000 simulated A/A and A/B trials on the Criteo proxy data, the sequential-testing and variance-control layers reduce realized Type-I error from 27.4% under unadjusted repeated testing to 5.1%, while lowering average sample consumed to 68.4% of the fixed-horizon budget at 81.2% power. Separately, in subgroup-screening simulations on the Hillstrom proxy data, estimated discovery precision reaches 82.3% against the held-out benchmark defined in Section 4.1. Because public datasets do not expose paid-search auction dynamics or native day-by-day novelty traces, the empirical contribution should be read as a proxy-based demonstration of digital marketing experimentation more broadly, not as a paid-search-specific validation. Even within these limits, the framework shows that lightweight statistical safeguards can preserve decision speed without sacrificing inferential validity, supporting more reliable ship decisions and more efficient experimentation budget allocation in real-world campaign optimization.
Suggested Citation
Download full text from publisher
Corrections
All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:dba:jsisia:v:2:y:2026:i:3:p:57-68. See general information about how to correct material in RePEc.
If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.
We have no bibliographic references for this item. You can help adding them by using this form .
If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.
For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Joseph Clark (email available below). General contact details of provider: https://pinnaclepubs.com/index.php/JSISI .
Please note that corrections may take a couple of weeks to filter through
the various RePEc services.