Experiment & Causal Study Design Workbook

An interactive, guided workbook to pre-specify, document, and save experiment and causal study designs: problem framing, hypothesis, outcomes, power and sample planning, assignment and randomization, stopping rules and ethics, analysis plan, rollout/monitoring, and results + decision template.

Interactive Tool

Experiment & Causal Study Design Workbook

Use this guided workbook to pre-specify a defensible experiment or causal study that answers a practical decision question. Pre-specification reduces researcher degrees of freedom, avoids selective reporting, and speeds operational adoption. Fill each section with the best available information and save a stable plan you can reference during analysis and rollout.

Describe the operational problem, the decision that the experiment will inform, affected stakeholders, timing constraints, and why this question matters. Keep it concise (3–6 sentences).
Example: 'Roll out change to all customers', 'Delay launch', 'Increase inspection frequency'. Specify who decides and the success threshold.
Write a measurable hypothesis: 'If we [intervention], then [measurable outcome will change] because [mechanism].' Example: 'If we add a confirmation step, then error rate will drop by X% because people notice mistakes.'
List each arm clearly and what users/units experience in each. Include any implementation constraints, rollout windows, and dependencies.
Choose the level at which assignment will occur. This affects sample size and inference.
Describe how you'll assign units to arms (simple random, stratified, block randomization, geo-based, deterministic rule). Mention any blocking/stratification variables and implementation details.
Name the single primary metric that will decide success. Use a clear operational definition.
Define numerator/denominator, time window, data source, units, smoothing or aggregation, and any filters. Example: '30-day retention = % users with any activity between day 1 and day 30 post-signup, measured in production DB table X.'
List metrics to monitor for unintended harms (e.g., error rates, complaint volume, latency, cost). Define thresholds that would pause or stop the experiment.
Enter the best available baseline estimate (proportion or mean). If unknown, note how you'll estimate it.
Enter the smallest practical effect size (absolute or relative) that would change the decision. Be realistic: smaller MDEs require much larger sample sizes.
Common defaults: 0.05. Choose lower if multiple comparisons are expected.
Common default: 0.8 (80%). Higher power increases sample requirements.
Record the sample size per arm, calculation method, assumptions (baseline, MDE, alpha, power, variance), clustering or ICC if applicable, and the calculator or script used (link or file name). If you will run a pilot, describe it here.
Provide a link to a spreadsheet, script, or internal tool, or paste the calculation reference.
Specify if you'll do interim analyses, what triggers pausing or stopping (efficacy, futility, safety), who reviews interim results, and the statistical correction or alpha-spending plan. If no interim analyses, state that explicitly.
List potential risks to participants, data privacy controls, informed consent if required, regulatory/IRB review status, and mitigation steps. Note any vulnerable groups affected.
Choose the principal statistical test or model you'll use for the primary outcome. Specify link function, adjustments, and software or script locations.
List covariates for adjusted analyses and the exact subgroups you will test. For subgroups, specify if tests are exploratory and how you'll report multiple comparisons.
Describe how you'll handle missing outcomes (imputation, complete-case), data validation steps, and who owns data quality during the experiment.
State whether you'll adjust for multiple tests (Bonferroni, Holm, FDR, pre-specify primary endpoint only, etc.).
Explain how you'll deploy the experiment, monitor technical health, observe guardrail metrics, and who will receive alerts. Include dashboards or logs to watch.
List thresholds that trigger investigation (e.g., error rate > X%, latency > Y ms) and the responsible person or team.
When will you lock data and run the final analysis? Will results be communicated in a post-mortem or pre-registered report?
Use this template after analysis: observed effect and CI, p-value, direction relative to MDE, guardrail outcomes, robustness checks, recommended action (adopt / iterate / stop), owner and timeline.
Describe precise decision criteria (e.g., 'Adopt if estimated lift >= MDE and no guardrail exceeded; Iterate if effect positive but CI includes zero; Stop if negative effect or guardrail threshold exceeded').
List primary owner, analytics lead, engineering contacts, and governance reviewers. Include names, roles, and emails.
Record assumptions that could invalidate inference (noncompliance, interference, seasonality) and how you'll check them.
You can explore this tool now. Sign in or create an account to save your responses and return to them later.
Make this tool part of your work

Save a personal copy, bring it to your team, or tailor the questions and workflow to fit what you are hungry to improve.

Member customization and team collaboration are coming soon.

Discussion

Comments and conversation will live here.