Experiment Runbook — Design, Run, Measure
A practical, repeatable runbook for designing, running, and measuring low-cost business experiments that produce clear yes/no decisions. Includes hypothesis templates, metric guidance, sample-size and duration heuristics, step-by-step run instructions, data collection template, strict decision rules, and three ready-to-use example experiments.
Purpose
This runbook helps one-person and small teams design low-cost experiments that produce clear decisions: iterate, pivot, or scale. It emphasizes simple, repeatable structure, strict decision rules to avoid noisy results, and practical heuristics you can use without a statistics degree.
Core principles
- Pre-register the hypothesis, primary metric, minimum sample/duration, and decision rule before you launch.
- One change at a time. Test a single variable or a clearly bundled set that you will treat as one change.
- Primary metric first. Pick one primary metric that answers the hunger behind the experiment; collect secondary metrics for context.
- Qualitative + quantitative. Use short interviews, recordings, or surveys to explain the numbers.
- Stop rules. Predefine what counts as success, failure, or inconclusive and how you’ll act in each case.
Hypothesis template
Write a concise, testable hypothesis using this format:
If we make this specific change, then primary metric will increase/decrease by expected amount or direction within timeframe because reason.
Example: If we show a simplified pricing table with a highlighted recommended plan, then trial signups will increase by 20% within two weeks because visitors will understand the offer faster.
Choosing a primary metric
Pick one metric that directly answers the experiment question. Good primary metrics are observable, measurable, and aligned with business value.
- Signups per unique visitor (conversion rate) for landing-page messaging experiments.
- Purchase rate, revenue per visitor, or average order value for pricing tests.
- Trial-to-paid conversion for product packaging or onboarding changes.
Track supporting metrics (traffic, click-throughs, churn, NPS) but avoid mixing them into the primary decision.
Sample size & duration guidance (practical heuristics)
Formal power calculations are ideal, but for quick, low-cost tests use these conservative heuristics to reduce noisy decisions:
- Aim for at least 7 days to cover weekday/weekend patterns; 14 days is safer for low-traffic sites.
- For experiments that rely on conversions (purchases, signups): aim for at least 50–200 conversions per variant when possible. If conversions are rare, prefer longer runs or measure nearer-term behaviors (clicks or form starts) that occur more frequently.
- When traffic is limited, prefer sequential small tests with qualitative follow-up rather than single large A/B tests. Run a pilot, collect qualitative feedback, then iterate.
- Don't trust results from tiny samples. If you get an early apparent win with fewer than ~30 conversions, treat it as a directional signal, not a decision.
If you want to be more rigorous later, connect the experiment plan to a simple sample-size calculator or ask an analyst to compute minimum detectable effect (MDE).
Experiment steps (actionable sequence)
- Write the pre-registration: hypothesis, primary metric, secondary metrics, minimum sample or minimum duration, assignment method, and clear decision rules.
- Setup tracking: ensure analytics capture unique visitors, events, and the primary metric. Tag variants or cohorts clearly.
- Run a small sanity check: preview pages, test events, and run internally for a day to validate instrumentation.
- Launch: start the experiment and note the start time and expected minimum end date.
- Collect qualitative signals: chat with early users, add a brief survey, or record sessions to explain why numbers move.
- Do not peek frequently: set a schedule (e.g., check once every 48–72 hours) unless you pre-specified early stopping rules for strong effects or harms.
- At minimum end: collect final data, check for instrumentation issues, and compare outcome against decision rules.
- Act: scale the winner, iterate on the variation, or pivot away based on the pre-registered decision.
Data collection template
Record these fields for every experiment run (use the platform’s interactive form when available):
- Experiment name and ID
- Owner(s)
- Hypothesis (pre-registered)
- Primary metric and how it’s measured (event name, filter)
- Secondary metrics
- Variant descriptions (control and test)
- Assignment method (random, geo, cohort)
- Start date and expected minimum end date
- Minimum sample or conversions required
- Decision rule (exact thresholds for success/failure/inconclusive)
- Notes on qualitative feedback
- Result summary and next action (iterate, pivot, scale)
Decision rules (iterate / pivot / scale)
Predefine thresholds and act on them.
- Scale: The variant exceeds the control by the pre-registered threshold (e.g., relative lift or business impact threshold) and you reached the minimum sample/duration. You will roll out the change broadly and monitor for regression.
- Iterate: The variant shows a directional improvement but did not meet the threshold or the sample was marginal. Use qualitative feedback to redesign and run a follow-up experiment.
- Pivot (stop): The variant fails to meet the threshold or harms the business (e.g., lower core conversions, negative feedback). Stop the change and consider different hypotheses.
- Inconclusive: Conflicting signals, instrumentation issues, or insufficient sample. Extend duration or redesign the test.
Record the chosen action and the rationale in the experiment log.
Examples
Pricing experiment (value-based packaging)
Hypothesis: If we introduce a mid-tier plan that bundles the most requested features, then trial-to-paid conversion will increase by 15% over four weeks because customers see clearer value.
Primary metric: trial-to-paid conversion rate. Minimum: 150 trials per variant or 28 days. Decision: >=15% relative lift → scale; 5–15% → iterate with price/positioning tweaks; <5% → pivot.
Landing-page messaging (headline + CTA)
Hypothesis: A benefit-first headline + simplified CTA increases signup rate by 25% in two weeks because it reduces friction.
Primary metric: signup rate per unique visitor. Minimum: 500 visitors per variant or 14 days. Use session recordings and a 2-question on-exit survey for qualitative context.
Service packaging (pilot cohort)
Hypothesis: Offering a 4-week guided pilot (limited seats) will increase paid onboarding by creating urgency and reducing perceived risk.
Primary metric: paid onboarding conversions from the pilot pool. Minimum: 20 pilot signups; decision based on conversion rate and qualitative client satisfaction.
Avoid common mistakes
- Changing multiple elements and then claiming a single cause.
- Using vanity metrics (pageviews, social likes) as primary decision metrics.
- Peeking and stopping on weak early signals without pre-registered rules.
- Failing to validate tracking and attribution before drawing conclusions.
- Not collecting qualitative explanations—numbers alone rarely tell you why something worked.
One-page quick checklist
- Write hypothesis & pre-register.
- Choose one primary metric and supporting metrics.
- Set minimum sample/duration and explicit decision rules.
- Implement and validate tracking.
- Launch and collect qualitative feedback.
- Hold to the pre-registered decision rule at the end.
- Document the result and action (iterate/pivot/scale).
Next steps and suggested tooling
For teams using this platform, consider converting this runbook into an interactive experiment form that pre-fills the data-collection template, stores submissions, and records results. Link the form to analytics or a simple agent that computes basic lift and flags instrumentation problems.
References & further reading
Keep a short library of experiment examples and post-mortems. Over time this becomes a valuable knowledge asset that increases your experimental leverage.
Discussion
Comments and conversation will live here.