Experiment Runbook — Design, Run, Measure

A practical, repeatable runbook for designing, running, and measuring low-cost business experiments that produce clear yes/no decisions. Includes hypothesis templates, metric guidance, sample-size and duration heuristics, step-by-step run instructions, data collection template, strict decision rules, and three ready-to-use example experiments.

Purpose

This runbook helps one-person and small teams design low-cost experiments that produce clear decisions: iterate, pivot, or scale. It emphasizes simple, repeatable structure, strict decision rules to avoid noisy results, and practical heuristics you can use without a statistics degree.

Core principles

  • Pre-register the hypothesis, primary metric, minimum sample/duration, and decision rule before you launch.
  • One change at a time. Test a single variable or a clearly bundled set that you will treat as one change.
  • Primary metric first. Pick one primary metric that answers the hunger behind the experiment; collect secondary metrics for context.
  • Qualitative + quantitative. Use short interviews, recordings, or surveys to explain the numbers.
  • Stop rules. Predefine what counts as success, failure, or inconclusive and how you’ll act in each case.

Hypothesis template

Write a concise, testable hypothesis using this format:

If we make this specific change, then primary metric will increase/decrease by expected amount or direction within timeframe because reason.

Example: If we show a simplified pricing table with a highlighted recommended plan, then trial signups will increase by 20% within two weeks because visitors will understand the offer faster.

Choosing a primary metric

Pick one metric that directly answers the experiment question. Good primary metrics are observable, measurable, and aligned with business value.

  • Signups per unique visitor (conversion rate) for landing-page messaging experiments.
  • Purchase rate, revenue per visitor, or average order value for pricing tests.
  • Trial-to-paid conversion for product packaging or onboarding changes.

Track supporting metrics (traffic, click-throughs, churn, NPS) but avoid mixing them into the primary decision.

Sample size & duration guidance (practical heuristics)

Formal power calculations are ideal, but for quick, low-cost tests use these conservative heuristics to reduce noisy decisions:

  • Aim for at least 7 days to cover weekday/weekend patterns; 14 days is safer for low-traffic sites.
  • For experiments that rely on conversions (purchases, signups): aim for at least 50–200 conversions per variant when possible. If conversions are rare, prefer longer runs or measure nearer-term behaviors (clicks or form starts) that occur more frequently.
  • When traffic is limited, prefer sequential small tests with qualitative follow-up rather than single large A/B tests. Run a pilot, collect qualitative feedback, then iterate.
  • Don't trust results from tiny samples. If you get an early apparent win with fewer than ~30 conversions, treat it as a directional signal, not a decision.

If you want to be more rigorous later, connect the experiment plan to a simple sample-size calculator or ask an analyst to compute minimum detectable effect (MDE).

Experiment steps (actionable sequence)

  1. Write the pre-registration: hypothesis, primary metric, secondary metrics, minimum sample or minimum duration, assignment method, and clear decision rules.
  2. Setup tracking: ensure analytics capture unique visitors, events, and the primary metric. Tag variants or cohorts clearly.
  3. Run a small sanity check: preview pages, test events, and run internally for a day to validate instrumentation.
  4. Launch: start the experiment and note the start time and expected minimum end date.
  5. Collect qualitative signals: chat with early users, add a brief survey, or record sessions to explain why numbers move.
  6. Do not peek frequently: set a schedule (e.g., check once every 48–72 hours) unless you pre-specified early stopping rules for strong effects or harms.
  7. At minimum end: collect final data, check for instrumentation issues, and compare outcome against decision rules.
  8. Act: scale the winner, iterate on the variation, or pivot away based on the pre-registered decision.

Data collection template

Record these fields for every experiment run (use the platform’s interactive form when available):

  • Experiment name and ID
  • Owner(s)
  • Hypothesis (pre-registered)
  • Primary metric and how it’s measured (event name, filter)
  • Secondary metrics
  • Variant descriptions (control and test)
  • Assignment method (random, geo, cohort)
  • Start date and expected minimum end date
  • Minimum sample or conversions required
  • Decision rule (exact thresholds for success/failure/inconclusive)
  • Notes on qualitative feedback
  • Result summary and next action (iterate, pivot, scale)

Decision rules (iterate / pivot / scale)

Predefine thresholds and act on them.

  • Scale: The variant exceeds the control by the pre-registered threshold (e.g., relative lift or business impact threshold) and you reached the minimum sample/duration. You will roll out the change broadly and monitor for regression.
  • Iterate: The variant shows a directional improvement but did not meet the threshold or the sample was marginal. Use qualitative feedback to redesign and run a follow-up experiment.
  • Pivot (stop): The variant fails to meet the threshold or harms the business (e.g., lower core conversions, negative feedback). Stop the change and consider different hypotheses.
  • Inconclusive: Conflicting signals, instrumentation issues, or insufficient sample. Extend duration or redesign the test.

Record the chosen action and the rationale in the experiment log.

Examples

Pricing experiment (value-based packaging)

Hypothesis: If we introduce a mid-tier plan that bundles the most requested features, then trial-to-paid conversion will increase by 15% over four weeks because customers see clearer value.

Primary metric: trial-to-paid conversion rate. Minimum: 150 trials per variant or 28 days. Decision: >=15% relative lift → scale; 5–15% → iterate with price/positioning tweaks; <5% → pivot.

Landing-page messaging (headline + CTA)

Hypothesis: A benefit-first headline + simplified CTA increases signup rate by 25% in two weeks because it reduces friction.

Primary metric: signup rate per unique visitor. Minimum: 500 visitors per variant or 14 days. Use session recordings and a 2-question on-exit survey for qualitative context.

Service packaging (pilot cohort)

Hypothesis: Offering a 4-week guided pilot (limited seats) will increase paid onboarding by creating urgency and reducing perceived risk.

Primary metric: paid onboarding conversions from the pilot pool. Minimum: 20 pilot signups; decision based on conversion rate and qualitative client satisfaction.

Avoid common mistakes

  • Changing multiple elements and then claiming a single cause.
  • Using vanity metrics (pageviews, social likes) as primary decision metrics.
  • Peeking and stopping on weak early signals without pre-registered rules.
  • Failing to validate tracking and attribution before drawing conclusions.
  • Not collecting qualitative explanations—numbers alone rarely tell you why something worked.

One-page quick checklist

  1. Write hypothesis & pre-register.
  2. Choose one primary metric and supporting metrics.
  3. Set minimum sample/duration and explicit decision rules.
  4. Implement and validate tracking.
  5. Launch and collect qualitative feedback.
  6. Hold to the pre-registered decision rule at the end.
  7. Document the result and action (iterate/pivot/scale).

Next steps and suggested tooling

For teams using this platform, consider converting this runbook into an interactive experiment form that pre-fills the data-collection template, stores submissions, and records results. Link the form to analytics or a simple agent that computes basic lift and flags instrumentation problems.

References & further reading

Keep a short library of experiment examples and post-mortems. Over time this becomes a valuable knowledge asset that increases your experimental leverage.


Discussion

Comments and conversation will live here.