Experimentation & Metrics Lab — Runbook

A practical runbook for designing, running, and interpreting small, low-cost business experiments. Includes a concise test design template (hypothesis, metric, sample target, duration), a ready-to-use tracking sheet, pragmatic statistical sanity checks for small samples, a decision guide (iterate/scale/stop), and two worked examples entrepreneurs can copy.

Purpose

This runbook helps you run small, focused experiments that produce clear directional learning without consuming a lot of time or budget. It’s built for entrepreneurs and one-person businesses that must decide quickly with limited data.

When to use this runbook

Use it when you want to test an assumption cheaply: a pricing change, a landing page, a new lead source, a service offer, or a small workflow change. The goal is not perfect statistical proof but repeatable, informative tests that reduce risk and guide action.

Quick Test Design Template

Fill these fields before you start. Treat them as a contract you can come back to when reviewing results.

  • Experiment name: concise, descriptive.
  • Hypothesis: If we [change X], then [primary metric] will [direction & expected magnitude].
  • Primary metric (one): the simplest measure that tells you whether the change helped (conversion rate, leads/day, revenue per visitor, average sale).
  • Secondary metrics: guardrails or signals (churn, customer satisfaction, support load).
  • Baseline: current value of the primary metric and recent variability.
  • Sample target: a practical target for observations or events (see guidance below).
  • Duration: calendar period you will run the test (days/weeks) and a fixed review date.
  • Success criteria: clear rules for Iterate / Scale / Stop (see decision guide).
  • Owner: who runs the test and who decides next steps.

Tracking sheet (copy this table)

ExperimentHypothesisPrimary metricBaselineSample targetStartEndActual sampleResultDecision
Example: Price testLower price -> more salesSales/day2/day50 orders2026-06-012026-06-1458+30% sales/dayIterate

Use a single spreadsheet or this content item’s future interactive form to capture each experiment. Keep the record brief but precise.

Practical sample & statistical sanity checks for small experiments

Small-sample experiments are noisy. These are pragmatic checks and rules of thumb to keep you honest.

  • Check counts first: For binary events (buys, signups) treat results with <20 events per variant as anecdotal. 20–100 events per variant are suggestive. >100 events gives stronger directional evidence. These are heuristics, not absolute rules.
  • Consider effect size, not just p-values: Ask: Is the observed change big enough to matter in your business? Even small percent changes can justify scaling if margin and volume make it profitable.
  • Look for consistent direction across repeats: If multiple small runs point the same way, confidence grows faster than a single noisy test.
  • Check for data quality issues: outliers, bot traffic, seasonal shifts, mis-tagged events, or changes in sample composition can create false signals.
  • Prefer practical rules: If your baseline conversion is very low (under 1%), you often need larger sample sizes to detect small relative changes. When in doubt, treat results as provisional and run a follow-up test designed to confirm the direction and magnitude.
  • Use simple analytics methods: For small counts, Fisher’s exact test or bootstrapped confidence intervals are often more reliable than large-sample z-tests. If you’re not a statistician, use the direction, magnitude, and repeatability as your primary guides.

Decision guide — Iterate / Scale / Stop

Use the pre-declared success criteria and the checks above to make a three-way decision.

  • Scale when: the test shows a clear, practically meaningful improvement in the primary metric, the effect is consistent across segments or repeats, secondary metrics (quality, margin) are acceptable, and the change is implementable at scale.
  • Iterate when: the result suggests a direction (improvement or problem) but the effect is small, sample is marginal, or there are quality questions. Adjust the design (bigger sample, narrower audience, different creative) and run a follow-up test with a clear confirmation plan.
  • Stop (and learn) when: there is no directional signal after a reasonably powered test, or the change harms a key secondary metric. Record the learning and move on—don’t waste cycles on marginal variations of losers.

Two short examples

1) Landing page headline change

Hypothesis: A clearer value-focused headline will lift signups by at least 20%. Primary metric: signup conversion rate. Baseline: 3% over last 2 weeks. Sample target: 1,200 visitors split (600 each) to have a reasonable chance of directional evidence. Duration: run for two weeks or until sample target met. Decision: if conversion goes up >15% and support tickets unchanged, iterate or scale; if noisy, repeat with larger sample or different segment.

2) Intro price experiment

Hypothesis: A limited-time lower price will increase purchase rate and overall revenue. Primary metric: revenue/day (not just conversion). Secondary: average order value, churn risk. Baseline revenue: $400/day. Sample target: 50 purchases across the test window. Decision: if revenue rises and AOV doesn’t fall too far, scale; otherwise iterate with an alternative offer.

Quick checklist before you start

  1. One clear hypothesis.
  2. One primary metric and 0–2 secondary metrics.
  3. Baseline measured and recorded.
  4. Sample target and fixed review date.
  5. Owner and decision rules documented.
  6. Instrumentation verified (events fire correctly).

How this runbook can be made interactive

This runbook is intentionally compact so it’s easy to follow. It would become more useful if paired with an interactive experiment form and a submission endpoint that stores outcomes. That would let you track experiments over time, auto-calculate basic statistics, and generate a dashboard of repeatability and outcomes.

Suggested platform features: a simple Experiment Intake form (template above), a results submission flow, and an experiment list/dashboard that highlights experiments ready for review and shows suggested decisions based on your declared criteria.

Final note

Small experiments are powerful when they’re deliberate and recorded. Focus on learning that changes what you do next. If a test is inconclusive, that’s also valuable information—treat it as a prompt to redesign and try again rather than as failure.


Discussion

Comments and conversation will live here.