Menu A/B Test Plan
A practical, low-risk playbook to design, run, measure, and decide on menu experiments that validate price, placement, portion, or description changes—plus hypothesis templates, POS tagging steps, sample-size guidance, KPIs, decision rules, and rollout guidance.
Welcome — test smart, not fast
If you want menu changes that reliably improve revenue and margin without surprising guests, run controlled experiments before you roll them out. This playbook helps you design low-risk A/B tests for price, placement, portion, description, or bundling changes so you can make decisions from evidence instead of opinion.
When to run an A/B test
- You expect a measurable change in sales, check size, mix, or margin from a single clear change (price, placement, portion, description, or add-on).
- You can identify the test and control in sales data (POS tagging, modifier, or item code).
- You can run the change long enough to collect meaningful samples (see sample size guidance).
- Change can be reverted quickly if guest feedback or performance is negative.
Quick test design checklist (short)
- Write a clear hypothesis and decide the primary metric.
- Choose randomization method (within-shift randomization, time-block rotation, or location split).
- Create POS tags/codes and train staff on execution and comps handling.
- Estimate sample size and test duration; set success and rollback rules.
- Run test; monitor daily for safety signals; analyze at completion; roll out or revert based on rules.
Hypothesis templates
Write hypotheses that name the change, the expected outcome, and the metric you'll use:
- Price: "If we increase Item X from $10 to $11, then weekly revenue from X will increase by at least 8% and contribution margin per check will increase, measured by item sales and margin per check."
- Placement: "If we move Item Y to the top of the lunch menu, then Y's share of item mix will increase by 20% and average check will increase by at least $0.50."
- Portion: "If we reduce Side Z by 10% and lower the price by $0.50, then margin per serving will improve without reducing repeat orders more than 5%."
- Description/Imagery: "If we add a brief descriptive blurb to Item W, orders for W will increase by 15% within two weeks without reducing guest satisfaction."
Primary and supporting KPIs
Choose one primary metric to avoid multiple-testing pitfalls; track supporting KPIs to understand mechanisms.
- Primary candidates: Item unit sales (count), item sales share (mix%), revenue per day or per shift for the item, contribution margin per check, or average check.
- Supporting: overall check size, total revenue, cannibalization (drop in related items), void/comp rate, guest feedback/complaints, and repeat purchase indicators if available.
Sample size and test duration (practical guidance)
Statistical power matters. Small expected effects require large samples. Use a power calculation when possible; otherwise use practical rules of thumb and clear minimums.
Rule-of-thumb minimums:
- For large effects (20%+ relative change) aim for at least 200–500 item sales per arm.
- For medium effects (10–20%) expect multiple thousands of observations per arm.
- Very small changes (single-digit %) often need tens of thousands of observations to detect reliably and may not be worth testing as an isolated experiment.
Simple sample-size formula (two-proportion comparison) — useful when you’re comparing proportions (e.g., purchase rate of an item):
n per group ≈ [ (z_alpha * sqrt(2 * p * (1-p)) + z_beta * sqrt(p1*(1-p1)+p2*(1-p2)))^2 ] / (p2 - p1)^2
where p1 is baseline conversion (current purchase share), p2 is expected conversion under variant, p = (p1+p2)/2, z_alpha is 1.96 for 95% confidence, and z_beta is 0.84 for 80% power. Example: detecting a 15% relative lift on a low baseline (10% → 11.5%) may require thousands of customers per arm—plan accordingly.
Randomization & test setup options (pros/cons)
- Within-shift randomization (ideal): Randomly assign control vs variant at transaction or terminal level. Requires POS support for tagging and staff training. Best statistical properties.
- Time-block rotation: Alternate weeks or shifts. Easier to implement but susceptible to time-of-day and day-of-week bias; balance by rotating across comparable days.
- Location split: Run control at one location and variant at another. Good when you have multiple similar sites; watch for local differences in customer base.
POS tagging & data capture instructions
Make the test visible in the POS so sales can be isolated reliably.
- Create a distinct SKU, modifier code, or menu item code for the test variant and for the control (do not overwrite the original SKU).
- If changing price, create a separate priced SKU or a consistent price-level modifier so accounting stays clean.
- Use consistent naming: e.g., "X_TEST_A" and "X_TEST_B" or modifiers "ME_AB_A" / "ME_AB_B".
- Record metadata (tester, start/end, randomization method) in your test log so results are auditable.
- Train staff with a one-page job aid and make sure comps/voids are flagged and noted with reason codes so they can be excluded or analyzed separately.
Data cleaning rules
- Exclude voids and staff meals from primary analysis unless testing guest-facing acceptance.
- Flag unusually large or comped checks; analyze sensitivity including and excluding them.
- Watch for menu stockouts or supply substitutions that could bias results.
Monitoring during the test
- Check daily for safety signals: large drops in sales, guest complaints, or kitchen issues. Stop immediately for serious quality or safety problems.
- Monitor conversion and item sales but avoid peeking frequently for statistical decisions—use pre-specified analysis points.
Analysis & decision rules
Define decision rules before you start. Combine statistical significance with practical significance and risk tolerance.
- Primary rule example: If variant increases primary metric with p < 0.05 and increases contribution margin per check by at least $0.30, roll out.
- Conservative rule: If statistically significant improvement in primary metric but margin falls, investigate cannibalization and run a follow-up test with price/portion adjustments.
- Reject or rollback immediately if guest complaints or negative satisfaction signals rise materially even if KPIs look neutral.
Rollout plan
- Small-scale rollout to a handful of locations or shifts for 2–4 weeks while monitoring KPIs and guest feedback.
- Full rollout in phases with POS updates, staff training, and updated recipe cards / plating guides if portion or presentation changed.
- Post-rollout audit after 30–90 days to verify sustained impact and operational fit.
Common pitfalls and how to avoid them
- Changing multiple variables at once — avoid unless you want to test the combined change only.
- Running during unusual weeks (holiday, construction, one-off events) — pick representative periods.
- Small sample sizes — don’t overinterpret noisy results; use practical minimums or increase test size.
- Poor POS tagging — ensure variants are tagged and staff follow instructions.
Pre-built templates
Hypothesis (one line)
"If we [change X], then [primary metric] will [direction and expected magnitude], measured by [metric], within [duration]."
Pre-test checklist
- Confirm hypothesis and primary metric
- Decide randomization method
- Create POS SKUs/modifiers and test codes
- Train staff and provide job aid
- Set test start/end and monitor plan
- Define decision rules and rollback thresholds
Post-test report (table to capture)
Include: test id, dates, randomization, total transactions, item sales per arm, primary metric results (delta & p-value), margin impact, cannibalization effects, guest feedback summary, and decision (rollout / iterate / reject) with rationale.
Next steps and suggested experiments
Begin with experiments that are easy to implement and reverse (menu placement, description tweaks, add-on prompts). Use price and portion tests after you’re comfortable with tagging and sample sizing. Combine A/B testing with brief guest feedback prompts to capture qualitative context for surprising results.
Capability opportunities (practical)
This playbook becomes more powerful when paired with tools: an interactive sample-size calculator, POS tagging templates you can import, an experiment registry for tracking hypotheses and decisions, and dashboards that pull POS + margin data to show results. See capability notes below for suggested enhancements.
Use this playbook as a living template. Capture learnings, iterate on your hypotheses, and codify successful experiments into standard menu engineering practice so your menu becomes a continuous improvement engine.
Discussion
Comments and conversation will live here.