Menu A/B Test Plan & Results Template
A practical playbook and ready-to-use templates to design, run, and evaluate menu A/B tests for placement, price, pictures, or descriptions. Includes hypothesis templates, segmentation guidance, measurement and sample-size considerations, a data collection CSV template, guardrails to protect margins and operations, an analysis plan, clear decision rules, and a runbook checklist.
Welcome — run safer, smarter menu experiments
This playbook helps you test menu changes with evidence, protect margins, and avoid confusing staff or guests. It takes you from a focused hypothesis to a clear decision: keep, iterate, pilot, or stop. Use the templates and the runbook checklist so experiments are easy to repeat and learn from.
Quick overview: the test flow
- Define a clear hypothesis and primary metric.
- Choose audience segmentation and test scope.
- Define measurement window and minimum sample size.
- Collect reliable data (automate where possible).
- Apply guardrails to protect operations and margins.
- Analyze using a pre-specified plan and apply decision rules.
- Act: launch, iterate, scale, or roll back.
1 — Hypothesis (template)
Write one concise test hypothesis that includes expected direction and measurable outcome. Include a clear primary metric and a business success threshold.
Template: If we [change X: price/placement/photo/description], then [measurable outcome: e.g., add-to-order rate, item attach rate, average check] will [increase/decrease] by [target amount or percent] during [daypart or date range]. Success = [business threshold: e.g., +5% attach rate AND no >2% drop in overall covers].
Example: If we move the new pasta to the top of the dinner section and add a plated photo, then the pasta attach rate will increase by at least 8% during dinner service (5–9pm). Success = ≥8% lift and no negative impact on average ticket or kitchen throughput.
2 — Segmentation and scope
Decide who and when the test applies to. Common segments:
- Dayparts (lunch, dinner, brunch)
- Channels (in-house dining, online ordering, delivery)
- Guest segments (new vs returning, loyalty members)
- POS terminals or locations (single-station, random table assignment, or location A/B)
Keep the scope tight for operational simplicity: single daypart and single channel is easier to run and interpret.
3 — Measurement window
Pick a start/end date long enough to reach the required sample and to smooth day-to-day variability (minimum one full week to cover weekday/weekend patterns, longer for slow items). Include a lead-in period for staff training if needed. Avoid running across major holidays, menu-price changes, or local events unless those are part of the test conditions.
4 — Sample size guidance (practical, not scary)
Sample size depends on baseline rates, the change you want to detect, and how confident you want to be. If you don’t have a data team handy, use these rules of thumb:
- For high-traffic items or sitewide tests, aim for at least several hundred to a few thousand orders per variation before trusting small percentage-point lifts.
- For lower-traffic restaurants, design tests to detect larger effects (5–10%+ relative lift) or run longer pilots.
- If you care about revenue per check rather than attach rate, capture average check and its variance; revenue tests usually need larger samples than binary attach tests.
When precise power calculations are needed, use a standard sample-size calculator (proportions or means) or ask a data analyst. Pre-specify the minimum detectable effect and acceptable risk (commonly 80% power, 5% alpha).
5 — Data collection: what to record (CSV-ready)
Automate via POS, online ordering tags, or analytics when possible. If manual, use a simple CSV with these columns:
variant_id,order_id,timestamp,channel,daypart,location_id,guest_segment,item_id,item_present (0/1),item_quantity,order_value,comps_or_refunds,notes
Variant_id distinguishes control vs treatment. Track order_value so you can compute average ticket and revenue impact. Record comps/refunds to avoid false signals.
6 — Guardrails to protect margins and operations
- Price guardrail: maximum allowed price change or acceptable margin impact.
- Inventory guardrail: ensure stock levels support the test without causing spoilage.
- Operational guardrail: limit the number of menu changes per shift to avoid staff confusion.
- Service guardrail: measure ticket times and remake rates; set thresholds that trigger immediate rollback.
- Guest experience guardrail: limit visual changes that could cause ordering confusion for loyalty members or reservations.
7 — Analysis plan (pre-register it)
Before you start, write down the primary metric, secondary metrics, and how you will compute uplift. Examples:
- Primary metric: item attach rate = (orders including item) / (total orders in segment)
- Secondary metrics: average check, revenue per order, ticket time, remake rate, refund rate.
Pre-specify how you'll handle: multiple comparisons (limit number of tested variants or declare one primary comparison), outliers (cap refund values), and early stopping rules (avoid peeking unless using sequential testing methods).
8 — Decision rules and recommended actions
Define clear outcomes and actions ahead of time so results aren’t interpreted selectively.
- Clear win: Statistically significant improvement in the primary metric AND no negative guardrail triggers. Action: roll out to remaining dayparts/locations or update menu.
- Mixed or small lift: Small but positive lift or improvement in secondary metrics. Action: pilot longer, collect more data, or iterate creative (photo/descriptor).
- No effect or negative impact: No measurable benefit or negative on guardrails. Action: revert to control and document learnings.
9 — Runbook checklist (practical steps)
- Assign owner (test lead), data lead, ops lead, and chef/trainer.
- Create POS/online variant tags and ensure staff know the plan (script + visual cue).
- Run training shift or quick checklist for servers and cooks.
- Start test; monitor daily for guardrail breaches (orders, ticket time, comps).
- At test end, export data, run the pre-specified analysis, and prepare results report.
- Decide using the pre-defined decision rules and communicate rollout or rollback steps.
10 — Results report template (what to include)
- Test summary: hypothesis, dates, locations, channels, owner.
- Primary metric result with absolute and relative lift, confidence intervals, and p-value if used.
- Secondary metric summary (average check, ticket time, comp rates).
- Operational notes: staff feedback, inventory issues, customer comments.
- Recommendation and next steps.
Appendix: Example filled plan (short)
Hypothesis: Moving the grilled salmon to the top of the mains and adding a plated photo will increase salmon attach rate by ≥10% at dinner. Primary metric: salmon attach rate during dinner (5–9pm). Measurement window: 14 days starting Tue 21st. Sample target: minimum 1,000 dinner checks per variant or run full 14 days and reassess. Guardrails: no >3% increase in remake rate; no >2% drop in average check. Owner: Ops Manager. Data lead: GM.
Final notes & recommended capability upgrades
This playbook preserves the practical steps and templates you need. To make tests easier and reduce manual work, consider adding an interactive planning form that stores test designs and test results (so teams can reuse, compare, and aggregate learnings). Automate tagging in your POS/online ordering where possible and capture variant_id at checkout. See Capability Enhancement Notes for concrete next steps.
Discussion
Comments and conversation will live here.