Feedback Calibration Meeting Agenda & Rubric

A practical, facilitator-ready agenda and a behavior-anchored rubric to align evaluators, reduce bias, and produce fairer, more consistent feedback and ratings. Includes prework, timing, facilitator script, bias-check questions, and a short practice case.

Purpose

Help evaluators reach shared standards for ratings and feedback so assessments are actionable, consistent, and psychologically safe. Use this template to run a focused calibration meeting, practice on one or two cases, and produce clear follow-up actions.

Intended outcomes

  • Shared understanding of rating bands and behavior anchors
  • Reduced unwarranted variability and bias in ratings
  • Concrete changes to guidance, training, or scoring rules when needed

Prework (complete before the meeting)

  1. Facilitator: select 3–5 real or anonymized assessment examples (recordings, notes, outputs).
  2. Each evaluator: independently review assigned examples and record an initial rating and 1–3 pieces of evidence that support your rating.
  3. All: read the rubric (below) and note any questions or unclear anchors.

Roles

  • Facilitator: keep time, surface disagreements, ask bias-check questions, capture decisions and action items.
  • Evaluators: bring evidence-based ratings and be prepared to explain the evidence behind your rating.
  • Scribe: record agreed rubrics, clarifications, and action items (can be the facilitator).

Suggested agenda (60–90 minutes)

  1. Opening (5–10 min) — Purpose, norms (psychological safety, focus on evidence not people), and desired outcomes.
  2. Rubric review (10–15 min) — Quickly confirm the rating bands and behavior anchors. Note any ambiguous language to resolve later.
  3. Case discussions (30–50 min) — For each selected case:
    1. Reveal each evaluator's initial rating (no names attached) and the top evidence items (2–3 min).
    2. Facilitated discussion: clarify facts, surface differing interpretations, run bias checks (8–12 min).
    3. Take a re-vote or reach consensus and record rationale (2–3 min).
  4. Norming & action items (10–15 min) — Consolidate clarifications into rubric updates, training needs, or documentation; assign owners for follow-up.
  5. Close (2–5 min) — Review decisions, confirm owners and deadlines, and schedule next calibration if needed.

Psychological safety norms to state at the start

  • We critique evidence and interpretations, not people.
  • Disagreement is expected; our job is to learn and improve consistency.
  • Speak from observable behavior and artifacts, not assumptions about intent.

Calibration Rubric (behavior-anchored)

Use the rubric as a living tool — capture examples during calibration and update anchors when patterns of disagreement emerge.

Rating bands (example: 1–5)

  1. 1 — Does Not Meet Expectations
    • Behavior anchors: Frequently misses critical requirements, shows repeated errors, or demonstrates unsafe/contradictory actions.
    • Evidence examples: multiple documented failures, customer complaints, missing critical deliverables.
  2. 2 — Below Expectations
    • Behavior anchors: Inconsistent performance, needs frequent coaching, important elements often incomplete.
    • Evidence examples: recent missed targets, partially resolved issues, reliance on others to correct work.
  3. 3 — Meets Expectations
    • Behavior anchors: Consistently performs core responsibilities with acceptable quality and occasional coaching.
    • Evidence examples: reliable task completion, stable metrics, positive stakeholder feedback on key items.
  4. 4 — Exceeds Expectations
    • Behavior anchors: Regularly delivers above standard, improves processes, or helps peers succeed.
    • Evidence examples: measurable improvements, proactive problem solving, positive testimonials.
  5. 5 — Outstanding
    • Behavior anchors: Role model; leads initiatives with sustained, demonstrable impact beyond scope.
    • Evidence examples: cross-functional leadership, repeatable successes, external recognition.

How to cite evidence during discussion

  • State the specific observable behavior or artifact (e.g., "On 2026-03-10 in the client call, they interrupted the client twice and did not follow up on the promised action").
  • Explain why that behavior aligns with a band anchor.
  • Indicate whether the evidence is one-off or part of a pattern.

Bias-check questions (use these aloud during disagreement)

  • Am I focusing on a recent event (recency bias) rather than longer-term pattern?
  • Would I rate someone the same if they were from another team/background (similarity bias)?
  • Is one strong positive/negative moment (halo/leniency) driving my whole judgment?
  • Am I attributing behavior to intent rather than situational factors (attribution error)?
  • Are my expectations consistent with the documented role or goal for this person (role inflation)?

Facilitator script — suggested prompts

  • "Let's list the observable evidence each person relied on; remember to be specific."
  • "Who has a different interpretation? What fact would change your view?"
  • "Which bias checks should we run on this case?"
  • "If we can't reach consensus, what would be a defensible provisional rating and what follow-up would we assign?"

Practice case (facilitator-ready)

Case: During a customer onboarding call, the evaluator being assessed (Alex) missed two customer questions, provided a partial answer, and promised follow-up within 24 hours. Follow-up was provided after 72 hours and required manager intervention to resolve a configuration issue. Evidence: call recording, follow-up email timestamp, support ticket showing manager intervention.

  1. Ask evaluators to share their initial rating and the 1–2 pieces of evidence they used.
  2. Discuss: Did Alex's initial mistakes materially harm the customer outcome? Was the late follow-up a sign of poor process or an isolated logistic issue?
  3. Run bias checks and re-vote. Capture a short rationale for the agreed rating.

Meeting outputs / deliverables

  • Updated rubric anchors or clarified scoring rules (document specific wording changes).
  • List of cases re-scored and their final ratings with brief rationales (stored for audit/training).
  • Assigned action items (owner, due date) — e.g., update training, add a checklist, re-train raters.

Follow-up checklist

  • Distribute the updated rubric to all raters and attach example evidence for each band.
  • Run short training sessions or micro-lessons on common disagreement patterns identified.
  • Schedule next calibration cadence (monthly/quarterly depending on volume and variability).

Notes on adapting this template

Make the rubric specific to your role, domain, and measurable outcomes. Use concrete artifacts (recordings, tickets, deliverables) as the primary evidence. Keep calibration meetings short and regular: frequent, lightweight norming beats occasional heavyweight debates.


Discussion

Comments and conversation will live here.