From The Blog · August 20, 2026
How to Calibrate Call Scores Across Dealership Locations
To calibrate call scores across multiple dealership locations, you need three things working together: a shared, behavior-anchored scorecard; a regular side-by-side scoring session where managers grade the same calls independently and then compare; and a documented standard for resolving scores when they diverge. Without all three, an 8 at your flagship store and an 8 at your satellite location will describe two completely different calls — and your QA data becomes useless for coaching or benchmarking.
The Real Problem With Multi-Rooftop Scoring
If you run more than one point, you’ve probably lived this: the GM at Store A says his team’s calls are averaging in the mid-8s. The GM at Store B says hers are in the low 7s. You pull sample calls from both stores expecting a clear skill gap — and the calls sound roughly the same.
The scores aren’t reflecting performance. They’re reflecting each manager’s personal interpretation of the scorecard.
That’s not a character flaw. It’s a calibration gap, and it’s extremely common. It happens because most scorecards include criteria that feel objective but aren’t — words like “professional,” “friendly,” “handled the objection well.” Without a shared definition of what those look like on a real call, every scorer fills in the blank differently.
The result: you can’t benchmark across stores, you can’t tell which location actually needs help, and when you move a rep from one store to another, their scores shift for reasons that have nothing to do with their performance.
Start With a Behavior-Anchored Scorecard
Before you can calibrate scorers, you need something concrete to calibrate to. Every criterion on your scorecard should describe an observable, audible behavior — not a judgment call.
Vague criterion: “Built rapport with the caller.”
Behavior-anchored version: “Asked at least one open question about what the caller was looking for or what mattered most to them, and listened without interrupting.”
The difference matters because the second version gives two scorers at two different stores a real thing to listen for. Either the rep asked an open question or they didn’t. Either they interrupted or they didn’t.
Do this for every row on your scorecard. If a criterion can’t be answered with a clear yes/no or a specific observable count, rewrite it until it can. This groundwork makes everything downstream easier.
For departments that haven’t formalized a scorecard yet, our post on how to build a call quality scorecard for parts and service covers the build process — the same principles apply to your sales and BDC scorecards. A structured call recording and scoring setup gives you the clean library you’ll need to pull calibration calls from.
How to Run a Calibration Session Across Rooftops
A calibration session is simple in structure. The discipline is in doing it consistently.
Step 1: Select the calibration calls. Pull three to five calls from a central pool — ideally a mix of quality levels: one strong, one weak, one in the middle. Avoid using calls from a store whose manager is in the room. That keeps the session focused on the scorecard, not on defending a store’s numbers.
Step 2: Score independently first. Send the calls and the scorecard to every participating manager or reviewer before the session. Each person scores every call on their own, no discussion. This step is non-negotiable — if people hear others’ scores first, they anchor to them instead of forming their own judgment.
Step 3: Compare scores out loud. Go criterion by criterion. Everyone reads their score for that item. Where scores match, move on. Where they diverge by more than one point, stop and talk it through.
Keep the conversation on the behavior: “What did you hear that made you score contact capture a 4? I had it as a 2.” Then pull up the exact moment and listen again together. This is where calibration actually happens — in the disagreement, not in a policy memo.
Step 4: Agree on a standard and write it down. When the group lands on the right score for a disputed moment, document the reasoning. Over time, these resolved disagreements become your calibration notes — a living record of what each score level sounds like. New managers can read them before scoring their first call.
Step 5: Set a cadence and stick to it. Monthly makes sense when you’re building or resetting a program. Once scorers are landing within one point of each other on most criteria, quarterly works as steady-state. Put it on the calendar like any standing meeting — “whenever we get around to it” calibration never happens.
What to Do When Scores Still Diverge
Even with good scorecards and regular sessions, you’ll get outliers. One manager consistently scores appointment-setting higher than everyone else. Another docks points for tone on calls where everyone else hears a warm rep.
A few practical fixes:
- Anchor scores to real examples. Build a small library of recorded calls the group agrees represent, say, a 4, a 6, and an 8 on your scale. When a scorer is uncertain, they compare against the anchor call.
- Add a tie-breaker role. Designate one person — a group QA director or an outside partner — whose score is the resolution standard when two scorers land more than two points apart.
- Track scorer variance as a metric. If one location’s QA manager consistently scores 1.5 points above the group average on identical calls, that’s data. Address it directly in coaching.
What a Calibrated Scorer Actually Hears
Calibration only matters if scorers are listening for the right phone moments. Take the leverage moment on a sales call — the point where a caller asks if a vehicle is available. C&M teaches reps to capture the name and number before handing out info:
Customer: “Is that silver Tahoe still available?” Rep: “Let me check on that for you right now — who do I have the pleasure of speaking with? … And the best number in case we get disconnected while I check?”
A calibrated scorer hears that and marks contact capture as a clear yes. An uncalibrated one might give partial credit for a rep who blurted out “Yep, it’s here!” and lost the lead. When your whole group agrees that the second version fails the criterion, your scores finally mean the same thing everywhere.
Department-Specific Calibration Checkpoints
Not every department’s calls are scored the same way — and they shouldn’t be. A sales call’s goal is an appointment; a service call’s goal is a booked visit, not a phone diagnosis. Group your calibration calls by department and use the right scorecard for each.
A few things to listen for during calibration, by department:
Sales / BDC calls:
- Did the rep capture name and contact info before answering the availability question?
- Was a specific day and time offered — or just a vague “come on by”?
- Did the rep recap at the end: day, time, vehicle, who to ask for, directions?
Service calls:
- Did the advisor work toward a booked appointment rather than diagnosing over the phone?
- Was the caller offered a specific time slot, or left open-ended? Our post on getting a service caller to commit to a specific time breaks this down.
Parts calls:
- Did the rep confirm fitment before quoting or pulling inventory?
- Was the caller’s order or visit captured?
Our quality assurance and call monitoring service is built to handle exactly these department-level distinctions — which is especially useful when internal managers are too close to their own teams to score consistently.
Turning Consistent Scores Into Coaching
Calibration isn’t just QA housekeeping. Once your scores mean the same thing across locations, they become a legitimate coaching tool.
Now you can genuinely compare stores. You can tell whether a skill gap — contact capture, appointment setting, call control — shows up at one location or systemwide, and prioritize coaching on data instead of a manager’s gut. Our post on how to use call scores to set rep coaching priorities walks through that.
Consistent scoring also protects rep morale. When a rep moves stores or gets a new manager, their scores shouldn’t move unless their behavior does. That builds trust — the team believes the number means something real.
Bringing in an Outside Perspective
One of the fastest ways to expose scoring drift is to have an outside set of ears score the same calls your managers score, then compare. The gap between an internal and external score shows you exactly where — and in which direction — your calibration has drifted. It’s the same discipline C&M uses calibrating scorers across high-call-volume operations well beyond auto, including collections and finance companies.
If you want a no-cost look at where your stores stand right now, request a free mystery shop for an objective baseline from calls your team doesn’t know are being evaluated. It’s a clean way to hear what a calibrated outside scorer catches before you invest in a full QA build.
For groups wanting broader operational context, NADA publishes ongoing guidance on dealership operations, and Cox Automotive offers research on how customer experience — phone handling included — ties to retention.
Calibration is unglamorous. It’s slower than handing out scorecards and telling managers to go. But when a call earns an 8, it should earn an 8 because of what the rep did — not because of which store they work at. That’s what makes a multi-rooftop QA program worth running.
Frequently Asked Questions
- What does it mean to calibrate call scores across dealership locations?
- Calibration means every manager or scorer at every rooftop applies the same standards when grading a call — so a call that earns an 8 at Store A would earn an 8 at Store B. It involves scoring the same sample calls, comparing results, and resolving disagreements until scorers align on what each number actually represents.
- How often should a dealership group run calibration sessions?
- For groups launching or fixing a QA program, monthly calibration sessions are a reasonable starting cadence. Once scorers are consistently within one point of each other on most calls, quarterly is appropriate — with a quick re-calibration any time you add a new location, new scorer, or update your scorecard.
- What is the most common reason call scores differ across locations?
- The most common reason is vaguely written scorecard criteria — phrases like 'was professional' or 'built rapport' mean different things to different managers. Without concrete, observable behaviors tied to each criterion, two scorers listening to the same call will legitimately arrive at different scores.
- Should every department use the same call scorecard?
- No. Sales, service, BDC, and parts calls have different goals, so their scorecards should reflect those differences. A sales call is scored on whether an appointment was set; a service call is scored on whether the advisor booked a visit rather than diagnosing over the phone. Calibration still applies within each department — you just calibrate scorers on the right scorecard for each role.
- Can recorded calls be used for calibration without demoralizing staff?
- Yes — the key is to frame the session around the scorecard, not around individuals. Use calls that illustrate specific scoring questions rather than calling out a rep, and when possible use calls from outside your stores or composite examples. The goal is to align scorers on standards, not to critique performers publicly.
Put this into practice
Related C&M Coaching training & services: