The essentials on procedural and relational quality monitoring
- Two schools answer the question "was this call good?" differently. The procedural school checks that the planned steps took place: introduction, rephrasing, summary, closing phrase. The relational school checks what the exchange produced for the customer: a problem solved, a clear next step, a reason to trust.
- Procedure is easier to measure, relationship matters more. A 2011 meta-analysis shows that loyalty depends at least as much on how customers were treated as on what they obtained. And as soon as a request falls outside the standard, a mostly scripted exchange gets the lowest perceived quality.
- Our stance, in the manner of the Agile Manifesto: interactions over phrases, the problem solved over the box ticked, the commitment kept over the script followed. With one exception: compliance is not a ritual, it stays in the scorecard.
- Measuring the relationship does not require measuring emotion. Three observable facts are enough: what the customer expresses on the way out compared with what they expressed on the way in, what the advisor promised, and whether the customer calls back about the same issue.
- A relational criterion is almost always conditional: it pairs a customer signal with the advisor's response. That is what sets it apart from a ritual, and what makes it fair.
Two ways to judge the same call
Among telecom operators, some stand out for their approach to customer relationships and, as a result, for the way they use conversation analytics. One of them recently summed it up for us in a single sentence: whether the advisor introduced themselves and rephrased the request is of no interest. What they want to know is whether the customer left with their problem solved, and with a reason to trust what happens next.
That sentence separates two ways of practising quality monitoring, which every contact centre knows without always naming them.
The procedural school judges a call by what the advisor did. Its scorecard describes the expected flow, step by step, and checks that each step took place. Its question: did the advisor do what was planned?
The relational school judges a call by what it produced for the customer. Its scorecard describes situations and the response they call for, then looks at how the exchange ends. Its question: does the customer leave in a better position than when they called?
| Procedural | Relational | |
|---|---|---|
| Question asked | Did the advisor follow the flow? | What did the exchange produce for the customer? |
| Typical criterion | "The advisor rephrases the customer's request." | "If the customer describes a concrete consequence of the problem, the advisor takes it into account in the solution." |
| When it applies | To every call, whatever the customer says | When the situation arises, otherwise N/A |
| What it sees well | Consistency, compliance, how well training is applied | Resolution, commitments, what the customer expresses |
| What it misses | A call that is perfect on paper but leaves the customer without a solution | What does not depend on any situation, such as a mandatory disclosure |
| Its trap | Phrases recited for the score | An outcome that also depends on what the advisor does not control |
No set-up is purely one or the other. But most scorecards in use lean clearly towards the procedural side, and that is no accident.
Why procedure won
The procedural scorecard was born of a measurement constraint. When quality control relies on listening to a few calls per advisor per month, each evaluator must be able to decide quickly, and in the same way as their colleague. "Did they introduce themselves?" is settled in three seconds, without debate. "Did they take into account what the customer told them?" requires understanding the whole call, and two evaluators will not always answer it the same way.
Procedure therefore took hold because it could be measured, not because it mattered most. It has real virtues: easy to train, consistent from one site to another, essential for anything that is a legal obligation. It also has a flaw that standards bodies know well. COPC, which publishes performance standards for contact centres, recommends separating the errors that matter to the customer, those that matter to the business and those related to compliance, because "quality scores and CSAT results often tell different stories" (May 2024). In the example it cites, a financial services company showed 86% overall quality: 100% on compliance, 70% on internal processes, and only 60% on what matters to the customer. The average hid precisely the dimension that makes customers come back or leave.
Automated analysis of every call removes the original constraint. A language model reads the whole call, identifies the situation, then judges the response it received, with the passage to back it up. Relational criteria, long considered too subjective for a scorecard, become measurable across all conversations. The question is no longer whether the relationship can be measured, but whether you choose to measure it.
What research says: treatment matters as much as outcome
Services marketing has been studying for nearly thirty years how customers judge the handling of a problem. The findings point in the same direction.
- Three judgements, not one. Faced with a complaint, customers judge the outcome they received, the procedure followed to reach it, and the way they were treated during the exchange. The resulting satisfaction has a direct impact on their trust in, and commitment to, the company (Tax, Brown and Chandrashekaran, Journal of Marketing, 1998).
- For loyalty, the manner weighs at least as much as the outcome. A meta-analysis of studies on complaint handling shows that the outcome mainly explains satisfaction with the transaction. Cumulative satisfaction, the kind that precedes loyalty, depends at least as much on the quality of treatment. The effect is stronger in services, and when the problem is not primarily financial (Gelbrich and Roschk, Journal of Service Research, 2011).
- Scripts hurt as soon as a request falls outside the standard. An experiment varied the degree of scripting in standardised and customised encounters. In a standardised encounter, more scripting makes no difference to perceived quality. In a customised encounter, a mostly scripted exchange gets the lowest perceived quality (Victorino, Verma and Wardell, Production and Operations Management, 2013).
- Being relational does not mean doing more. According to CEB, a firm since absorbed into Gartner, "96 percent of customers who endure a high-effort experience" report being more likely to be disloyal, "compared to only 6 percent" of those whose interaction was easy (press release of 16 October 2013). A relationship is built less through warmth than through what the customer is spared: repeating themselves, being transferred, calling back, doubting.
The third point deserves a closer look. The calls that matter most for the relationship, complaints, cancellations, incidents, are never standard. That is precisely where a scorecard that imposes set phrases pushes the advisor towards the script, and therefore towards the worst perception.
Our stance: a manifesto for relational quality monitoring
In 2001, seventeen software development practitioners published the Agile Manifesto. Its first value fits on one line: "Individuals and interactions over processes and tools". The text ends with a sentence that is quoted less often: "That is, while there is value in the items on the right, we value the items on the left more."
The manifesto did not abolish processes. It reversed the order of priorities, at a time when method had become an end in itself. Quality monitoring has reached the same point. So we borrow its form.
To judge the quality of a call, we value:
- interactions with the customer over phrases recited;
- the problem solved over the box ticked;
- the commitment kept over the script followed;
- what the customer says on the way out over what the advisor says on the way in.
That is, while there is value in the items on the right, we value the items on the left more.
One clarification sets us apart from the original manifesto. There is one kind of procedure whose value is not up for debate: compliance. Authenticating the account holder before opening their file, announcing that a call is commercial, informing the customer of a right of withdrawal or cancellation: these are not rituals. Leaving them out creates a dispute, a void contract or a penalty. These criteria stay in the scorecard, with a weight that reflects their seriousness.
A ritual is what gets checked because it can be checked: the first name, the greeting, the "is there anything else I can help you with?". The procedural approach puts rituals and obligations on the same level. The relational approach separates them.
Measuring the relationship without measuring emotion
The relational need is often expressed like this: "knowing whether the customer's emotion was positive". The wording is understandable, but it leads to a dead end.
Legally, first. The European AI Act defines emotion recognition as identifying or inferring the emotions or intentions of a person on the basis of their biometric data, such as their voice. It has banned it in the workplace since 2 February 2025, and its recital 18 explicitly lists satisfaction among emotions. Applied to customers, it is not prohibited, but it falls under high-risk systems. Analysis based solely on the text of the transcript falls outside the prohibition, according to the guidelines published by the Commission in February 2025. Even so, as soon as an emotion indicator is used to judge how an employee handled a call, the question comes back.
Practically, above all. The same regulation points out the limited reliability of these systems. An emotion score varies with the nature of the calls more than with the quality of the handling: a complaint starts from an unfavourable situation by definition. It points to no passage that can be replayed, and therefore to nothing that can be discussed with an advisor. Part of the market answers with "predicted" satisfaction based on tone and vocabulary: it signals that a call went badly, it does not say what should have been done differently.
The good news is that the relationship does not need emotion to be measured. It leaves three observable traces.

1. The shift: what the customer expresses on the way in, and on the way out
We do not measure the customer's state, we note what they say, at the start and at the end. A customer who opens the call with a complaint and closes it with explicit agreement or heartfelt thanks has changed position, and the transcript shows it, with the passage to back it up. It is a fact, not an interpretation.
The criterion is written conditionally: "If the customer expresses dissatisfaction at the start of the exchange, they express agreement or thanks at the end of the exchange, beyond a mere polite formula." It is N/A when the customer does not start out dissatisfied. Calculated only on calls that started from an unfavourable situation, its rate gives a turnaround rate, which does not depend on the mix of call reasons.
2. The commitment: what was promised, and how the customer received it
Trust cannot be decreed, it can only be observed, in both directions. On the advisor's side: did they announce a next step, a timeframe, and who is handling it? Did they guarantee a result that does not depend on them? On the customer's side: do they rely on what they are told ("if you say so"), or do they bring up an earlier promise that was not kept?
3. The follow-up: does the customer call back about the same issue?
This is the only measure that does not rely on the call itself. A customer who thanks the advisor warmly and then calls back three days later for the same reason was not helped, whatever the transcript says. Linking the calls of a single customer, by their phone line or customer number, is enough to find out.
The difference between measuring emotion and measuring the relationship often comes down to a few words:
| Wording to avoid | Observable wording |
|---|---|
| The customer's emotion was positive | The customer, dissatisfied at the start, expresses agreement or thanks at the end |
| The customer left satisfied | The customer expresses agreement with the proposed solution at the end of the exchange |
| The advisor built trust | The advisor announces the next step, its timeframe and who is handling it |
| The customer is distrustful | The customer brings up an earlier promise that was not kept |
| The problem is solved | The customer does not call back about the same issue within the week |
Each line on the right points to a precise passage or a dated fact. It can be shown to an advisor, challenged, corrected. None of them attributes an inner state to anyone.
Writing a relational scorecard: the conditional criterion
A procedural criterion checks that a step is present, whatever the customer. A relational criterion pairs a customer signal with the advisor's response. It is almost always a conditional criterion: it starts with "If" or "When", and is N/A when the situation does not arise.
| Procedural criterion | Relational criterion to write instead |
|---|---|
| The advisor introduces themselves by first name. | If the customer says they have already called about the same issue, the advisor refers to what has already been done without asking them to explain everything again. |
| The advisor rephrases the customer's request. | If the customer describes a concrete consequence of the problem for them, the advisor takes it into account in the solution they propose. |
| The advisor summarises the exchange before closing. | If the request cannot be resolved during the call, the advisor announces the next step, its timeframe and who is handling it. |
| The advisor asks the customer whether they have any other questions. | The customer asks a question that the advisor does not answer. |
| The advisor uses the closing phrase. | If the advisor cannot grant what the customer asks for, they explain why. |
Three properties make this type of criterion fairer:
- It does not penalise a call in which nothing happened. A customer who asks a simple question and leaves with the answer offers no opportunity to "take a consequence into account": the criterion is N/A, not zero.
- It does not reward a ritual. A rephrasing recited out of habit validates nothing. The only thing that counts is the response to what the customer actually said.
- It can be verified. The verdict points to two moments in the call: the one where the customer describes their situation, and the one where the advisor responds to it.

In Raisetalk, a criterion that accepts N/A is answered N/A when the situation it describes does not arise in the call, and each verdict comes with its justification and the relevant passages. Before deploying a rewritten scorecard, run it on calls that have already been evaluated to compare the two readings: that is the purpose of re-evaluating with another scorecard.
Score the behaviours, count the outcomes
One delicate question remains: should the advisor's score depend on the outcome of the call? Our answer is no. The outcome also depends on what the advisor does not control: an outage, a price increase, a commercial decision made elsewhere. An advisor assigned to complaints would get a low score without being able to change anything, and one who does everything right with a customer who leaves anyway would be penalised for a decision that is not theirs.
The rule we recommend: the advisor's behaviours go into the score, the outcomes are counted separately. Turnaround, agreement expressed at the end of the exchange, doubt about what happens next, repeat calls about the same issue are read as rates, by team, by site and by call reason, never as a ranking of people. In Raisetalk, a criterion that neither adds nor removes any points keeps its verdict without affecting the score: it remains filterable, feeds reporting and can trigger an alert.
Don't decree relational behaviours: measure them
The relational approach has its own trap, and it closely resembles the one it criticises: replacing one list of set phrases with another. "I completely understand your situation", "I'll do everything I can to help you". An empathy phrase recited for the scorecard is a ritual like any other, and the customer hears it as one.
The way out is to reverse the approach. Rather than deciding in advance which behaviours build the relationship, start from the outcomes and look, in your own calls, for the behaviours that precede them. Which advisor behaviours are more frequent in calls that lead to no repeat call? Which are missing from calls that end in a complaint? And above all, which ones make a difference at equal score, so as not to confuse a behaviour that explains the outcome with one that merely accompanies it?
That is the job of the "mechanics" page of the Raisetalk management report. For a given outcome, a repeat call from the customer, a complaint, an expected result such as an appointment booked, it compares the rate of the outcome when a behaviour is present and when it is absent, repeats the comparison at equal overall score, and separates the levers from mere symptoms. For a first exploration, the Understand page accepts the question as it stands: "In calls where the customer was dissatisfied at the start and thanks the advisor at the end, what does the advisor do in between?". The answer comes with the excerpts it is based on.
The result is a shorter scorecard, which keeps only what matters to your customers. Procedure does not disappear: it is put at the service of the relationship.
The test: is your set-up procedural or relational?
Eight statements. Count, in each column, those that describe your current set-up.
| Statement | Leaning |
|---|---|
| Your scorecard contains a criterion on the advisor's introduction or on the closing phrase. | Procedural |
| More than half of your criteria apply to every call, whatever the customer says. | Procedural |
| An advisor can score 100% on a call in which the customer leaves without a solution. | Procedural |
| Your quality score is going up, and your complaints are not going down. | Procedural |
| A good share of your criteria start with "If the customer...". | Relational |
| You track what the customer expresses at the end of the exchange, separately from the advisor's score. | Relational |
| You know how many customers call back about the same issue within the week, team by team. | Relational |
| You know which of your advisors' behaviours precede a resolution, and which make no difference. | Relational |
Three or four procedural statements, one or no relational one. Your set-up faithfully measures whether the flow was followed. It is solid for compliance, and tells you almost nothing about what makes your customers come back or leave.
As many of one as of the other. You are in transition. The risk is adding the two together in a single score, where rituals dilute what matters. Separate them.
Three or four relational statements. Your set-up measures what the call produces. Check that it has not rebuilt a list of empathy phrases, and that compliance keeps the weight it deserves.
Moving from one to the other in four steps
1. Sort the scorecard into three families
Go through each criterion and place it in one of these families: compliance (legal or contractual obligation), relationship (what changes something for the customer), ritual (what is checked because it can be checked). Compliance stays, with a weight that reflects its seriousness. Rituals come out of the score.
2. Rewrite as conditional whatever can be
Some rituals hide a sound intention. Rephrasing is meant to check understanding: the useful criterion is about the response given to what the customer actually said. The summary is meant to set out what happens next: write the next step itself, with its timeframe and the person handling it.
3. Add three outcomes, outside the score
Agreement expressed at the end of the exchange, turnaround, doubt about what happens next. Add repeat calls about the same issue if you can recognise a customer from one call to the next. Read these outcomes as rates, call reason by call reason, to compare like with like.
4. Measure, then prune
After a few weeks, look at which behaviours precede good outcomes at equal score, and which make no difference. Keep the former, remove the latter, and train on what makes the difference, with examples taken from your own calls, during the evaluation debrief.
Start with cancellation or complaint calls. They are the least standard, the ones where a scripted exchange most damages the customer's perception, and the ones where the relationship is worth the most. A relational scorecard of about ten criteria on this scope alone is enough to see, within a month, what the procedural scorecard was not showing.
What the relational approach does not replace
- Compliance. It remains procedural by nature, and rightly so: a mandatory disclosure is checked by its presence.
- Human judgement. Deciding that a problem is really solved sometimes remains a human decision. A criterion can be reserved for manual evaluation, alongside automated analysis: that is the principle of hybrid analysis.
- Customers who do not call. Those who leave without a word leave no trace in the calls. That is the limit of any voice of the customer built from calls.
In Europe, French customer experience leaders are the ones who most often put loyalty and empathy at the top of their priorities, as the Genesys 2026 study showed. The choice of the relationship has therefore already been made. What remains is to measure it, rather than measuring what is easy.
Key terms
- Procedural quality monitoring: an approach that evaluates a call by the presence of the planned steps, whatever the customer
- Relational quality monitoring: an approach that evaluates a call by what it produced for the customer: resolution, commitments, what the customer expresses on the way out
- Ritual: a step checked because it is easy to check, with no proven link to the customer's experience
- Conditional criterion: a criterion that only applies if a specific situation arises in the exchange, and is N/A otherwise
- Turnaround: a call in which the customer, dissatisfied at the start, expresses agreement or thanks at the end
- Observable outcome: a fact noted at the end of the exchange or after it (agreement expressed, doubt about what happens next, repeat call about the same issue), read as a rate rather than built into the score
Frequently asked questions
What is relational quality monitoring?
It is a way of evaluating calls that starts from what the exchange produced for the customer, rather than from the steps the advisor followed. Its criteria describe situations and the response they call for: a customer calling back about the same issue, a request that cannot be resolved straight away, a refusal that needs explaining. Alongside them, it tracks observable outcomes, such as agreement expressed at the end of the exchange or repeat calls about the same issue.
Should introduction and rephrasing criteria be removed?
Not necessarily from the scorecard, but at least from the score. They are steps that are easy to check, whose link with the customer's experience has not been established, and which turn into set phrases as soon as they are scored. When they carry a useful intention, such as checking understanding, rewrite them as a conditional criterion that judges the response given to the customer.
Can a customer's emotion be measured during a call?
It is neither necessary nor desirable for evaluating quality. The European AI Act strictly regulates emotion recognition from the voice, and bans it in the workplace. Above all, an emotion score points to no verifiable passage. What the customer expresses at the start and at the end of the exchange, noted in the transcript, says what matters in a verifiable way.
Should the advisor's score depend on customer satisfaction?
No. The outcome of a call also depends on factors the advisor does not control, such as an outage or a price increase. The score should cover their behaviours, and the outcomes should be tracked separately, as rates by team and by call reason.
How can you tell whether a call really solved the customer's problem?
The most reliable sign is that there is no repeat call about the same issue in the following days. This requires being able to recognise a customer from one call to the next, by their phone line or customer number. Within the call itself, the customer's expressed agreement with the solution and the announcement of a dated next step are the best indicators.
Getting started
Your scorecard says what your advisors do. To find out what your calls produce, try Raisetalk on your own conversations.

