Key takeaways
- A quality scorecard is tuned, not written in one sitting. Until a criterion has run on real conversations, nobody truly knows what it measures.
- Re-evaluation closes the loop: you replay a scorecard you are still building on conversations you already know, as many times as needed, with no re-import.
- The existing score is never touched. Re-evaluation adds a second evaluation alongside it, with its own score and its own history.
- Only the criteria analysis is billed again. Transcription, summaries and tags are carried over as they are.
- A scorecard overhaul can be measured before you switch: re-score a handful of conversations with the new version, compare with the old one, and decide on numbers rather than on a hunch.
A quality scorecard isn't written, it's tuned
Ask a quality manager how they built their evaluation framework. They will rarely describe a writing session. They will describe months of adjustments: a criterion so vague it always returned the same verdict, a weighting that crushed everything else, a wording that meant one thing to them and something else to their evaluators.
That scorecard is the foundation of any quality monitoring programme: it decides what gets measured, and therefore what gets managed and coached afterwards. When an AI applies the scorecard, the tuning work does not disappear. It moves. The question is no longer "do my evaluators read this criterion the same way?" but "does this criterion, worded the way I wrote it, actually detect what I am looking for?". And there is only one way to settle that question: run it on real conversations, then read what it says about them.
That is what re-evaluation makes possible. From a conversation's page, you create a second evaluation, on the scorecard of your choice, including the one you are still writing.
The short cycle: write, test, correct
In practice, tuning a scorecard looks like this.
You pick a few reference conversations, five or ten, that you know inside out. An excellent call, a poor one, two or three borderline cases: the ones where you already know what the score should say.
You write a first version of your scorecard, then run it on those conversations. You don't just look at the score. You open each criterion and read the justification produced by the analysis: the sentence explaining why the criterion was considered met or not, together with the passage of the conversation that triggered it. That is where a badly worded criterion shows itself. The verdict can be right for the wrong reason, or wrong for a reason that jumps out at you the moment you read the explanation.
You fix the assertion in your criteria library, then run the same scorecard again on the same conversations. This is why re-evaluation explicitly accepts a scorecard that has already been used: it is not a tolerated edge case, it is the main way of working.
Because each run creates a distinct evaluation, you keep the successive versions side by side: you see the effect of your correction instead of assuming it.
A good habit for your trials. Test evaluations count towards your indicators like any other. Once your scorecard has settled, deactivate the trial evaluations: they disappear from reporting and exports while remaining available for consultation. You can also set them aside on the fly with the conversation list filter described below.
Overhauling your quality framework without flying blind
The short cycle applies to a single criterion. The same mechanism changes the nature of a far heavier project: a full scorecard overhaul.
Overhauling a scorecard means accepting a discontinuity in your quality monitoring programme. Scores from before and after are no longer comparable, and the question that comes straight back from the leadership team is always the same: did our advisors get worse, or is the scorecard tougher? Without a point of comparison, nobody can answer.
Re-evaluation lets you prepare that switch with numbers. Take a sample of conversations already scored with the old scorecard, re-score them with the new one, and read the gap. A fifteen-point gap on the same sample does not mean the floor has declined: it means your new framework is fifteen points more demanding. You know your new baseline before you have switched anything.
This works because re-evaluation overwrites nothing. The old score stays intact, your reporting history doesn't move, and both evaluations carry the date of the original call, which keeps them comparable over the same period.
An example: short calls
On an outbound floor, a considerable share of calls last under a minute: voicemails, gatekeepers, wrong numbers, immediate refusals. Scoring them with a full interview scorecard makes no sense. Almost every criterion is either not applicable or mechanically failed, and the advisor is penalised for a situation they do not control.
Building a short scorecard for these calls is a typical partial overhaul. And the question that decides everything is empirical: on my sub-one-minute calls, does this twenty-criteria scorecard produce a score that means anything? The only way to find out is to run it on real short calls, already analysed, whose verdict under the long scorecard you already know.
How do you actually re-evaluate a conversation?
From a conversation's page, in the action bar of the Analysis block, a branch icon opens the re-evaluation window. There you pick the target scorecard among all the active scorecards in your organisation, with their criteria count. Scorecards already used by another evaluation of the conversation remain available, simply flagged as such.

A justification comment is mandatory. It is not there for form's sake: a re-evaluation creates an additional score and consumes analysis credits. That comment explains why this second pass was requested, and it remains available afterwards, which is valuable when you chain trials on a scorecard under construction.
The new evaluation appears immediately, its criteria under analysis, with a countdown that refreshes the screen on its own. A few dozen seconds later, the score is there.
The button stays greyed out until the original analysis is entirely finished: a re-evaluation carries over the conversation's summaries and tags, so they need to exist first.
What is not replayed, and what that changes
This is what makes a cycle of trials economically reasonable.
| Analysis stage | Replayed? | Why |
|---|---|---|
| Transcription | No | The text of the exchange does not depend on the scorecard |
| Summaries and sentiment | No | Copied over as they are, with no new model call |
| Tags | No | They are configured at organisation level, not scorecard level |
| Prosody | No | The measurement applies to the audio, not to the framework |
| Criteria and assertions | Yes | This is precisely what the scorecard defines |
A re-evaluation therefore costs only the assertions of the chosen scorecard. On a short scorecard of around twenty criteria, that is a fraction of the cost of a full analysis, and nowhere near a re-import that would pay for the transcription all over again. Three rounds of trials on ten reference conversations remain a modest expense compared with a framework you will use for the next two years.
Each evaluation carries its own billing, available in its dedicated panel.
Two evaluations, two lives
A re-evaluation is not a variant, nor an annex. It is an evaluation in its own right.
It has its own score, computed on the criteria of its scorecard. It has its own debrief, which the manager opens and closes independently. It can be contested by the advisor just like the other one, because a score that counts must be open to a right of reply. It has its own version history. And it can be deactivated without affecting the evaluation it came from, which is exactly what you want to do with a series of trials.
That autonomy extends to the files: the re-evaluation owns its copy of the recording and of the transcript. A purge applied to one can never cut into the other.
A banner links the evaluations of a same conversation, each with its date, its scorecard, its score and the justification behind it. One click moves from one to the other.

A conversation whose recording has been purged remains re-evaluable. Purging the audio and purging the transcript are two distinct settings. As long as the transcript exists, the analysis can be replayed: it works on the text, not on the sound.
Reading your reporting during a trial phase
Two evaluations count as two in your indicators. During a scorecard testing campaign, that is an effect to be aware of.
Two safeguards. Each evaluation reports under its own scorecard, so a filter by scorecard naturally separates your trials from your day-to-day management. And the conversation list offers a two-group filter: debrief statuses on one side, evaluation scope on the other, with four choices crossing original versus re-evaluation, active versus deactivated. You display what you want to measure, and the export takes exactly the same scope.
In the list, a re-evaluation is recognisable by its branch icon, in place of the usual progress gauge.

Traceability, by design
Re-evaluation joins the family of human supervision gestures that frame quality monitoring in Raisetalk: modifying an evaluation, contesting it, deactivating it, debriefing it. They all follow the same principle: the AI proposes, the human decides, and the platform keeps track of who decided what.
Here, that record takes three forms. The original evaluation is never modified, whatever happens to the new one. The justification is archived as the first version of the re-evaluation, with its author and its timestamp. And that first version holds a complete snapshot of the evaluation it came from: opening the history of a re-evaluation shows you where it came from, in the exact state it was in at the time.
Re-evaluation is finally protected by a dedicated permission, "Re-evaluate a conversation", in the Conversations family of the roles and permissions matrix. It is granted to administrators by default, and you open it to other roles as your organisation requires.

Getting started
- Pick three to five reference conversations whose expected verdict you already know.
- Open one of them and click the branch icon in the action bar of the Analysis block.
- Select the scorecard you are building, and note in a few words the hypothesis you are testing.
- Read the justifications criterion by criterion, not just the score. Fix your assertions, then run the same scorecard again.
What about tomorrow: several scorecards from import?
One question comes up as soon as you handle several scorecards on the same conversation: why not apply two frameworks to every call from the start, one commercial and one compliance-oriented, for instance?
It is a fair question, and re-evaluation is not the answer to it. A systematic need calls for a systematic answer: the ability to designate several scorecards at integration or import time, so that every conversation arrives with both readings, with no manual gesture. It is worth adding that a single well-built framework often covers two dimensions perfectly well, and that splitting them requires real reasons: two distinct audiences and two separate action plans.
Re-evaluation answers a different and perfectly complementary need: the one-off gesture, decided after the fact, on a specific conversation. Putting a scorecard to the test, preparing an overhaul, targeting a debrief, correcting a mistake. That is already a great deal, and it is available today.

