Skip to content
RBRitwik Bisht

002 / Compliance call evaluation

Koodoo Evaluate

Koodoo Evaluate scored mortgage compliance calls from the recording alone. A structured feedback loop with the compliance team kept surfacing the same complaint: scores came back wrong, on checks that had clearly happened somewhere the tool wasn't looking.

LIVE / COMPLIANCE CALL EVALUATIONINTERACTIVE
CONVERSATION / SOURCE
  • 0:00Call opens
  • 2:14Affordability discussion
  • 6:40Call ends
  • Training and monitoring disclaimerconfirmed
  • Initial rate typeunconfirmed

One check can't confirm on voice alone. Switch chat back in — that's where the answer was.

A working miniature of the interaction. The frames below are from the project itself.

PLATE 01 / EVALUATEREPORT
Koodoo Evaluate's report page. A conversation named 'Conversation Name' with two failed compliance checks on the left — a training and monitoring disclaimer, and an initial rate type — each collapsible with a comments thread, and the entire conversation transcript on the right with an audio scrubber pinned along the bottom.
A failed check reads as the rule and the reason, next to the transcript it was read from. Comments and the audio scrubber sit one layer down, not hidden behind a summary.

CONTEXT

Koodoo Evaluate scored a call. It never saw the conversation that continued after the call ended.

The compliance team ran two feedback loops: structured check-ins and unstructured notes dropped into Slack. Both kept surfacing the same complaint, that evaluations were coming back inaccurate, often on checks that had clearly happened somewhere the tool wasn't looking.

Two outcomes came out of those sessions as the ones that actually moved evaluation scores: trust and accuracy, and how fast a reviewer could pull a conversation to start with. Speed lost that trade before the project began.

FROM THE BRIEF

We could improve the time to grab conversations, but if we were grabbing and processing conversations inaccurately, then the value would be low.

PROBLEM / 01

The tool was told to score a call. Compliance officers were following a conversation.

A single client interaction routinely crossed channels: a call, then a follow-up over chat minutes after it ended. Koodoo Evaluate only ever looked at the call.

PROBLEM / 02

A coverage gap read as a tool failure, not a data failure

When a check came back unconfirmed because the answer lived in a channel nobody had uploaded, the compliance officer's takeaway was that the tool couldn't be trusted, not that it was missing an input.

PROBLEM / 03

Ordering a mixed conversation isn't just a timestamp sort

The working assumption was that voice and non-voice could be merged by time. Testing found conversation order was more complex than that once a real multi-channel, multi-call thread was in front of a reviewer.

GOALS

Before I drew a single screen, I had to know what success looked like.

Success criteria came before any design, so the work had something to be judged against other than taste.

COMMITTED TO

  1. 01Support non-voice communication, chat and Slack, as first-class input to an evaluation, not an afterthought
  2. 02Prove accuracy actually improves with coverage before building anything past a manual upload
  3. 03Give a reviewer confidence about which channels were included in a score, not just the score itself
  4. 04Order a mixed voice and non-voice conversation in a way a reviewer can actually follow

RESEARCH

Before touching the design, we went looking for where trust in the tool was actually breaking.

The hypothesis: conversational accuracy was capped by coverage. If non-voice channels carried compliance-relevant information the tool never ingested, no amount of better voice analysis would close the gap.

Ten side-by-side comparisons tested it directly: the same conversations scored by a human reviewer against Koodoo Evaluate, voice-only versus voice-plus-chat, to see whether coverage moved accuracy or just moved the number.

THE ASSUMPTION

50%
Of conversations in the discovery sample included a non-voice component the tool never ingested

The tool matched the reviewer less often on voice alone

One completed comparison: the reviewer scored a call at 95%, Koodoo Evaluate at 83% on voice alone, flagged as not matching. The note on the mismatch: a compliance check read as passed because the information that would have failed it had moved to chat.

Multi-call conversations broke the tool's own logic

A second opportunity surfaced independently of coverage: conversations spanning more than one call confused the evaluation logic regardless of channel, which meant coverage was necessary but not sufficient.

IDEATION / 01

I let people manually attach the half of the conversation the tool couldn't reach.

Rather than building a live chat integration up front, a manual upload path let compliance officers attach non-voice communications to a conversation before running the test that would prove whether coverage was even worth building further.

IDEATION / 02

I tested the hypothesis before designing around it.

Ten side-by-side comparisons ran ahead of any interface work, on the same conversations, coverage on and off. The design that followed was aimed at a validated gap, not an assumed one.

IDEATION / 03

I made channel visible, not just content.

Reviewers needed to know which parts of a conversation were voice and which were chat before they could trust a score built from both, so the interface marks the source on every segment rather than presenting one merged transcript.

EVIDENCE

The interface, screen by screen

Screens and research artifacts from the project itself, not recreations.

PLATE 02 / OUTCOMESPRIORITISED
Two outcome cards side by side: 'Trust and accuracy', marked chosen with a green checkmark, and 'Reduce time to grab conversations', left unmarked.
Speed lost to accuracy in the room. A slower, correct answer was worth more than a fast, wrong one.
PLATE 03 / RESEARCHDISCOVERY BOARD
A research board: experience-mapping notes on the left, five opportunity-discovery interview clusters in the middle (product and customer-ops interviews), an opportunity solution tree and a known/unknown, important/unimportant prioritisation grid, and a column of individual research questions on the right.
Every opportunity downstream of this project traces back to a labelled cluster on this board, not a hunch.
PLATE 04 / OPPORTUNITYCOVERAGE
Three opportunities branching from 'Trust and accuracy'. Opportunity 1, bolded: coverage was limited, and fifty percent of conversations were non-voice. Opportunity 2: a multi-call issue was feeding the logic inaccurate conversations. Opportunity 3: validate data against Acre.
The bolding is in the original board. Half the evidence a reviewer needed was arriving in a channel the tool never opened.
PLATE 05 / TEST10× COMPARISON
A side-by-side comparison spreadsheet, voice-only against voice-plus-non-voice, ten rows deep. The one completed row shows a human reviewer's score of ninety-five percent against Koodoo Evaluate's eighty-three percent on voice alone, marked as not matching, with notes on exactly which checks the tool missed and why.
The mismatch has a named cause in the notes column: the information the check needed had moved to chat. Not a modelling problem, a coverage problem.
PLATE 06 / USABILITYSESSION
A usability-testing session in progress: the Figma prototype of Evaluate's report page under a browser toolbar reading 'Evaluate → Working file' and a Share prototype button, with a video-call window open in the corner.
Tested live with the compliance team who'd have to trust the score, not just reviewed internally.

THE PIVOT

Testing agreed with most of the plan and broke one assumption we didn't know we'd made.

Three tasks structured the session with the compliance team: basic navigation, using the audio player, and moving through a transcript that now mixed voice and chat. Each question was aimed at one of the assumptions the design had been built on.

ASSUMED

Coverage plus time-ordering would be enough

Non-voice channels carried compliance-relevant information, and merging by timestamp would keep a mixed conversation readable.

VALIDATED

Coverage was the right bet

Non-voice information mattered, most conversations actually had some, and reviewers needed to see which channel a segment came from before they'd trust the score.

CHALLENGED

Time-ordering wasn't enough on its own

Ordering a mixed conversation by timestamp alone undersold how confusing a real multi-channel thread reads. Conversation order needed more structure than a sort.

FINAL SOLUTION

Three pieces, in service of one question: does coverage change accuracy?

  • Manual upload for non-voice communications, attached to a conversation ahead of evaluation
  • A conversation view marking every segment by channel, voice or chat, instead of one merged transcript
  • A ten-comparison side-by-side test rig used to validate coverage against a human reviewer's score

IMPACT

TODO
Accuracy delta, voice-only vs voice-plus-chat, after the fix
TODO
Share of conversations now including a non-voice upload

Numbers pending employer sign-off. A placeholder beats a figure I cannot source.

REFLECTION

Accuracy has a coverage problem before it has a modelling problem

The instinct when a score looks wrong is to fix the scoring. Here the scoring was fine; the input was half of what it should have been. Validating that with ten side-by-sides before designing anything meant the manual upload wasn't a bet, it was the smallest thing that could prove the bet right. The harder lesson came after: once coverage stopped being the gap, ordering did. Fixing what one assumption gets wrong usually just uncovers the next one.

THE LINE THAT MATTERED

The scoring was fine. The input was half of what it should have been.