002 / Compliance call evaluation
Koodoo Evaluate
Koodoo Evaluate scored mortgage compliance calls from the recording alone. A structured feedback loop with the compliance team kept surfacing the same complaint: scores came back wrong, on checks that had clearly happened somewhere the tool wasn't looking.
- 0:00Call opens
- 2:14Affordability discussion
- 6:40Call ends
- Training and monitoring disclaimerconfirmed
- Initial rate typeunconfirmed
One check can't confirm on voice alone. Switch chat back in — that's where the answer was.
A working miniature of the interaction. The frames below are from the project itself.

CONTEXT
Koodoo Evaluate scored a call. It never saw the conversation that continued after the call ended.
The compliance team ran two feedback loops: structured check-ins and unstructured notes dropped into Slack. Both kept surfacing the same complaint, that evaluations were coming back inaccurate, often on checks that had clearly happened somewhere the tool wasn't looking.
Two outcomes came out of those sessions as the ones that actually moved evaluation scores: trust and accuracy, and how fast a reviewer could pull a conversation to start with. Speed lost that trade before the project began.
FROM THE BRIEF
We could improve the time to grab conversations, but if we were grabbing and processing conversations inaccurately, then the value would be low.
PROBLEM / 01
The tool was told to score a call. Compliance officers were following a conversation.
A single client interaction routinely crossed channels: a call, then a follow-up over chat minutes after it ended. Koodoo Evaluate only ever looked at the call.
PROBLEM / 02
A coverage gap read as a tool failure, not a data failure
When a check came back unconfirmed because the answer lived in a channel nobody had uploaded, the compliance officer's takeaway was that the tool couldn't be trusted, not that it was missing an input.
PROBLEM / 03
Ordering a mixed conversation isn't just a timestamp sort
The working assumption was that voice and non-voice could be merged by time. Testing found conversation order was more complex than that once a real multi-channel, multi-call thread was in front of a reviewer.
GOALS
Before I drew a single screen, I had to know what success looked like.
Success criteria came before any design, so the work had something to be judged against other than taste.
COMMITTED TO
- 01Support non-voice communication, chat and Slack, as first-class input to an evaluation, not an afterthought
- 02Prove accuracy actually improves with coverage before building anything past a manual upload
- 03Give a reviewer confidence about which channels were included in a score, not just the score itself
- 04Order a mixed voice and non-voice conversation in a way a reviewer can actually follow
RESEARCH
Before touching the design, we went looking for where trust in the tool was actually breaking.
The hypothesis: conversational accuracy was capped by coverage. If non-voice channels carried compliance-relevant information the tool never ingested, no amount of better voice analysis would close the gap.
Ten side-by-side comparisons tested it directly: the same conversations scored by a human reviewer against Koodoo Evaluate, voice-only versus voice-plus-chat, to see whether coverage moved accuracy or just moved the number.
THE ASSUMPTION
- 50%
- Of conversations in the discovery sample included a non-voice component the tool never ingested
The tool matched the reviewer less often on voice alone
One completed comparison: the reviewer scored a call at 95%, Koodoo Evaluate at 83% on voice alone, flagged as not matching. The note on the mismatch: a compliance check read as passed because the information that would have failed it had moved to chat.
Multi-call conversations broke the tool's own logic
A second opportunity surfaced independently of coverage: conversations spanning more than one call confused the evaluation logic regardless of channel, which meant coverage was necessary but not sufficient.
IDEATION / 01
I let people manually attach the half of the conversation the tool couldn't reach.
Rather than building a live chat integration up front, a manual upload path let compliance officers attach non-voice communications to a conversation before running the test that would prove whether coverage was even worth building further.
IDEATION / 02
I tested the hypothesis before designing around it.
Ten side-by-side comparisons ran ahead of any interface work, on the same conversations, coverage on and off. The design that followed was aimed at a validated gap, not an assumed one.
IDEATION / 03
I made channel visible, not just content.
Reviewers needed to know which parts of a conversation were voice and which were chat before they could trust a score built from both, so the interface marks the source on every segment rather than presenting one merged transcript.
EVIDENCE
The interface, screen by screen
Screens and research artifacts from the project itself, not recreations.





THE PIVOT
Testing agreed with most of the plan and broke one assumption we didn't know we'd made.
Three tasks structured the session with the compliance team: basic navigation, using the audio player, and moving through a transcript that now mixed voice and chat. Each question was aimed at one of the assumptions the design had been built on.
ASSUMED
Coverage plus time-ordering would be enough
Non-voice channels carried compliance-relevant information, and merging by timestamp would keep a mixed conversation readable.
VALIDATED
Coverage was the right bet
Non-voice information mattered, most conversations actually had some, and reviewers needed to see which channel a segment came from before they'd trust the score.
CHALLENGED
Time-ordering wasn't enough on its own
Ordering a mixed conversation by timestamp alone undersold how confusing a real multi-channel thread reads. Conversation order needed more structure than a sort.
FINAL SOLUTION
Three pieces, in service of one question: does coverage change accuracy?
- Manual upload for non-voice communications, attached to a conversation ahead of evaluation
- A conversation view marking every segment by channel, voice or chat, instead of one merged transcript
- A ten-comparison side-by-side test rig used to validate coverage against a human reviewer's score
IMPACT
- TODO
- Accuracy delta, voice-only vs voice-plus-chat, after the fix
- TODO
- Share of conversations now including a non-voice upload
Numbers pending employer sign-off. A placeholder beats a figure I cannot source.
REFLECTION
Accuracy has a coverage problem before it has a modelling problem
The instinct when a score looks wrong is to fix the scoring. Here the scoring was fine; the input was half of what it should have been. Validating that with ten side-by-sides before designing anything meant the manual upload wasn't a bet, it was the smallest thing that could prove the bet right. The harder lesson came after: once coverage stopped being the gap, ordering did. Fixing what one assumption gets wrong usually just uncovers the next one.
THE LINE THAT MATTERED
The scoring was fine. The input was half of what it should have been.