AI-assisted conversation analysis

Turning thousands of recorded sales calls into evidence managers could coach from

When
Jan–Sep 2026
Client
Ponder, across client work
Method
Research design and intake, traceability, answer-key evaluation
My role
Co-designed from the research side; about 30 pipelines run

Working with my business partner, I co-designed a research pipeline that reads hundreds of customer conversations per question. Every finding links back to the exact moment in the conversation where it came from. Runs went from 5 a month to 50–80 once standing questions ran every night. The output now feeds regular coaching reports, win-loss research and an automated, personalised referral request sent to customers in the rep's name when the rep missed the moment to ask.

The situation

Ponder helps B2B companies learn from the conversations their Sales and Client Success teams already record: calls, texts and emails. Our clients had thousands of them and a CRM that showed where deals stalled but not why. A language model can read every conversation, but its output sounds equally confident whether or not it's right, and a sales manager won't coach a rep on a claim they can't check. My business partner built the pipeline. I co-designed it from the researcher's side: I set what it had to deliver, ran it on client work and reported from it, and fed what I found back to him, so each version improved on the last.

What I decided

Agree the questions and freeze the scope before anything runs. It's tempting to point a pipeline at all the data and see what comes out. I designed an intake step instead: before any run, we agree with the client who the output is for, what decision it serves, which conversations are in scope and what to look for, and we write the scope down and freeze it. I also ask for hypotheses outright, the client's and my own, rather than leaving the analysis to infer them. I added that after a win-loss study where the AI-assisted analysis had worked from ideas it picked up along the way, never asked me for mine, and so tested only part of what mattered. Freezing scope matters because early counts move: on one study, the share of losses put down to one reason fell by about a third once the scope was agreed.

Make every finding point back to the moment it came from. The pipeline could have summarised themes across the whole set of conversations, which reads well but can't be checked effectively at scale. I made traceability a requirement: each finding carries its category, its strength and a link to the exact passage in the conversation. That is what lets a manager open the call at the right second, and what let me check the pipeline's output by hand, one moment at a time.

Where the method sits
The pipeline my partner built, and the points where I set the rules and checked the output.

Measure prompt changes against an answer key, not by eye. At first, prompt changes were judged by reviewing output by hand, which can't tell a real improvement from noise. I moved us to hand-labelled answer keys: a fixed set of conversations, drawn on mechanical properties rather than on what looked wrong, where I had reviewed each one and decided the right answer, scored in both directions. Both directions matters. Most of the fixes suppress false alarms, and on a set of flagged moments alone, a pipeline that flags nothing scores perfectly. Every change now runs on a copy of the live pipeline first and has to beat the last version on the key before it goes live.

The answer-key loop
How a prompt change is tested before it reaches a client.

Where it was harder than planned. My manual reviews showed that the detectors behind the coaching and referral work were over-flagging, and not for one reason. Some moments the pipeline marked as missed had been handled, but just outside the passage it looked at. Some were never the kind of moment it was looking for. Others were missed before any prompt could see them, so no prompt change could recover them. I now stop tuning when changes stop improving the score and hand the problem to my partner as a build change, with the labelled cases attached.

Time, depth, team

January to September 2026. About 30 pipelines run by me, for clients on calls, texts and emails. Me: research design and intake, traceability and evaluation requirements, answer keys, running pipelines, reporting, and turning output into coaching, marketing and decisions at the client. Business Partner: pipeline design and implementation, output review.

What the organisation did with it

What I'd do differently. I could change prompts; only my partner could change code. From June to September most of his time went into building the client automations, so fixes that needed a build change waited: the misses no prompt could reach, and a record of which model version produced each finding. That second gap mattered when the main pipeline had to change model mid-year and the earlier months were never rescored, so month-to-month movement in 2026 can't be read as a change in behaviour; the coaching page claims none for that reason. Next time I'd agree up front a standing share of build capacity for evaluation fixes, stamp every finding with the version that produced it, and rescore an overlap sample before any model switch.

← All case studies