We did not set out to build a coaching product. We built a way for people to rehearse difficult conversations with an AI that talks back and pushes back. In the industry this is called roleplay, which is a slightly unfortunate name for something quite serious: you practice the conversation before you have to have it for real, and you get a score, detailed report, and a transcript at the end.
Then customers started using it for something we had never sold them, and telling us about it afterwards. It happened often enough, and consistently enough, that it has now reshaped what we are building.
So this is a piece about how that happened, why demand for AI coaching showed up faster than most people expected, and what it has actually been worth to the businesses that got there early.
Why AI Coaching Has a Cost Problem Worth Solving
Coaching works. Nobody in corporate learning seriously argues otherwise. A good coach watching you have a real conversation, then telling you what to do differently, is the most reliable way anyone has found to change how a person behaves at work.
The problem has never been efficacy. It has been price.
A human coach costs hundreds of dollars an hour. At that rate you can coach your senior leadership team, and if the budget is generous you can go a few hundred people deeper. What you cannot do, at any budget I have ever seen, is coach the tens of thousands of people in a large company who actually speak to customers every day. The field technician on a doorstep. The agent on a support line. The rep in a store on a Saturday afternoon.
Which means coaching gets spent almost entirely at the top of the building, and the conversations that decide whether a customer stays or goes happen at the bottom of it. That is a strange way to allocate a resource, and everyone involved knows it. They just did not have another option.
The interesting thing about a cost curve is not the saving. It is what becomes possible on the other side of it.
AI Coaching in Practice: A Telecoms Case Study
I want to be specific here, because generalities in this market are cheap.
A few years ago a large telecoms operator came to us with a single compliance use case. Small, contained, easy to say no to. Their situation was the one you find in almost every incumbent business: the CEO had announced a customer-first strategy, it had landed beautifully with the executive team, and then it had died somewhere on the road to the frontline. Their field technicians were genuinely excellent engineers. Not one of them had ever been given an hour of preparation for an angry customer on a third repeat visit to the same fault.
Within a year that one use case had turned into five functions. Field readiness. How network operators communicate on repeat incidents. Sales enablement. New-hire onboarding. And, entirely without us, virtual simulations of high risk field work like ladder safety and gas detection, which their own operations team built from their own incident data.
That last part matters more than it sounds. They were not waiting for a vendor to ship them content. They had built an internal capability, the kind that normally takes an enterprise two or three years to develop, and they had done it in months. When the business later went through a major structural change, the sort of moment where software contracts usually get quietly cancelled, the platform did not get rationalised away. It became the template. An internal champion took a year of real usage data into the executive suite and made the case to extend it across the whole customer-facing workforce.
Nobody at that company describes what they bought as a training tool anymore.
From AI Roleplay Scoring to AI Coaching: What the Data Showed
The signal that we were in a different business than we thought came from how people used the scored sessions.
A score tells you what happened. It does not tell you what to do differently on Tuesday morning. Customers started asking for the conversation after the conversation, the debrief: what went wrong, why it went wrong, what to try next time, and a system that remembered the last four attempts so the advice actually built on itself.
That is not roleplay. Roleplay is an event with a beginning and an end. Coaching is a relationship with memory and a point of view about where you are going.
At the same time the use cases were drifting commercial, which we did not plan for either. Sales enablement was not our pitch. It was the thing that spread fastest inside accounts, because the value is embarrassingly easy to see. A rep who has run thirty practice conversations against a well-modelled sceptical buyer sells differently than a rep who has run none.
In one multinational pilot, hundreds of sales reps across three regions voluntarily completed thousands of sessions. Thirty-two each, on top of their day jobs, with nobody making them. Average scores went from 56% on first attempt to 67% on best attempt. In comparable contact centre work we have seen people reach competency 20 to 30% faster, and in one financial services deployment the QA supervision hours per new agent dropped by up to 80%.
Of all of those numbers, the one I trust most is thirty-two. Nobody voluntarily does thirty-two sessions of anything they consider training.
How AI Coaching Is Transforming Leadership Development
Earlier this year we were brought into a large enterprise alongside a leadership development partner. The senior leadership cohort was working through a set of behavioural shifts, the most important being a move away from inward-looking measures of success and toward decisions grounded in what customers actually say.
The program itself was a series of facilitated sessions where leaders take a real, live business decision and rework it. Our job was to put the customer in the room. Not a slide about the customer, and not a research summary about the customer. An AI persona the leaders could interrogate directly: a long-tenured residential customer who is tired of a reliability problem, and a business buyer who is openly weighing up alternatives. Both of them built to push back the moment someone in the room reached for an internal metric to justify a decision.
Two things about that engagement changed how I think about our product.
The first is that the session lasts two hours and the persona does not leave. Leaders keep access afterwards and bring the customer into their own team meetings. The scarce resource in any development program is the facilitator’s time, and this is the first mechanism I have seen that genuinely separates the practice from it. The programme stops being a day in a hotel and becomes something people can use on a Wednesday afternoon.
The second is that every one of those conversations is scored against the behaviors the program is trying to shift. Which means you can finally answer the question every executive sponsor asks about leadership development and never really gets an answer to. Is it landing? Not “did people enjoy it”, which is what a feedback form measures, but is the behaviour actually changing, with evidence at the level of the whole cohort. In my experience that evidence is what turns a pilot into a mandate.
And then customers keep asking the harder version of the question. If you can score a practice conversation, can you score the real ones against the same standard? There are thousands of them every day and nobody reads them.
That is exactly where we are taking this.
What Enterprise AI Coaching Looks Like in 2026
Five themes, and I am happy to be held to them.
A coaching experience rather than a scoring experience. Coaching frameworks built in, a coaching summary in place of a score, a knowledge bank the coach can draw on, and session memory so it knows your history. Coach agents that can debrief a learner on any scored session, and that turn up where teams already work instead of in one more portal nobody logs into.
One standard for practice and for reality. You build an evaluation once, keep it in one place, attach it to as many scenarios as you like, and use the same standard to score real conversation transcripts. Rubrics, meaning the written criteria you are marking against, can be generated from frameworks you already have rather than written from scratch. Practice and real performance end up in one comparable picture, which as far as I know nobody has properly had before.
Scoring your QA team can audit. Every judgement comes with a verbatim quote from the transcript, so you can see what the score was based on. Calibration against your own team’s scored examples, so the AI learns your standards instead of a generic idea of good. And repeatability, meaning the same transcript gets the same score run after run. AI scoring is only useful if a human being can challenge it and win.
Personas with real range. A library of distinct characters, frustrated and sceptical and passive-aggressive and coldly analytical, with the surface details varied so learners cannot memorize their way through, and difficulty that holds up under pressure. Plus a lighter text-based version for the very high volume cases where video is more than you need.
Insight across the whole population. Automatic summaries across every evaluation, sliced by team, site, language or cohort. Not so you can rank people. So you can see where coaching is genuinely needed and send your humans there.
Some of that is live today, including natural realtime voice conversations, hints that appear while a learner is speaking, custom avatars, and forty-plus languages. The rest is in active development.
The Real Value of AI Coaching at Scale
AI coaching does not replace your best coaches. It makes them affordable to deploy, which is a different and better claim.
The repetition, the measurement and the scale get carried by something with almost no marginal cost, which means the scarce thing, human judgement, gets concentrated where it changes an outcome. Your best coach stops spending Tuesday listening to a rep fumble an opening line for the first time, because the rep has already fumbled it forty times in private and worked it out.
The companies that figured this out first did not do it because a vendor convinced them. They did it because they had tens of thousands of people having conversations they could not see, could not measure, and could not improve.
That is not a training problem. It is a performance problem that happened to have no tooling until fairly recently.
Kurt Kratchman – connect with me!
Virti is an AI-powered performance and coaching platform used by enterprise teams to build measurable behavioural change through realistic practice. If you want to see what that looks like in your industry, request a demo.
