Difficult Conversations

The Practice Paradox: Measure Capability Without Turning Training Into Surveillance

October 4, 2026

If professional training is meant to expose weaknesses, what happens when every weakness becomes visible to management?

That is the tension inside a new generation of AI coaching systems.

The technology can analyse more conversations, identify repeated patterns and give people feedback that was previously too expensive to provide. That is valuable.

The same technology can also create a permanent record of hesitation, mistakes, weak moments and low-confidence behaviour.

At that point, a development system can start to feel like a performance-monitoring system.

The distinction matters because people behave differently when practice feels safe to fail.

The same analysis can support coaching or surveillance

A recent dispute at training company Multiverse makes the issue unusually concrete.

On 25 September 2026, Multiverse published an explanation of Compass, its AI-supported instructor-development system. The company describes the problem it is trying to solve: managers could previously observe only a small sample of instructor sessions, making feedback inconsistent and expensive to scale. Compass analyses sessions more broadly so managers can identify patterns and target development.

Three days later, The Guardian reported concerns from instructors who described the experience very differently. They reported stress from continuous monitoring and concern about AI-generated assessments being visible to managers. Multiverse said the system does not make performance-management decisions and that human managers remain responsible.

Those accounts are not mutually exclusive.

A system can genuinely be designed to improve coaching and still create a surveillance experience for the person being observed.

That is why the governance question cannot be reduced to whether a human remains "in the loop."

The more important question is what evidence is collected, who can see it and what consequences can follow from it.

Better measurement changes behaviour

Our article on why training completion does not prove capability argues that observed performance is more useful than course completion when the goal is to understand whether someone can actually execute a skill.

That creates an obvious follow-on problem.

If every observed practice attempt becomes management evidence, the practice environment changes.

Suppose a buyer knows that every simulation is scored and that every low score is visible to their manager.

Which scenario will they choose?

The difficult supplier who threatens continuity?

Or the easier scenario they already know how to handle?

Will they test an unfamiliar response?

Or use the safest behaviour that protects the score?

Will they expose the weakness they most need to work on?

Or avoid creating evidence that the weakness exists?

The measurement system can become accurate about the wrong thing.

It starts measuring how well people protect themselves inside the training environment.

Practice requires permission to be temporarily bad

Deliberate practice works because people attempt something they cannot yet do reliably.

That means failure is not an exception.

Failure is part of the mechanism.

A buyer working on concession discipline may give something away too quickly in the first attempt.

A manager practising difficult feedback may soften the message until it becomes unclear.

A category manager rehearsing an escalation may lose structure when the simulated supplier becomes aggressive.

Those moments are useful because they identify exactly what should be practised next.

If each one becomes durable employee-performance evidence, the rational incentive changes.

The learner is no longer asking:

"What do I need to expose so I can improve it?"

They are asking:

"What should I avoid exposing because somebody may judge me for it?"

A development environment that punishes visible weakness will eventually become a place where people demonstrate strengths they already have.

Managers do need evidence

The answer is not to hide everything from leadership.

A procurement leader needs to know whether a capability programme is being used.

They need to know whether the team is improving.

They may need to know that a common weakness exists across the function, such as unreciprocated concessions, poor questioning after supplier pushback or loss of structure during escalation.

Without evidence, capability investment becomes difficult to manage.

The important distinction is between evidence about the system and evidence about the individual attempt.

A manager can reasonably see that:

  • 70 percent of the team has practised a price-increase scenario this quarter,
  • handling of continuity threats is a recurring team-level weakness,
  • repeated practice is reducing early concessions,
  • one category has had very little practice before a major renewal window.

None of that requires a permanent management feed showing that Buyer A froze at minute six, Buyer B revealed a target too early and Buyer C had a poor session on Friday afternoon.

The leadership question is whether the capability is developing.

The learner question is exactly where their own behaviour needs work.

Those views do not have to be identical.

Aggregate visibility can be more useful than individual surveillance

There is a common assumption that more granular management data must be better.

For learning, it can be worse.

A manager with access to every transcript and every score can easily overinterpret noise.

One weak simulation can reflect an unfamiliar scenario, a bad day, an experimental technique or a deliberate attempt to stretch beyond current competence.

Aggregated patterns are often more decision-useful.

If ten buyers repeatedly concede after a continuity threat, that is a capability-design signal.

If one buyer does it once, it may mean very little.

This is also why longitudinal development matters more than ranking. The question is not who had the highest score this week. It is whether a repeated weakness becomes more stable after targeted practice.

Training data should have a declared purpose

A useful governance rule is simple:

Do not collect practice data first and decide what to do with it later.

The purpose should be clear before the learner starts.

Is the conversation being analysed for private feedback?

Can the manager see the detailed transcript?

Can individual scores be used in formal performance management?

How long is the data retained?

Can the learner practise without creating a permanent record?

Can aggregated team patterns be separated from identifiable individual analysis?

These questions affect behaviour inside the training environment, not only privacy compliance.

The learner needs to know whether the system is a coach, an assessor or both.

Ambiguity creates the worst combination: people behave as if they are being assessed even when the organisation says the tool is mainly for development.

Assessment and development can be separated

There are situations where formal assessment is legitimate.

A company may want to verify that somebody can perform a safety-critical procedure, meet a regulatory requirement or demonstrate a role-specific skill before taking on responsibility.

That is different from everyday practice.

The distinction should be designed, not implied.

A development mode can prioritise private feedback, repetition and experimentation.

A formal assessment mode can make the criteria, audience, consequences and evidence requirements explicit.

Mixing the two weakens both.

If practice feels like assessment, people stop experimenting.

If assessment feels like informal practice, the evidence may not be rigorous enough for the decision being made.

What this means for negotiation capability building

Negotiation training is particularly sensitive because the behaviours worth improving are often the ones people least want managers to see.

Freezing after an aggressive threat.

Conceding without getting anything back.

Talking too much because silence feels uncomfortable.

Failing to challenge a senior-sounding counterpart.

Losing the plan after an unexpected objection.

These are exactly the behaviours a realistic simulation should surface.

They are also exactly the behaviours employees may hide if the training environment becomes a reputation-management exercise.

Voice2Evolve separates those purposes deliberately. Detailed individual analysis is primarily for the learner. Leadership can see usage and higher-level development patterns without needing a permanent feed of every person's mistakes.

That is not only a privacy decision.

It is a learning-design decision.

If the goal is stronger behaviour under pressure, people need a place where weak behaviour can become visible before it becomes costly.

The organisation still needs evidence that capability is developing.

But it does not need to turn every failed repetition into an employee record to get it.

Sources

Why training falls shortWhy Traditional Negotiation Training Falls Short

Train the moment, not the theory.

Voice2Evolve puts you in the scenario repeatedly until your reaction under pressure is no longer panic.