← Back to insights

BI & Analytics

Is AI Making Good Work Harder to Measure?

Is AI Making Good Work Harder to Measure?

AI can make people faster and raise the quality of visible work, yet that creates a management problem we rarely discuss: the result may now tell us less about the capability behind it. As AI takes on more execution, judging performance may depend less on counting output and more on understanding the human judgment that made the work worth trusting.

Download 1-page brief

A while ago, I wrote that as data and AI become more embedded in work, people's value should move away from routine task execution and closer to the outcomes they create. I still believe that, although a recent AI-enabled BI project made me add an important qualification because the answer a user sees can now appear in seconds, while much of the work that makes that answer reliable has already happened somewhere else.

In that project, I deliberately kept KPI definitions, business rules, query boundaries and validation checks in a governed layer, while AI was used to interpret results, explain patterns and continue the conversation around them. From the user's side, the experience can feel almost immediate, yet the quality of the result still depends on decisions about context, data meaning, where the model is allowed to operate and when its answer should be questioned; after many years in BI, that separation feels natural to me because a fluent answer and a trusted answer are not the same thing.

Working this way changed how I think about performance more generally, because the visible result and the human capability behind it no longer move in lockstep. If AI can improve the output much faster than a person's judgment develops, then a polished report, a strong analysis or a fast answer may tell us less about the person behind it than it did a few years ago, which leaves managers with a harder question than whether employees are using AI effectively: when AI does part of the work, what exactly are we measuring when we measure the individual?

When the output stops telling the whole story

Take a fairly ordinary analytical task in which a manager wants to know why margin fell in one business unit. One analyst uses AI to inspect the data, identify the major movements and prepare an explanation quickly, while another takes longer because something in the result does not fit what she knows about the business, so she checks the classifications and discovers that a change in the underlying data has distorted part of the comparison; the first output may look cleaner and arrive earlier, yet the second person may have contributed the judgment that prevented the wrong decision.

On another day, the opposite could easily be true, because the first analyst may understand the business just as deeply and simply know how to use AI well enough to remove hours of routine work, while the second may be slower because they are still doing manually what no longer needs to be manual. That is why neither speed nor AI usage tells us enough on its own; the more interesting distinction is what each person added to the result, especially when that contribution is hidden inside the way the problem was framed, the assumptions that were challenged, the context that was supplied and the decision to accept or reject what the system produced.

For years, the finished work gave managers a useful, if imperfect, signal about the person behind it because strong work usually reflected accumulated knowledge, technical ability, familiarity with the business and a growing capacity to handle exceptions. AI does not remove that relationship, although it can weaken it by allowing people with very different levels of experience to reach a similar visible standard, so good work remains valuable while becoming less informative as a shortcut for judging the capability underneath it.

This matters to BI as much as it matters to HR because performance systems naturally prefer what can be observed and counted. We can measure cases closed, reports delivered, sales results, cycle time or error rates with increasing precision, yet the number alone cannot tell us whether a fast result came from genuine expertise, good use of AI, uncritical acceptance of a generated answer or careful review that happened to leave no visible trace; the metric can be accurate while the interpretation of the person is wrong.

What speed is actually telling us now

There is good evidence that AI can make people more productive, but the effect changes considerably depending on who is doing the work and what kind of work they are doing. A large NBER study of 5,179 customer-support agents found an average productivity increase of about 14 percent, with much larger gains among less experienced and lower-skilled workers, and the researchers found evidence that the system was helping newer employees absorb some of the practices of stronger agents more quickly.

I find that result more interesting than a simple productivity headline because it suggests that AI can compress part of the old experience curve. A person with only a few months in a role may suddenly perform closer to someone who previously needed much longer to reach the same visible level, and that is a real benefit if the employee is learning along the way; from a management perspective, however, it also means that tenure, output and apparent independence begin to separate from one another sooner than they used to.

Research in other settings makes the picture more complicated because speed cannot be read in only one direction. Harvard and BCG found that consultants using GPT-4 became faster and produced stronger work on tasks that sat within what researchers called the *jagged technological frontier*, yet performance deteriorated when the work moved outside that frontier, while METR later found that experienced open-source developers working in repositories they knew well took 19 percent longer when early-2025 AI tools were available even though they believed the tools had made them faster.

Those studies involve different professions, different tasks and different generations of AI, so I would not combine them into a universal claim that AI makes experts slower or juniors faster. What they show more usefully is that speed now depends on the relationship between the person, the task and the tool, which means a manager who rewards throughput without understanding that relationship can easily reward the wrong behavior; sometimes expertise produces dramatic acceleration, while in another situation the same expertise produces more checking because the experienced person sees reasons to distrust an answer that looks perfectly acceptable to someone else.

The junior employee may look more senior before becoming more senior

This is where the subject becomes less about technology and more about people development. Early in a career, work has always been part of the learning process itself because simpler tasks expose people to the language of the business, recurring exceptions, customer behavior, technical constraints and the small mistakes that gradually build pattern recognition; many of those tasks are now exactly the ones AI can help complete almost immediately.

That can be a very good thing when the technology accelerates learning rather than replacing it. The customer-support research is encouraging here because there was evidence that newer workers were not merely producing more while the assistant was present, they were also retaining some of what they had learned, which suggests that AI can shorten the route to competence when it transfers useful patterns instead of simply hiding the work from the employee.

The management challenge is that the visible output may improve before the underlying judgment is mature enough to be obvious. A junior employee can now produce a management summary, an analysis or a piece of code that looks far more experienced than their tenure would normally suggest, and a manager who judges development mainly through the polish of the final result may miss whether the person can explain the reasoning, handle an unfamiliar exception, recognize a weak assumption or adapt when the AI has no useful pattern to offer.

From the employee side, the question is just as personal because better output can create the feeling that capability is improving at the same pace. Real development becomes easier to see when the context changes, the obvious pattern no longer fits or the AI produces an answer that needs to be challenged, and those moments may tell us more about professional growth than another polished first draft ever could; what matters is whether the person can carry the thinking into a situation where the machine has less to offer.

I would be equally uncomfortable with the opposite response, where managers deliberately restrict AI so they can observe what someone can do unaided. That would preserve an old measurement environment at the cost of real productivity, and it would ignore the fact that working effectively with AI is itself becoming part of professional capability; the better question is whether the employee is using the technology in a way that expands their own understanding or simply borrowing a level of output they cannot yet support when the situation changes.

What managers may need to see differently

Microsoft's 2026 Work Trend Index gives an interesting view of how workers themselves are experiencing this shift. In a survey of 20,000 AI users across 10 countries, 58 percent said they were producing work they could not have produced a year earlier, while 86 percent said they treated AI output as a starting point and remained responsible for the thinking behind it; quality control and critical thinking also ranked among the human capabilities respondents believed were becoming more important as AI took on more execution.

I would not treat a vendor survey as a final model for performance management, although the direction matches what I see in analytical work. When execution becomes easier, more of the human value can move into understanding what the problem actually is, bringing the missing context, knowing where the technology becomes unreliable, deciding when further investigation is worth the time and taking responsibility when an analytical answer turns into a real business action; these activities have always mattered, but they were easier to overlook when they were mixed with hours of visible manual work.

This is also where my BI instincts make me cautious about solving the problem with more metrics. A dashboard can show output, quality, cycle time, customer results and other measurable outcomes, and those numbers remain useful, but adding prompt counts, AI-usage percentages or another productivity score would not suddenly reveal the quality of the thinking behind them; if anything, we would risk measuring the new activity instead of understanding the contribution.

Performance conversations may therefore need to move closer to the reasoning behind the work without turning into surveillance. A manager can learn a great deal from asking what changed during the analysis, which assumption created uncertainty, where AI helped, where it did not and what the employee would do differently next time, because those conversations reveal whether the person is developing judgment rather than merely becoming efficient at obtaining polished output.

That does not mean every job needs a philosophical discussion after every task, and measurable results should remain part of performance because businesses still need people to deliver. It means that when AI changes the amount and type of execution required, the context around those results matters more, especially if we are making decisions about development, promotion, responsibility or who is ready to handle work where the answer is not already waiting inside a model.

I wrote before that data and AI would move the workforce from task execution toward outcome creation, and I would still make that argument today, although I would add one more layer now: once AI participates heavily in producing the outcome, good performance is no longer only about what was produced; it is also about the human contribution that made the result useful, reliable and worth acting on. That addition changes the management question because an outcome can be excellent while the capability behind it is still developing, or it can look ordinary even though someone made the judgment that prevented a much worse result.

That makes performance harder to reduce to one clean number, which may be uncomfortable for anyone who has spent years trying to make management more measurable. Still, I would rather accept that complexity than reward the person who produces the most visible work, overlook the person who quietly prevents the wrong decision, or confuse AI-supported polish with capability that is still developing, because as the machine makes more of the output look easy, understanding the person behind it becomes more important, not less.