Ten Reasons Why We Won’t See Productivity Improvements from GenAI
At least not measurable ones...
Source: Ben Hickey, The Economist
I have co-authored several articles over the last year or two that together suggest organizations are not likely to get business value with genAI if their primary focus is improving individual productivity. But everybody doesn’t read Harvard Business Review, and I have never put all the reasons for this in one place. So here goes—a laundry list of explanations for why you’re probably not going to achieve measurable productivity gains from genAI.
You need to redesign end-to-end processes around AI capabilities if you’re going to get substantial value from AI. And that, as I wrote about a couple of months ago, is hard. Even if you do it successfully it’s going to take years.
We’re not training people well. I saw a good summary of this argument in Robert Scoble’s newsletter last week. To get productivity from genAI demands that individuals incorporate AI into their daily work. But if you only get generic education—unrelated to your industry or your function—that’s unlikely. AI education should be highly segmented, with each business and job class getting their own version if it.
If you use genAI in the “right way,” you don’t save a lot of time and effort. The right way, at least if you want to create high-quality content and avoid rotting your brain, is to employ multiple prompts until you get something close to what you want, and then to review it for possible hallucinations. Then edit it to take out the cliché phrases and to add your unique perspective. When I insist that my students do this with their AI-assisted essays, they often tell me it doesn’t save them any time over writing it all themselves.
Measuring the aggregate impact of individual productivity gains is really hard. You have to measure how long it took you or someone else to perform a task pre-AI. Then you have to measure post-AI. Then you have to do that for multiple tasks and determine how all that rolls up to entire jobs and organizations. Not surprisingly, few organizations do this type of measurement.
It’s unclear what people are doing with the saved time. Let’s assume that you do the needed measurement and figure out that the average employee in a particular job saves an hour a day. What are those people doing with that saved hour? Maybe they’re doing more work, or maybe they are streaming a World Cup game. Theoretically you could lay off 1/8 of the people who do the job, but that’s hardly ever done. Most organizations aren’t confident enough about the measurements to take that step.
Token costs are going up, which makes getting productivity improvements harder. Now that tokens are being measured and billed by vendors more assiduously, to claim any real productivity gains you need to compare the savings from any labor time reductions to the cost of the tokens needed to perform that task. Again, it’s a measurement hassle that hardly anyone undertakes.
To really get productivity requires that we explore all the capabilities of genAI to transform our jobs. But most people don’t invest a lot of time in personal infrastructure development. I read a lot about individuals who create agents to do much of their work. But I don’t know any such people and I strongly suspect there aren’t too many of them. It takes a substantial investment in time, tokens (see above), and learning to automate significant aspects of your life. Most of us would rather watch the World Cup, Desperate Housewives, or YouTube videos in our spare time.
Organizations don’t do controlled experiments. To really figure out whether, for example, genAI improves the productivity (and quality) of blog post creation, you should do an experiment. You might assemble a group of bloggers who use a lot of genAI help to write their posts, a group that just uses it for brainstorming, and a control group that doesn’t use it at all. You then measure their average time to produce outputs and have some independent group score the output quality. Again, a few university researchers have done this sort of experiment, but they are rare in real organizations.
Poor genAI outputs lower productivity. This is the “workslop” problem that many observers have begun to notice and write about. Your own productivity may not suffer when you use AI to create workslop, but other people down the line may need to improve your output, and their productivity will suffer. Matthias Hollweg and I recently described the “process slop” problem, in which interorganizational use of AI leads not only to productivity loss, but reduced trust in the integrity and value of the entire process.
AI agents will help with the individual productivity problem, but probably won’t solve it. Teams of agents are typically for enterprise use cases more than individual ones. Even then, somebody is going to have to monitor the quality of agents’ work, and that will detract from any productivity gains. It’s true that some aggressive individual users are already adopting agents, but their efforts will run into some of the same issues involved in enhancing productivity that I’ve described above.
What does this mean for our AI-driven economy? It probably indicates that AI providers are over-valued by the market, that GDP growth that’s largely driven by data center construction is unhealthy, and that generative AI is not enough to power the economy on its own.
These factors may also mean that we are unlikely to see substantial productivity gains from AI as an overall economy. Thus far we certainly haven’t. Over the last seven years—half of which period includes the genAI era—average annual nonfarm productivity in the US (as measured by the Bureau of Labor Statistics) has increased by 2.1%. That is exactly the amount that productivity increased between 1947 and 2026. In the first quarter of 2026 productivity increased by only 0.3%. If genAI is increasing our productivity, it’s apparently not easy to measure at the overall economy level. Perhaps it’s just too early for us to see AI-driven productivity gains, but how long do we need to wait?



This resonates in so many levels:
- Process transformation was already hard before GenAI. Few do it consequentially.
- Tutorials look great on YouTube, but they fail in a business environment. Governance, security, and knowledge management are needed, but hardly covered.
- Expert skills remain scarce, because few have done the AI work in a business. Don’t confuse influencers with actual experts.
- Instead of productivity, measuring KPIs and PPIs is the better approach that is actually tied to a goal that business leaders measure, too, and care about.
- Business users are not developers. The bar is pretty high as the agentic platforms are not exactly self-service or provide the dead simple, reliable results that would be needed. Discoverability and interoperability remain behind expectations.
These are just a few of my thoughts on this topic…
"Perhaps it’s just too early for us to see AI-driven productivity gains, but how long do we need to wait?" Maybe longer than we expect given our history of mis-understanding technology trends. I met Hal Varian for the first time in late 1999 when he came to Boston to participate in the meeting of a committee charged with solving the "productivity paradox." I told him they were wrong to start measuring the gap between spending and the (expected) increase in productivity from the advent of PCs in the early 1980s. I argued that the real productivity impact occurred much later with the implementation and widespread use of local area networks, getting so much more from PCs by connecting them. That took time--when the first LAN was installed at NORC (in 1985) and I asked why do we need it, the answer was: So we can print on the laser printer in Dick's office (rather than walk a few yards and stick a floppy disk in his PC). It took another ten years for organizations to find and adopt productivity-enhancing uses of LANs, resulting in the increase in productivity in the second half of the 1990s.