Research, translated
We are poor witnesses to our own productivity
Sixteen experienced developers thought AI would make them faster. The measurement showed the opposite. And afterwards, they still thought so.
You have probably received a document like this. It looks right. The structure is there, the paragraphs are well formed, nothing is outright wrong. But when you have finished reading it, you do not know much more than before — and you spend the next hour working out what the sender actually meant.
American researchers have given the phenomenon a name: workslop. BetterUp Labs and Stanford Social Media Lab surveyed 1,150 desk workers in the United States in autumn 2025. Four in ten had received such content in the past month. The average time spent clearing up each incident was two hours.
That is an irritating finding, but not a surprising one. The surprising part comes now.
The number that is about us
The research institute METR ran a controlled experiment in 2025 with sixteen experienced developers. Not beginners — people who had contributed for years to large codebases they knew well. They worked through 246 issues, half with AI tools, half without.
It is that last part I cannot get out of my head. They did not guess wrong beforehand and then correct themselves. They guessed wrong, experienced the opposite, and believed the same thing afterwards.
Why we miss
This is not a story about AI not working. It is a story about how poor our instrument is when we try to measure our own work.
We rarely judge our own productivity in hours. We judge it in effort. Work that was heavy felt time-consuming; work that went easily felt fast. That is a reasonable rule of thumb across most of working history, because effort and time largely moved together.
AI breaks that link. The heavy part disappears. What remains is reading, checking and correcting — and we register that kind of work poorly, because it does not feel like production.
The blank page, the first sentence, getting started at all — all of it goes. What is left is the tidying up, and tidying up we do not count. So: less effort, the same or more time. And we report the effort.
Harvard Business Review put it in February this year as AI not reducing work but intensifying it. I think it is more precise to say the work moves — from making to verifying. And that we have not become any better at seeing the work that comes after.
What we do not actually know
The METR study has sixteen participants. That is few. The researchers say so themselves, loudly and first, along with the likelihood of a selection bias in who volunteered. They later found problems with their own experimental design, changed it, and published that they had done so.
When they ran a survey of 349 users in May this year in which people estimated the benefit themselves, and the median landed at 1.4 to 2 times more value, they wrote in the same breath that there is reason to be sceptical — and that it is a live possibility people report higher figures than they would on reflection.
The workslop figures are from American desk workers. I do not know whether four in ten holds in Norway, and I have not seen anyone measure it. And all of this concerns writing and programming. Whether it applies to casework, teaching or clinical work, we do not know.
This is how you read a study: not only what it found, but how well it knows it, and what the people who did it think about the matter themselves.
What we can say is that the distance between measured and perceived effect is large enough, and well enough documented, to worry anyone who has promised someone a gain they have not measured.
What it means for anyone doing the sums
Most organisations now estimating what AI has saved them do it by asking their employees. That is the cheapest method, and after this it is also the least reliable — it measures precisely the variable we know is skewed.
If you want to find out whether it actually works where you are, here are three things that cost very little:
-
Measure throughput time, not perceived time spent.
How long does a case take from arrival to completion — including the rounds of correction? Not how long someone felt they were working on it.
-
Count the work that comes after.
How many rounds does a document go through before it is good enough? That is the interesting number, and almost nobody records it.
-
Ask the recipient, not the sender.
Whoever made something with AI is systematically the wrong person to ask how good it turned out. Whoever received it knows a great deal more.
A small matter of language
Norwegian has no word for workslop. Some have probably tried already — I have not found a suggestion that has stuck.
My own suggestion is skinnarbeid: work that looks like work. The form already exists in Norwegian, in skinnhellig (sanctimonious) and skinnprosess (show trial), and it points at what the problem actually is — not that the content is bad, but that it passes itself off as something.
It is not a large matter. But it is easier to agree to stop doing something when you have a word for it.
Sources
- METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (July 2025), plus an update on experimental design (February 2026) and a survey (May 2026). Main study ↗
- BetterUp Labs and Stanford Social Media Lab, survey of 1,150 American desk workers, September 2025. betterup.com/workslop ↗
- AI Doesn't Reduce Work — It Intensifies It, Harvard Business Review, February 2026.